Pittsburgh, Pennsylvania
Projects

GAN-Based Synthetic Portrait Generation with Pix2Pix

image
February 21, 2024
I wanted to teach a neural network to draw anime faces. Sounds simple enough. It is not. Anime faces break every rule that makes photorealistic face generation tractable. Eyes can take up a quarter of the face. Hair shows up in electric blue, cherry blossom pink, silver, violet, green. Proportions shift wildly between art styles. A model that generates believable anime faces has to learn not just anatomy but an entire artistic tradition. The approach: generative adversarial networks. Two neural networks competing against each other. A generator that learns to create images. A discriminator that learns to spot fakes. Neither one has an explicit definition of "good art." Quality emerges from the competition itself. I built three implementations. Two Deep Convolutional GANs (one in Keras, one in PyTorch) and a conditional Pix2Pix model. The goal was not just passable outputs. I wanted to understand how architectural decisions shape the creative range of a generative model and push a 64x64 pixel canvas to its limits. No pre-packaged dataset here. I built the training set from scratch through a multi-stage pipeline that turns raw anime video into cropped character faces. The first stage extracts frames from anime video files using concurrent processing. A single anime episode has roughly 34,000 frames at 24 fps. Processing dozens of episodes one at a time would take hours. The concurrent pipeline distributes video files across multiple threads, each reading frames, filtering for quality, and writing outputs. Throughput scales nearly linearly with CPU cores. Next comes face detection. Raw frames contain backgrounds, action sequences, text overlays. I used OpenCV's cascade classifier with lbpcascade_animeface.xml, a model trained specifically for anime-style faces. It understands the unique geometry of drawn faces: oversized eyes, simplified noses, distinctive jawlines. The classifier scans each frame at multiple scales, sliding a detection window and applying increasingly strict feature tests. Only regions that pass every stage survive. After detection, the pipeline filters out partial faces, tiny detections, and heavily occluded results. Surviving faces get resized to 64x64 pixels. That resolution is a deliberate tradeoff: large enough to capture eyes, hair, expression, and skin tone, small enough to allow extensive experimentation without blowing up GPU costs. Pixel values get normalized from [0, 255] to [-1, 1]. This lines up with the tanh activation on the generator's output layer, so the model learns a direct mapping without scale mismatches. The final dataset is thousands of 64x64x3 RGB images covering a wide range of anime character designs, hair colors, eye styles, and expressions. The generator starts with noise and tries to make a convincing face. The discriminator looks at an image and decides: real or fake. They train together. The generator gets better at fooling the discriminator. The discriminator gets better at catching fakes. Round and round. Early in training, the generator outputs pure static. Random pixels, no structure. The discriminator catches everything easily. Then the generator starts discovering patterns. It learns that faces have consistent color palettes. It finds the oval shape of a head, the dark regions where eyes go, the varied colors of hair. Each improvement forces the discriminator to look harder, checking edge quality, color consistency, anatomical plausibility. The beautiful part: nobody programs the rules. The rules emerge from competition, refined over millions of exchanges. Two failure modes haunt every GAN project. Mode collapse is when the generator finds a handful of outputs that fool the discriminator and stops exploring. It just makes the same face over and over. Training divergence is worse: if one network gets too far ahead of the other, the whole thing falls apart. Keeping the game balanced is the real challenge.
DCGAN adversarial loop: a 100-dim noise vector runs through the generator (dense, transposed convs, BatchNorm/ReLU, tanh) to a fake 64×64×3 image; the discriminator judges real versus fake and its loss flows back to train the generator
The Generator takes a 100-dimensional noise vector sampled from a standard normal distribution. One hundred random numbers that encode everything about the face that will appear. This vector passes through a dense layer, gets reshaped into a small spatial feature map, then flows through four transposed convolution blocks. Each block doubles the spatial resolution, progressively building structure from coarse shapes to fine details. Batch normalization stabilizes each layer. ReLU activations keep gradients flowing. The final layer uses tanh to produce 64x64x3 RGB output in the [-1, 1] range. The Discriminator mirrors the generator in reverse. It takes a 64x64x3 image and pushes it through strided convolution layers. Each layer halves spatial resolution while increasing feature channels, compressing the image into an abstract "real or fake" judgment. Batch normalization appears everywhere except the first layer (where it would mess with learning low-level input statistics). LeakyReLU with a 0.2 negative slope prevents dying neurons, which is a real problem for discriminators during adversarial training. Two critical design choices from the DCGAN paper. First: no pooling layers anywhere. Strided convolutions in the discriminator and transposed convolutions in the generator replace all pooling ops. The network learns its own downsampling and upsampling instead of relying on fixed operations. Second: no fully connected hidden layers. The architecture is convolutional end to end, which preserves spatial relationships throughout. The only dense layer is at the generator's input, mapping the noise vector into a spatial feature map. Keras DCGAN was the foundation build. High-level abstractions made the architecture quick to define and easy to modify. Good for iterating on ideas fast. PyTorch DCGAN reproduced the same architecture but exposed more of the training mechanics: manual gradient zeroing, explicit forward passes, direct loss computation. Same model, deeper understanding. One key addition was explicit weight initialization. All conv and batch norm weights initialized from Normal(0, 0.02). This sounds minor but has an outsized impact on stability. Too-large weights cause exploding gradients on the first batch. Too-small weights cause vanishing signals. The 0.02 standard deviation is a sweet spot the GAN community found through extensive experimentation. Having both side by side was genuinely useful. The Keras version taught the architecture. The PyTorch version taught the mechanics. Pix2Pix moved into conditional generation. Instead of creating faces from noise, it learns to translate between paired image domains. The U-Net generator features skip connections linking encoder layers to their matching decoder layers. Fine spatial details bypass the bottleneck and flow directly to the output, preserving high-frequency information that would otherwise get lost. For faces, this means precise alignment of eyes, eyebrows, and hair boundaries. The PatchGAN discriminator evaluates images as a grid of overlapping 70x70 patches, each getting an independent real/fake classification. This forces the generator to produce consistent textures everywhere, not just where the discriminator happens to look. The loss function combines adversarial loss with L1 reconstruction. Adversarial loss pushes toward realism. L1 anchors the output to ground truth. Too much L1 and you get blurry-but-safe results. Too much adversarial loss and the model hallucinates sharp-but-wrong details. Balancing the two is part engineering, part taste. The optimizer across all three models is Adam with a learning rate of 0.0002 and beta1 of 0.5. Both values are deliberately conservative. Large learning rates in adversarial training cause the two networks to overshoot each other, leading to oscillating losses and eventual divergence. The lower beta1 (compared to the typical 0.9) reduces momentum, which can amplify oscillations in adversarial setups. Batch sizes ranged from 128 to 256. The Keras DCGAN trained for 15,000+ epochs. The PyTorch DCGAN hit comparable quality in 50 epochs, reflecting differences in how epochs were defined and optimizations in the training loop. Binary cross-entropy loss drove both networks: the discriminator minimizes classification error on real and fake images, the generator maximizes the discriminator's error on fakes. Stability looked like this: discriminator accuracy hovering around 85%. That number matters. Too close to 100% means the discriminator is dominating and the generator can't learn. Too close to 50% means the adversarial signal is too weak. 85% is the sweet spot where the discriminator provides useful gradient information without crushing the generator. I tracked training by generating faces from the same noise vectors at regular intervals. Early on, blobs. Midway through, recognizable features. By the end, distinct characters. Comparing these timelines across hyperparameter configurations revealed which setups converged faster and produced better final quality. No mode collapse observed. Output batches of 64 faces showed 64 distinct identities: different hair colors, eye shapes, expressions, overall aesthetics. The generator learned a meaningful mapping from noise space to face space rather than memorizing a few safe outputs. PSNR came in at 28.5 dB on Pix2Pix outputs. For context, 30+ dB is considered high quality for natural images. 28.5 dB for synthetic anime faces is solid given the complexity of the task. SSIM hit 0.91, which means the generated faces preserve the structural relationships that make a face recognizable: relative feature positions, contrast patterns that define edges, luminance gradients that create depth. The outputs showed real diversity. Hair ranged from short and spiky to long and flowing. Colors covered the full anime palette: natural blacks and browns, vivid blues, pinks, greens. Eye designs varied from large and expressive to narrow and intense. The DCGAN models produced the most raw variety, sampling freely from the learned distribution. Pix2Pix outputs were more structurally consistent thanks to the skip connections and conditional training signal, but operated in a narrower creative range. Where things got rough: fine facial details. At 64x64, each eye occupies maybe 8 to 12 pixels across. That is a tiny canvas for the complexity anime eyes demand. Iris patterns blur together. Characteristic eye highlights merge or shift. Mouths and noses sometimes look smudged. Hair boundaries, where strands meet skin or background, can show a soft halo instead of the clean line you'd expect from hand-drawn art. These aren't architectural failures. They're resolution constraints. The generator receives 100 random numbers and produces 12,288 pixel values (64 x 64 x 3 channels). That's a 122x information expansion where every output value must be spatially consistent with every other. The nose has to sit between the eyes. The hair has to frame the face. Colors have to cohere. Doing that from a 100-dimensional seed is a remarkable feat of learned compression. StyleGAN is the obvious next step. It separates the latent space into style vectors that control different aspects at different scales. Coarse styles handle head shape and pose. Medium styles control features and hair. Fine styles handle color and micro-details. You could build a tool where an artist adjusts sliders for hair length, eye color, and expression, and the model generates a matching character in real time. Progressive growing would solve the resolution ceiling. Start training at 4x4, gradually double the resolution. Each stage builds on the stable foundation of the previous one. The model learns coarse structure first, then medium detail, then fine detail. This could push outputs from 64x64 to 256x256 without the instabilities that come from training at high resolution directly. Super-resolution post-processing is the pragmatic alternative. Take the 64x64 outputs and run them through a dedicated upscaling network like Real-ESRGAN. Anime-specific variants already exist. This separates the problem of learning face structure from the problem of learning fine detail, making each part more tractable. Higher resolution, better control, sharper details. The foundation is here. The data pipeline works. The training intuition is built. The next version would be significantly more capable. Python, TensorFlow/Keras, PyTorch, OpenCV, NumPy. OpenCV handled the anime face detection pipeline with the lbpcascade cascade classifier. Training ran on GPU with Adam optimization across all three architectures.