← Personal projects Machine learning Graduate research · 2025–26

From DCGANs to Vision Transformers

A run through the modern vision stack, implemented rather than imported: DCGANs for generation, Fast R-CNN detection on PASCAL VOC, UNet segmentation, and ViT classification. Written to understand the architectures rather than to hit a leaderboard.

  • PyTorch
  • TensorFlow
  • CUDA
  • NumPy
4
Architectures built
PASCAL VOC
Detection benchmark
Generation, detection, segmentation, classification
Domains

Architecture

  1. 01DCGANadversarial generation
  2. 02Fast R-CNNdetection on VOC
  3. 03UNetdense segmentation
  4. 04ViTattention-based classification

Why build them instead of calling them

You can solve most vision tasks in a dozen lines by loading a pretrained checkpoint. That’s usually the correct engineering decision and it teaches you almost nothing about why the model works or how it fails.

These four were chosen because together they cover the actual conceptual span of the field — generative versus discriminative, sparse output versus dense, convolutional inductive bias versus learned attention — and because each one fails in a characteristic, instructive way.

DCGAN — generation, and the instability underneath it

Adversarial training is the only one of the four where the loss going down tells you almost nothing. Generator and discriminator losses move against each other, so a “converging” curve is as likely to mean one network has overpowered the other as it is to mean anything good.

What actually mattered: the architectural constraints in the original DCGAN paper are not stylistic. Strided convolutions instead of pooling, batch norm in both networks, and bounded activations exist specifically to keep the two networks in a fight neither one wins outright. Relaxing them produces mode collapse quickly and visibly — the generator finds one output that fools the discriminator and stops exploring.

The lesson that generalized: when a training procedure has no reliable scalar to watch, you need to be looking at samples throughout, not at the end.

Fast R-CNN — where the speedup actually came from

Detection on PASCAL VOC. The interesting part of Fast R-CNN isn’t the detection head, it’s the insight that made it fast: R-CNN ran the convolutional backbone once per region proposal, which is enormously wasteful because the proposals overlap. Fast R-CNN runs the backbone once over the whole image and then pools features per region out of that shared map.

Implementing the RoI pooling step is where that stops being an abstract idea. The quantization in mapping a proposal’s coordinates onto a coarser feature grid is a real source of misalignment, and it’s exactly what RoI Align later fixed.

Multi-task loss — classification and box regression trained jointly — was the other thing worth building by hand, because the balance between the two terms visibly changes what the model prioritizes.

UNet — dense prediction and the case for skip connections

Segmentation needs a per-pixel answer, which means recovering spatial precision that downsampling threw away. The encoder-decoder shape alone doesn’t do it — upsampled features are semantically rich and spatially vague.

The skip connections are the whole architecture. Concatenating high-resolution encoder features into the decoder at matching depth is what puts the edges back. Ablating them is the fastest way to see it: the model still finds the right region and completely loses the boundary.

ViT — what happens when you drop the inductive bias

A Vision Transformer treats an image as a sequence of patches and uses no convolution at all. The tradeoff is stark and shows up immediately: without the locality and translation-equivariance assumptions baked into a CNN, ViT has to learn them, and learning them takes far more data.

On a small dataset a comparable CNN wins comfortably. That’s not a mark against ViT — it’s the clearest practical demonstration I’ve encountered of what an inductive bias actually buys you, and why “the architecture with fewer assumptions” is only better past a certain data scale.

What I took from it

Four architectures, four different failure modes, and one pattern: in each case the component that looks like an implementation detail — the normalization choice, the pooling quantization, the skip connection, the patch embedding — is the component doing the real work. Reading the paper tells you the architecture. Building it tells you which part is load-bearing.

Contact

Want the longer version?

Happy to walk through any of this in detail — the parts that broke are usually the interesting bit.