Loading...
VLM 1: Vision-Language Models — all 50 problems
Fifteen hands-on exercises covering the arithmetic that actually runs a vision-language model: patch tokenisation and image-token budgets, the CLIP and SigLIP contrastive objectives, retrieval and zero-shot transfer, bridge modules from linear projectors to the Perceiver Resampler, and the metrics for efficient adaptation and hallucination. Every problem takes its arrays as input - no vision encoder is ever run.
Bridge Modules: Projectors, Q-Former & Cross-Attention
- Q-Former Output Shape and Parameter CountEasy
- MLP Projector Parameter CountEasy
- Attention Pooling to a Single TokenEasy
- Scaled Dot-Product Cross-AttentionMedium
- Multi-Head Cross-Attention OutputMedium
- Gated Cross-Attention (Flamingo tanh Gate)Medium
- Learned Query Compression RatioMedium
- RMSNorm Before the ProjectorMedium
- Perceiver Resampler Forward PassHard
- One Perceiver Resampler BlockHard
Contrastive Objectives: CLIP and SigLIP
- L2-Normalize an EmbeddingEasy
- Contrastive Accuracy at the DiagonalEasy
- CLIP Contrastive (InfoNCE) LossMedium
- SigLIP Pairwise Sigmoid LossMedium
- Symmetric CLIP Loss with a Learnable TemperatureMedium
- SigLIP Loss with the Learnable BiasMedium
- Effective Negatives from Sharded Contrastive BatchesMedium
- Modality Gap Between Image and Text EmbeddingsMedium
- Softmax vs Sigmoid: Batch Dependence of the Match ProbabilityHard
- Hard Negative Mining in a Contrastive BatchHard
Efficient Adaptation & Hallucination Metrics
- LoRA Trainable Parameter SavingsEasy
- Attention FLOPs and Visual Token PruningEasy
- LoRA Adapter Parameter CountEasy
- POPE Object Hallucination AccuracyEasy
- POPE Precision, Recall, and F1Medium
- QLoRA Memory Footprint EstimateMedium
- Merge a LoRA Update into the Base WeightMedium
- CHAIR-i and CHAIR-s Hallucination RatesMedium
- CHAIR Hallucination RateHard
- Visual Token Pruning by Attention MassHard
Patch Embedding & Visual Tokens
- Patch Embedding Token CountEasy
- Patch Grid for a Resized ImageEasy
- CLS and Register Token Sequence LengthEasy
- AnyRes Tiling Token BudgetMedium
- Interleaved Multimodal Sequence LayoutMedium
- Patchify an Image into Flattened VectorsMedium
- Positional Embedding Interpolation CountMedium
- 2D Sine-Cosine Positional EmbeddingMedium
- Pixel-Shuffle Token DownsamplingMedium
- NaViT Sequence Packing of Variable-Size ImagesHard
Retrieval & Zero-Shot Transfer
- Cosine Similarity Top-K RetrievalEasy
- Zero-Shot Prediction from Class LogitsEasy
- Mean Reciprocal Rank for RetrievalEasy
- Recall@K for Image-Text RetrievalMedium
- Zero-Shot Classifier by Prompt EnsemblingMedium
- Prompt-Ensemble Zero-Shot Classifier WeightsMedium
- Recall@K and Median Rank TogetherMedium
- Temperature-Scaled Zero-Shot ConfidenceMedium
- Deduplicate a Retrieval Result ListMedium
- Weighted k-NN Zero-Shot TransferHard