Advanced Concepts
Deep dive into influential research papers that shaped modern computer vision. Learn key concepts through interactive visualizations and animations.
Research Papers
Qwen3
Unified Thinking and Non-Thinking Modes, Thinking Budgets, and Strong-to-Weak Distillation in an Open Dense + MoE LLM Family
Reasoning and chat used to be separate models. Qwen3 folds both into one set of weights, switched per turn by /think and /no_think, and adds a thinking budget that emerges without being trained. Behind it: dense and 128-expert MoE models with QK-Norm, 36T tokens in 119 languages, a four-stage post-training pipeline, and strong-to-weak distillation at a tenth of RL's GPU hours. The 235B-A22B flagship scores 85.7 on AIME'24 and beats DeepSeek-R1 on 17 of 23 benchmarks.
The Llama 3 Herd of Models
A 405B Dense Transformer, Its Scaling Laws, 16K-GPU Training Stack, and SFT + Rejection Sampling + DPO Post-Training
Meta's 92-page recipe for an open frontier model, and its argument that no new architecture is needed. A standard dense Transformer (GQA, RoPE base 500K, 128K vocab) is sized by IsoFLOP scaling laws to 405B, trained on 15.6T curated tokens with 3.8e25 FLOPs across 16K H100s in 4D parallelism, stretched to 128K context in six gated stages, then aligned with six rounds of SFT, rejection sampling and DPO. It lands on par with GPT-4, and compositional adapters add vision and speech to the frozen LLM.
Qwen2.5-VL
A Flagship Vision-Language Model for Documents, Grounding, Long Video and GUI Agents
Today's large vision-language models are like the middle of a sandwich cookie — competent across tasks, exceptional at none. Qwen2.5-VL bets the missing base layer is fine-grained perception: a native-resolution ViT with window attention, MRoPE aligned to absolute time, and a 4.1T-token data effort. The 72B matches GPT-4o and Claude 3.5 Sonnet, leads on documents, and turns grounding into agency — a 1.6 → 43.6 leap on ScreenSpot Pro.
Dream-RSI
Recursive Self-Improvement through Evolving Worlds
When AI agents improve their own exploration, judging a strategy means paying for a whole long rollout — the meta-exploration bottleneck. Dream-RSI's insight: a completed discovery tree already stores every outcome, so it can be replayed as a zero-cost simulator. The agent "dreams" over these worlds to refine its policy off-policy, then redeploys it — competitive quality at up to 162x fewer discovery-agent calls across algorithm, math, and GPU kernel engineering.
Composing Continual Learning Mechanisms
Continual Learning Mechanisms Compose for Long-Horizon Memorization
Can a language model still recall what it learned 100 updates ago? Naive continual fine-tuning retains just 1.2%, and no single mechanism fixes it. Organizing forgetting into data/function/weight anchors and shared vs merged LoRA — then composing them — lifts average final retention to 34.9%, a 28-fold gain, driven by a super-additive replay × merged-LoRA core.
ZGCM-1
A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search
A fully open 7.39B model that couples deliberate thinking with tool use instead of memorizing the web. Hybrid gated sliding-window/global attention (6.4x KV cache, 3.94x throughput at 256K), a stable FP8 + Muon + TWEO stack (~4.2x time-to-loss), and progressive 16K→64K→256K MDP mid-training — first on average across 14 reasoning benchmarks at the 7B-8B scale.
WMRL — World Model RL
Scaling Automatic Research Agents via World Models
RL for AutoResearch agents is bottlenecked by execution, not generation: every solution needs its own sandbox. Replace it with a world model that predicts outcomes in a few forward passes, correct its bias and noise with a thin anchor stream — 3-4x faster training, and 4B/9B agents that beat off-the-shelf 48B/120B.
Agentic Game Development
A Verifiable Trajectory Data Engine for Scaling World Models
Spatial generation is capped by fuzzy proxies (CLIP, FVD, MLLM-as-judge). Game development supplies the missing reward engine: RLHEV pairs dense engine checks with human acceptance, and AWoMo turns build traces into training data — 0.681 primary on UnitySceneBench with cross-engine and embodied transfer gains.
UrbanGround
From Local Perception to Spatial Agency in a Real-Scale City
The first sandbox to test whether MLLM agents can turn local urban perception into sustained action, built as a physically constrained, georegistered replica of Hong Kong. Grounding holds, but navigation collapses past a few blocks — 75.0% short-range success falls to ~0% long-range across GPT-5.5, Claude-Opus-5, Gemini-3.6-Flash and Kimi-K3.
StudentSim
Training LLM-based Student Simulators
Microsoft Research decomposes student simulation into behavioral fidelity (F) and guidance responsiveness (R), then trains per-student simulators via a pooled-then-specialized LoRA pipeline. On chess, StudentSim reaches F=0.51 / R=0.91 vs GPT-5.4's 0.23 / 0.72 — and a frozen simulator reward yields the best-rated chess tutor in an expert human study.
Repo-To-Skill
Distilling GitHub Repositories Into AI4AI Skills
Operational knowledge is the missing layer beyond a model and a harness. DisCo distills repos and papers into verified skill graphs — the 5,000+-skill AREX-Skill Library — lifting a fixed GPT-5.5 Codex agent +134.3% on MLE-bench, +34.4% on PaperBench, +9.2% on FrontierCS and +14.0% on PassNet.
FlashAttention
Fast and Memory-Efficient Exact Attention with IO-Awareness
Attention is memory-bound, not compute-bound. Tiling plus an online softmax rescaling trick removes the N×N round trip to HBM entirely — 7.6x faster attention, up to 20x less memory, and exact rather than approximate.
Multi-View Foundation Models
Turning DINO, CLIP and SAM into Multi-View Consistent Variants
Inference-time 3D consistency without NeRF or Gaussian Splatting: multi-view adapters with Plücker ray embeddings cut DINOv2's location error from 0.1029 to 0.0247 while keeping 0.9376 cosine similarity to the base model.
EUPE
Efficient Universal Perception Encoder
Meta's two-stage knowledge-distillation recipe (scale up to a 1.9B proxy, then scale down) yields edge-deployable encoders that match same-size domain experts on classification, dense prediction, and VLM — all in 7–60ms on iPhone CPU.
TinyLoRA
Learning to Reason in 13 Parameters
Achieving 91% GSM8K accuracy with only 13 trainable parameters (26 bytes) through extreme parameter sharing and RL training.
Nemotron 3 Super
Hybrid Mamba-MoE for Agentic Reasoning
NVIDIA's 120B MoE hybrid Mamba-Attention model with LatentMoE, Multi-Token Prediction, and NVFP4 training — 2.2x faster than GPT-OSS-120B.
Attention Residuals
Learned Depth-Wise Attention for Residual Connections
Replacing fixed residual accumulation with softmax attention over depth, enabling content-dependent information routing across layers.
TurBoQuant
Near-Optimal Vector Quantization with Zero Overhead
Training-free, data-oblivious vector quantization achieving 6x KV cache compression and 8x attention speedup via polar coordinate transforms and 1-bit QJL error correction.
The Illusion of Thinking
Strengths and Limitations of Reasoning Models via Problem Complexity
Apple/NeurIPS 2025: controllable puzzle environments reveal three complexity regimes and a counter-intuitive thinking-token collapse in Large Reasoning Models.
YOLOv10
Real-Time End-to-End Object Detection
Eliminating NMS with consistent dual assignments and optimizing efficiency-accuracy tradeoffs through holistic design.
LLaVA
Visual Instruction Tuning
Connecting a CLIP vision encoder to a large language model via a simple projection layer, trained in two stages with GPT-4-generated instruction data.
BLIP-2
Bootstrapping Language-Image Pre-training with Frozen Encoders and LLMs
Freeze the image encoder, freeze the LLM, and train only a 188M-parameter Q-Former between them — beating Flamingo80B by 8.7 points on zero-shot VQAv2 with 54x fewer trainable parameters.
DINOv3
7B Parameters, Universal Visual Features
Self-supervised learning at scale with Gram Anchoring. 7B parameters trained on 1.7B images achieving SOTA on 60+ benchmarks.
SigLIP 2
Multilingual Vision-Language Encoders
Unified training recipe combining sigmoid loss, decoder objectives, self-distillation, and masked prediction for improved multimodal understanding.
Segment Anything (SAM)
Promptable Segmentation for Any Image
Foundation model for image segmentation that can segment any object using points, boxes, or masks as prompts.
3D Gaussian Splatting
Real-Time Radiance Field Rendering
Replaces neural radiance fields with explicit 3D Gaussians and tile-based rasterization, achieving real-time (>100 FPS) rendering with state-of-the-art visual quality.
Masked Autoencoders
Self-Supervised Visual Pre-training
Pre-training vision transformers by masking 75% of image patches and learning to reconstruct them.
CLIP
Contrastive Language-Image Pre-training
Learning visual representations from natural language supervision using contrastive pre-training on 400M image-text pairs.
Diffusion Models
Denoising Diffusion Probabilistic Models
Learning to generate images by reversing a gradual noising process, achieving state-of-the-art image synthesis.
Vision Transformer (ViT)
An Image is Worth 16x16 Words
Applying Transformers directly to image patches, achieving state-of-the-art results on image classification.
DETR
End-to-End Object Detection with Transformers
Treating object detection as a set prediction problem using Transformers and bipartite matching loss.
AdamW vs Adam
Decoupled Weight Decay Regularization
Why AdamW decouples weight decay from gradient updates, and when to use Adam vs AdamW in modern deep learning.
Attention Is All You Need
The Transformer Architecture
Introducing the Transformer architecture with self-attention mechanisms that revolutionized NLP and later computer vision.
Mask R-CNN
Instance Segmentation Framework
Extending Faster R-CNN with a parallel mask branch for pixel-precise instance segmentation.
Normalization Techniques
BatchNorm, LayerNorm, GroupNorm & InstanceNorm
Visual guide to normalization types: which dimensions each operates over and when to use which.
ResNet
Deep Residual Learning
Skip connections enabling training of very deep networks, solving the degradation problem in neural networks.
U-Net
Convolutional Networks for Biomedical Image Segmentation
Encoder-decoder architecture with skip connections for precise localization in image segmentation tasks.
More papers coming soon: Mask R-CNN, GANs, Diffusion Models, and more.