PIXELBANKv9.1.0
Menu
Interactive Learning from Research Papers

Advanced Concepts

Deep dive into influential research papers that shaped modern computer vision. Learn key concepts through interactive visualizations and animations.

37 Papers
Interactive Demos
Research Insights

Research Papers

Large Language Models

Qwen3

Unified Thinking and Non-Thinking Modes, Thinking Budgets, and Strong-to-Weak Distillation in an Open Dense + MoE LLM Family

Reasoning and chat used to be separate models. Qwen3 folds both into one set of weights, switched per turn by /think and /no_think, and adds a thinking budget that emerges without being trained. Behind it: dense and 128-expert MoE models with QK-Norm, 36T tokens in 119 languages, a four-stage post-training pipeline, and strong-to-weak distillation at a tenth of RL's GPU hours. The 235B-A22B flagship scores 85.7 on AIME'24 and beats DeepSeek-R1 on 17 of 23 benchmarks.

One Model, Two ModesThe Qwen3 Model Family — Dense and MoEPre-training — 36T Tokens in Three Stages+7
Start Learning
Large Language Models

The Llama 3 Herd of Models

A 405B Dense Transformer, Its Scaling Laws, 16K-GPU Training Stack, and SFT + Rejection Sampling + DPO Post-Training

Meta's 92-page recipe for an open frontier model, and its argument that no new architecture is needed. A standard dense Transformer (GQA, RoPE base 500K, 128K vocab) is sized by IsoFLOP scaling laws to 405B, trained on 15.6T curated tokens with 3.8e25 FLOPs across 16K H100s in 4D parallelism, stretched to 128K context in six gated stages, then aligned with six rounds of SFT, rejection sampling and DPO. It lands on par with GPT-4, and compositional adapters add vision and speech to the frozen LLM.

Data, Scale, and Managing ComplexityA Deliberately Standard Dense TransformerCompute-Optimal Sizing of the 405B+6
Start Learning
Vision-Language

Qwen2.5-VL

A Flagship Vision-Language Model for Documents, Grounding, Long Video and GUI Agents

Today's large vision-language models are like the middle of a sandwich cookie — competent across tasks, exceptional at none. Qwen2.5-VL bets the missing base layer is fine-grained perception: a native-resolution ViT with window attention, MRoPE aligned to absolute time, and a 4.1T-token data effort. The 72B matches GPT-4o and Claude 3.5 Sonnet, leads on documents, and turns grounding into agency — a 1.6 → 43.6 leap on ScreenSpot Pro.

The Fine-Grained Perception GapNative-Resolution ViTMRoPE Aligned to Absolute Time+5
Start Learning
AI Research Agents

Dream-RSI

Recursive Self-Improvement through Evolving Worlds

When AI agents improve their own exploration, judging a strategy means paying for a whole long rollout — the meta-exploration bottleneck. Dream-RSI's insight: a completed discovery tree already stores every outcome, so it can be replayed as a zero-cost simulator. The agent "dreams" over these worlds to refine its policy off-policy, then redeploys it — competitive quality at up to 162x fewer discovery-agent calls across algorithm, math, and GPU kernel engineering.

The Meta-Exploration BottleneckHistory as a Replay SimulatorDiscovery Trees & the Decision Interface+4
Start Learning
Continual Learning

Composing Continual Learning Mechanisms

Continual Learning Mechanisms Compose for Long-Horizon Memorization

Can a language model still recall what it learned 100 updates ago? Naive continual fine-tuning retains just 1.2%, and no single mechanism fixes it. Organizing forgetting into data/function/weight anchors and shared vs merged LoRA — then composing them — lifts average final retention to 34.9%, a 28-fold gain, driven by a super-additive replay × merged-LoRA core.

Long-Horizon MemorizationThree AnchorsLow-Rank Allocation+3
Start Learning
Efficient Foundation Models

ZGCM-1

A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search

A fully open 7.39B model that couples deliberate thinking with tool use instead of memorizing the web. Hybrid gated sliding-window/global attention (6.4x KV cache, 3.94x throughput at 256K), a stable FP8 + Muon + TWEO stack (~4.2x time-to-loss), and progressive 16K→64K→256K MDP mid-training — first on average across 14 reasoning benchmarks at the 7B-8B scale.

The Scale & Opacity BarriersHybrid Attention ArchitectureFP8 + Muon + TWEO Co-Design+5
Start Learning
Agentic Reinforcement Learning

WMRL — World Model RL

Scaling Automatic Research Agents via World Models

RL for AutoResearch agents is bottlenecked by execution, not generation: every solution needs its own sandbox. Replace it with a world model that predicts outcomes in a few forward passes, correct its bias and noise with a thin anchor stream — 3-4x faster training, and 4B/9B agents that beat off-the-shelf 48B/120B.

The Execution BottleneckWorld Model as EnvironmentThe Price: Bias & Noise+4
Start Learning
World Models

Agentic Game Development

A Verifiable Trajectory Data Engine for Scaling World Models

Spatial generation is capped by fuzzy proxies (CLIP, FVD, MLLM-as-judge). Game development supplies the missing reward engine: RLHEV pairs dense engine checks with human acceptance, and AWoMo turns build traces into training data — 0.681 primary on UnitySceneBench with cross-engine and embodied transfer gains.

The Verifiability BottleneckHuman-Engine VerificationAWoMo — The Agentic World Model+4
Start Learning
Embodied Agents

UrbanGround

From Local Perception to Spatial Agency in a Real-Scale City

The first sandbox to test whether MLLM agents can turn local urban perception into sustained action, built as a physically constrained, georegistered replica of Hong Kong. Grounding holds, but navigation collapses past a few blocks — 75.0% short-range success falls to ~0% long-range across GPT-5.5, Claude-Opus-5, Gemini-3.6-Flash and Kimi-K3.

The Spatial Agency ProblemClosed-Loop Urban InteractionA Real-Scale Hong Kong+4
Start Learning
AI Tutoring & Simulation

StudentSim

Training LLM-based Student Simulators

Microsoft Research decomposes student simulation into behavioral fidelity (F) and guidance responsiveness (R), then trains per-student simulators via a pooled-then-specialized LoRA pipeline. On chess, StudentSim reaches F=0.51 / R=0.91 vs GPT-5.4's 0.23 / 0.72 — and a frozen simulator reward yields the best-rated chess tutor in an expert human study.

The Student Simulation ProblemFidelity & ResponsivenessTwo-Stage Training Pipeline+3
Start Learning
AI Research Agents

Repo-To-Skill

Distilling GitHub Repositories Into AI4AI Skills

Operational knowledge is the missing layer beyond a model and a harness. DisCo distills repos and papers into verified skill graphs — the 5,000+-skill AREX-Skill Library — lifting a fixed GPT-5.5 Codex agent +134.3% on MLE-bench, +34.4% on PaperBench, +9.2% on FrontierCS and +14.0% on PassNet.

The Missing LayerSkills and Skill GraphsSkill Distillation+4
Start Learning
Efficient Attention

FlashAttention

Fast and Memory-Efficient Exact Attention with IO-Awareness

Attention is memory-bound, not compute-bound. Tiling plus an online softmax rescaling trick removes the N×N round trip to HBM entirely — 7.6x faster attention, up to 20x less memory, and exact rather than approximate.

The GPU Memory HierarchyWhat Standard Attention CostsTiling and the Online Softmax+4
Start Learning
3D Foundation Models

Multi-View Foundation Models

Turning DINO, CLIP and SAM into Multi-View Consistent Variants

Inference-time 3D consistency without NeRF or Gaussian Splatting: multi-view adapters with Plücker ray embeddings cut DINOv2's location error from 0.1029 to 0.0247 while keeping 0.9376 cosine similarity to the base model.

Viewpoint InconsistencyMulti-View AdapterPlücker Ray Embeddings+4
Start Learning
Efficient Vision Encoders

EUPE

Efficient Universal Perception Encoder

Meta's two-stage knowledge-distillation recipe (scale up to a 1.9B proxy, then scale down) yields edge-deployable encoders that match same-size domain experts on classification, dense prediction, and VLM — all in 7–60ms on iPhone CPU.

Scale Up Then Scale DownMulti-Teacher AggregationMulti-Resolution Finetuning+1
Start Learning
Parameter-Efficient Fine-Tuning

TinyLoRA

Learning to Reason in 13 Parameters

Achieving 91% GSM8K accuracy with only 13 trainable parameters (26 bytes) through extreme parameter sharing and RL training.

Sub-Rank AdaptationRank Below OneRL vs SFT Efficiency+2
Start Learning
Large Language Models

Nemotron 3 Super

Hybrid Mamba-MoE for Agentic Reasoning

NVIDIA's 120B MoE hybrid Mamba-Attention model with LatentMoE, Multi-Token Prediction, and NVFP4 training — 2.2x faster than GPT-OSS-120B.

Pipeline SummaryLatentMoE ArchitectureHybrid Mamba-Attention+4
Start Learning
Large Language Models

Attention Residuals

Learned Depth-Wise Attention for Residual Connections

Replacing fixed residual accumulation with softmax attention over depth, enabling content-dependent information routing across layers.

Pipeline SummaryPreNorm DilutionFull AttnRes+3
Start Learning
Quantization & Compression

TurBoQuant

Near-Optimal Vector Quantization with Zero Overhead

Training-free, data-oblivious vector quantization achieving 6x KV cache compression and 8x attention speedup via polar coordinate transforms and 1-bit QJL error correction.

Quantization OverheadPolarQuantQJL Error Correction+2
Start Learning
LLM Reasoning & Evaluation

The Illusion of Thinking

Strengths and Limitations of Reasoning Models via Problem Complexity

Apple/NeurIPS 2025: controllable puzzle environments reveal three complexity regimes and a counter-intuitive thinking-token collapse in Large Reasoning Models.

Pipeline SummaryBenchmark ContaminationControllable Puzzles+4
Start Learning
Object Detection

YOLOv10

Real-Time End-to-End Object Detection

Eliminating NMS with consistent dual assignments and optimizing efficiency-accuracy tradeoffs through holistic design.

Dual AssignmentsNMS-Free InferencePSA Module+4
Start Learning
Multimodal

LLaVA

Visual Instruction Tuning

Connecting a CLIP vision encoder to a large language model via a simple projection layer, trained in two stages with GPT-4-generated instruction data.

Visual TokenizationVision-Language ProjectionTwo-Stage Training+4
Start Learning
Vision-LanguagePRO

BLIP-2

Bootstrapping Language-Image Pre-training with Frozen Encoders and LLMs

Freeze the image encoder, freeze the LLM, and train only a 188M-parameter Q-Former between them — beating Flamingo80B by 8.7 points on zero-shot VQAv2 with 54x fewer trainable parameters.

The Modality GapQ-FormerStage 1 — Representation Learning+4
Pro Only
Self-SupervisedPRO

DINOv3

7B Parameters, Universal Visual Features

Self-supervised learning at scale with Gram Anchoring. 7B parameters trained on 1.7B images achieving SOTA on 60+ benchmarks.

Self-DistillationGram Anchoring7B Scale+3
Pro Only
Vision-LanguagePRO

SigLIP 2

Multilingual Vision-Language Encoders

Unified training recipe combining sigmoid loss, decoder objectives, self-distillation, and masked prediction for improved multimodal understanding.

Sigmoid LossDecoder ObjectivesSelf-Distillation+3
Pro Only
SegmentationPRO

Segment Anything (SAM)

Promptable Segmentation for Any Image

Foundation model for image segmentation that can segment any object using points, boxes, or masks as prompts.

Promptable SegmentationModel ArchitectureZero-Shot Transfer+3
Pro Only
3D ReconstructionPRO

3D Gaussian Splatting

Real-Time Radiance Field Rendering

Replaces neural radiance fields with explicit 3D Gaussians and tile-based rasterization, achieving real-time (>100 FPS) rendering with state-of-the-art visual quality.

3D Gaussian PrimitivesDifferentiable SplattingSpherical Harmonics+3
Pro Only
Self-SupervisedPRO

Masked Autoencoders

Self-Supervised Visual Pre-training

Pre-training vision transformers by masking 75% of image patches and learning to reconstruct them.

Asymmetric Encoder-DecoderHigh Masking RatioRandom Masking+3
Pro Only
Vision-LanguagePRO

CLIP

Contrastive Language-Image Pre-training

Learning visual representations from natural language supervision using contrastive pre-training on 400M image-text pairs.

Contrastive Pre-trainingDual EncoderZero-Shot Transfer+3
Pro Only
GenerativePRO

Diffusion Models

Denoising Diffusion Probabilistic Models

Learning to generate images by reversing a gradual noising process, achieving state-of-the-art image synthesis.

Forward ProcessReverse DenoisingNoise Schedule+4
Pro Only
ArchitecturePRO

Vision Transformer (ViT)

An Image is Worth 16x16 Words

Applying Transformers directly to image patches, achieving state-of-the-art results on image classification.

Patch EmbeddingsPosition EmbeddingsClass Token+3
Pro Only
Object DetectionPRO

DETR

End-to-End Object Detection with Transformers

Treating object detection as a set prediction problem using Transformers and bipartite matching loss.

Object QueriesBipartite MatchingEncoder-Decoder+4
Pro Only
OptimizationPRO

AdamW vs Adam

Decoupled Weight Decay Regularization

Why AdamW decouples weight decay from gradient updates, and when to use Adam vs AdamW in modern deep learning.

MomentumAdaptive Learning RatesWeight Decay vs L2+2
Pro Only
ArchitecturePRO

Attention Is All You Need

The Transformer Architecture

Introducing the Transformer architecture with self-attention mechanisms that revolutionized NLP and later computer vision.

Self-AttentionMulti-Head AttentionPositional Encoding+4
Pro Only
Instance SegmentationPRO

Mask R-CNN

Instance Segmentation Framework

Extending Faster R-CNN with a parallel mask branch for pixel-precise instance segmentation.

Two-Stage PipelineRoIAlignMask Branch+3
Pro Only
TrainingPRO

Normalization Techniques

BatchNorm, LayerNorm, GroupNorm & InstanceNorm

Visual guide to normalization types: which dimensions each operates over and when to use which.

Internal Covariate ShiftBatch NormalizationLayer Normalization+2
Pro Only
ArchitecturePRO

ResNet

Deep Residual Learning

Skip connections enabling training of very deep networks, solving the degradation problem in neural networks.

Skip ConnectionsGradient FlowBottleneck+4
Pro Only
SegmentationPRO

U-Net

Convolutional Networks for Biomedical Image Segmentation

Encoder-decoder architecture with skip connections for precise localization in image segmentation tasks.

Encoder-DecoderSkip ConnectionsTransposed Convolution+4
Pro Only

More papers coming soon: Mask R-CNN, GANs, Diffusion Models, and more.