DETR
End-to-End Object Detection with Transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, Sergey Zagoruyko
Read the Paper on arXivPaper Overview
DETR (DEtection TRansformer) reformulates object detection as a direct set prediction problem, eliminating the need for hand-designed components like non-maximum suppression (NMS), anchor generation, or region proposal networks that have dominated detection architectures for a decade.
The architecture is elegantly simple: a ResNet-50 CNN backbone extracts a feature map of shape (e.g., for a standard COCO image), which is projected to 256 channels, flattened to a sequence of tokens, and processed by a standard transformer encoder-decoder. The encoder (6 layers, 8-head self-attention, ) performs global reasoning over all spatial positions. The decoder (6 layers) takes 100 learned object queries and transforms them into 100 output embeddings via cross-attention to the encoded image features. Two feed-forward network (FFN) heads predict a class label (92 COCO classes + "no object") and a normalized bounding box for each query.
The key insight is that detection can be viewed as predicting a set of objects -- and the Hungarian algorithm provides a differentiable-compatible way to compute the optimal one-to-one matching between the predictions and ground-truth objects (where typically ). Unmatched predictions are supervised to output the "no object" () class. This set-based loss is permutation-invariant -- the model never needs to learn an ordering over detections.
Why this matters architecturally:
- No anchors: Faster R-CNN uses 15 anchor templates per position across 5 FPN levels, generating ~200K candidates. DETR uses 100 learned queries -- period.
- No NMS: The set loss naturally teaches queries to avoid predicting the same object (the Hungarian matching assigns each GT to exactly one query), so there are no duplicate detections to suppress.
- No hand-tuned thresholds: IoU thresholds for anchor assignment (0.3/0.7), NMS thresholds (0.5), proposal counts (300) -- all eliminated.
- Global reasoning: Transformer self-attention enables each spatial position to reason about the entire image, capturing long-range dependencies that CNNs miss (e.g., context: a tennis racket near a person on a court).
COCO benchmark results (ResNet-50 backbone, 500 epochs):
- 42.0 AP overall (vs. Faster R-CNN's 42.0 AP with the same backbone -- competitive)
- 20.5 AP-S on small objects (vs. Faster R-CNN's 24.1 AP-S -- the main weakness)
- 61.1 AP-L on large objects (vs. Faster R-CNN's 54.0 AP-L -- significant advantage)
- Total parameters: 41M (ResNet-50: 23.5M backbone + 17.5M transformer + FFN heads)
- Inference: 28 FPS on a V100 GPU (vs. Faster R-CNN's 26 FPS -- comparable speed, but no NMS latency variance)
DETR's strength is on large objects (where global attention shines) but weakness on small objects (where the downsampled feature map loses detail). Deformable DETR, introduced later, fixes this with multi-scale features and converges 10x faster.
Chapter Roadmap
Click any topic to jump in
Encoder-Decoder
Transformer backbone replacing region proposals and anchor boxes.
Object Queries
Learned slot embeddings, one per potential object in the image.
Positional Encoding
2D sinusoidal encodings injecting spatial structure into attention.
Cross-Attention
Queries gather visual evidence from encoder features for localization.
Hungarian Matching
Optimal one-to-one assignment of predictions to ground truth.
Set Prediction
Permutation-invariant loss removing the need for NMS.
Deformable DETR
Sparse attention variant solving slow convergence and high resolution cost.
DETR replaces the entire proposal/anchor machinery of traditional detectors with a fixed set of N=100 learnable embeddings called object queries. Each query is a 256-dimensional vector that learns to specialize in detecting certain types of objects in certain spatial regions. Through 6 decoder layers of self-attention (among queries) and cross-attention (to image features), each query independently produces at most one detection or outputs 'no object'.
The Problem
The proposal/anchor paradigm and its accumulated complexity:
Every major object detector before DETR generates a massive number of candidate detections, then filters them down:
Faster R-CNN's proposal pipeline:
- Backbone produces features at 5 FPN levels (P2-P6)
- At each of ~200K spatial positions, 3 anchor ratios are evaluated
- RPN predicts objectness + box offsets for each of ~200K anchors
- NMS reduces to ~1000 proposals (IoU threshold 0.7)
- Stage 2 classifies and refines ~300 proposals
- Final NMS (IoU threshold 0.5) produces ~100 detections
- Hand-designed hyperparameters: anchor sizes (32, 64, 128, 256, 512), aspect ratios (0.5, 1.0, 2.0), NMS IoU thresholds (0.7 for RPN, 0.5 for final), positive/negative IoU thresholds (0.3/0.7), proposal counts (1000/300), score thresholds
Single-stage detectors (YOLO, SSD, RetinaNet):
- Dense anchors at every spatial position across multiple scales
- RetinaNet: 9 anchors per position x 5 levels = ~100K anchors for a 600x800 image
- Focal loss to handle extreme foreground/background imbalance (99.9% of anchors are background)
- NMS remains mandatory to remove duplicate detections
The fundamental problems:
- NMS is non-differentiable: It is a hard selection operation that cannot be optimized during training. The model learns to produce overlapping detections, then relies on a separate heuristic to remove them.
- Anchor design requires domain knowledge: The choice of anchor scales, ratios, and assignment thresholds significantly affects performance. RetinaNet's 9 anchor templates were tuned on COCO; different datasets may need different anchors.
- No global reasoning about duplicate suppression: Each anchor/proposal is processed independently. The model has no mechanism to "know" that another anchor is already detecting the same object. NMS is a post-hoc fix for this lack of global coordination.
- Two-stage overhead: RPN + NMS + RoIAlign + second-stage NMS adds both compute and engineering complexity. Each component has its own hyperparameters and failure modes.
- Arbitrary prediction ordering: Detections are ordered by confidence score, which has no semantic meaning. There is no principled connection between the set of predicted objects and the set of ground-truth objects.
The Solution
Object queries: N=100 learnable embeddings that directly produce the detection set:
DETR replaces the entire anchor/proposal/NMS pipeline with 100 learned vectors, each of dimension :
These are randomly initialized and learned via backpropagation, just like any other network parameter. They are the only input to the decoder (besides the encoded image features).
How object queries work through the decoder:
Each of the 6 decoder layers performs three operations:
1. Self-attention among queries (queries reason about each other):
This is the key to NMS-free detection: queries can "see" what other queries are predicting and learn to avoid duplicates. If query 7 is already detecting a dog, query 23 learns (through self-attention) to detect a different object or output .
2. Cross-attention to image features (queries look at the image):
where is the encoder output (850 tokens for a typical COCO image) and is the 2D positional encoding. Each query attends to all spatial positions and learns to focus on regions relevant to its predicted object.
3. FFN refinement (per-query feature refinement):
where the FFN has hidden dimension 2048 (8x expansion from ).
After 6 decoder layers, each query produces a 256-dim output embedding that is fed to two parallel prediction heads:
Why N=100?
COCO images contain at most ~63 annotated objects (the maximum in the dataset). Setting N=100 provides comfortable headroom. The authors tested:
- N=50: 40.8 AP (some images with many objects lose detections)
- N=100: 42.0 AP (optimal)
- N=200: 41.8 AP (more unmatched queries provide weaker supervision signal per matched query)
- N=300: 41.5 AP (training becomes harder with many targets)
Query specialization (emergent behavior):
Despite having no explicit spatial initialization, trained queries develop clear specialization:
- Certain queries consistently detect objects in the bottom-left of images
- Other queries specialize in large objects (cars, buses) while others focus on small objects (bottles, cups)
- Some queries learn to detect specific object categories preferentially
- This specialization emerges purely from the set-based loss and cross-attention mechanism -- no hand-designed assignment rules
No NMS needed because:
- Self-attention allows queries to coordinate and avoid duplicate detections
- The Hungarian matching loss assigns each GT object to exactly one query, so the model is never rewarded for predicting the same object twice
- Unmatched queries learn to output with high confidence, producing clean output without filtering
Key Points
N=100 learned 256-dim embeddings replace ~200K anchors, RPN, NMS, and all associated hyperparameters -- the entire proposal machinery reduced to a single parameter matrix
Self-attention among queries enables duplicate suppression without NMS: each query 'sees' what others are predicting and learns to output distinct objects or empty-set
Cross-attention to encoder output (HW x 256 tokens) allows each query to attend to any spatial position -- attention maps become sparse and object-focused after training
N=100 is optimal for COCO (max 63 objects per image): N=50 loses 1.2 AP from capacity limits, N=200 loses 0.2 AP from training difficulty with many empty-set targets
Query specialization emerges naturally: specific queries learn to detect objects in particular spatial regions, at particular scales, or of particular categories -- without any explicit assignment
Each query produces a 92-class prediction (91 COCO + empty-set) and a normalized box (x_c, y_c, w, h) via separate FFN heads after 6 decoder layers
Mathematical Formulation
Query Output (per decoder layer)
Each decoder layer transforms query q_i through three sub-layers: (1) self-attention with all other queries (for duplicate suppression), (2) cross-attention to the HW encoded image features (for localization), (3) FFN with 2048 hidden units (for refinement). After 6 layers, the final q_i is decoded into class logits and box coordinates.
Box Prediction Head
Bounding boxes are predicted as normalized center coordinates and dimensions relative to the image. The sigmoid activation constrains all values to [0,1]. The 3-layer MLP has hidden dimension 256 with ReLU activations. This direct regression avoids the anchor-relative parameterization of Faster R-CNN.
Object queries are learned embeddings that act as 'slots' the decoder fills with object hypotheses. Each query attends to the encoder output via , gathering features from spatial locations relevant to its slot. After training, each query specializes to a particular size/location prior — like learned anchor boxes but in feature space.
DETR uses the Hungarian algorithm to find the optimal one-to-one matching between the N=100 predictions and M ground-truth objects, then computes loss only on matched pairs. This set-based loss formulation is what makes DETR end-to-end trainable -- it replaces the heuristic IoU-threshold-based anchor assignment of traditional detectors with an optimal matching that considers class, location, and box size simultaneously.
The Problem
The assignment problem: connecting predictions to ground truth
With N=100 predictions and M ground-truth objects (typically M=5-15 for COCO images, always ):
- There are possible assignments of predictions to GT objects -- for N=100 and M=10, that is approximately possibilities
- Some predictions should match objects, the remaining N-M should predict "no object" ()
- The loss function must automatically determine which prediction is responsible for which GT object
- Crucially: the assignment must be one-to-one (each GT matches exactly one prediction) to avoid duplicate detection
How traditional detectors handle assignment (and why it is suboptimal):
Faster R-CNN (IoU-threshold assignment):
- An anchor is "positive" if IoU with any GT > 0.7, "negative" if IoU with all GT < 0.3, ignored otherwise
- Problem: Multiple anchors match the same GT object (intentionally), creating redundant detections that NMS must later remove. The model is explicitly trained to produce duplicates.
- Problem: The 0.3/0.7 thresholds are hand-tuned on COCO. A different dataset with different object size distributions may need different thresholds.
- Problem: Assignment considers only spatial overlap (IoU), not class -- an anchor is positive even if its predicted class is wrong.
RetinaNet (similar IoU-based assignment):
- 9 anchors per position, positive if IoU > 0.5
- Focal loss helps with class imbalance but assignment is still many-to-one
DETR's requirement: We need a one-to-one assignment that is:
- Optimal: minimizes total matching cost across all prediction-GT pairs
- Joint: considers class, location, and size simultaneously (not just IoU)
- Permutation-invariant: the loss should not depend on which output slot (query index) a prediction comes from
- Compatible with backpropagation: the matching itself is not differentiable, but the loss computed on matched pairs must be
The Solution
Bipartite matching via the Hungarian algorithm:
DETR formulates assignment as a bipartite matching problem solved in two steps:
Step 1: Compute the N x M cost matrix
For each prediction-GT pair where and :
where:
- = prediction 's softmax probability for the GT class (higher is better, so negated)
- = L1 distance between GT box and predicted box in normalized coordinates (sum of absolute differences of )
- = Generalized IoU, which handles non-overlapping boxes better than standard IoU (range [-1, 1])
- Weights: , , (tuned on COCO val set)
Why both L1 and GIoU? L1 loss has equal magnitude regardless of box size: a 10-pixel offset on a 1000-pixel box is the same L1 as a 10-pixel offset on a 30-pixel box. GIoU is scale-invariant -- it penalizes relative misalignment. Together, they handle objects of all sizes.
Step 2: Hungarian algorithm finds the minimum-cost one-to-one matching
The Hungarian algorithm solves:
where is the set of all permutations of . This finds the optimal assignment that minimizes the total matching cost.
- Time complexity: -- for N=100, this takes ~0.5ms per image, negligible compared to the forward pass (~35ms)
- The algorithm is not differentiable -- but it does not need to be. It is run as a preprocessing step to determine the assignment, then gradients flow through the loss computed on matched pairs.
Step 3: Compute the loss on matched pairs
Once the optimal matching is found:
where:
- For the M matched queries: is the GT class, and both classification and box losses apply
- For the N-M unmatched queries: , and only classification loss applies (training them to predict "no object")
- Loss weights: ,
Class imbalance handling: With N=100 and M~7 (COCO average), ~93% of predictions are trained on the class. The authors down-weight the class loss by a factor of 10 to prevent the model from trivially predicting for everything.
Auxiliary losses: The matching loss is applied at every decoder layer (not just the final one), using the same matching . Prediction FFN heads at intermediate layers share weights with the final layer. This intermediate supervision significantly accelerates training and improves final performance by +1.2 AP.
Why this is fundamentally better than IoU-threshold assignment:
| Property | IoU-Threshold (Faster R-CNN) | Hungarian (DETR) |
|---|---|---|
| Assignment | Many-to-one (multiple anchors per GT) | Strictly one-to-one |
| Optimality | Greedy (each anchor independently) | Globally optimal (min-cost) |
| Criteria | IoU only (spatial) | Class + L1 + GIoU (joint) |
| NMS needed? | Yes (to remove many-to-one duplicates) | No (one-to-one by construction) |
| Hyperparameters | IoU thresholds 0.3/0.7, NMS 0.5 | None for assignment |
Key Points
Hungarian algorithm finds the globally optimal one-to-one matching in O(N^3) = O(100^3) ~0.5ms -- negligible compared to the ~35ms forward pass
Cost matrix combines class probability, L1 box distance, and GIoU (weights 1:5:2) -- considering identity, location, and scale simultaneously, unlike IoU-only anchor assignment
One-to-one matching eliminates duplicate detections by construction: each GT object is assigned exactly one prediction, so the model is never rewarded for predicting the same object twice
Unmatched queries (N-M ~93 of 100) are supervised with the empty-set class, down-weighted by 10x to prevent the model from trivially predicting empty-set everywhere
Auxiliary losses at all 6 decoder layers (shared FFN heads, same matching) provide intermediate supervision, improving final AP by +1.2 and stabilizing training
The matching itself is non-differentiable, but this is fine: the matching is a fixed assignment for each forward pass, and gradients flow normally through the loss computed on the matched pairs
Mathematical Formulation
Hungarian Matching Cost
The cost of matching prediction i to GT object j combines: (1) negative class probability (lower cost for correct class), (2) L1 box distance in normalized coordinates (penalizes absolute displacement equally for all sizes), (3) GIoU loss (penalizes relative misalignment, handles non-overlapping boxes). The 1:5:2 weighting was tuned on COCO val.
Set Prediction Loss
After the Hungarian algorithm finds the optimal permutation sigma-hat, the loss is computed on all N predictions: matched queries receive class + box loss (classification cross-entropy + L1 + GIoU), unmatched queries receive only class loss targeting the empty-set label (down-weighted by 10x). Box loss uses the same L1 + GIoU combination as the matching cost.
The Hungarian algorithm solves in time, finding the unique permutation matching predictions to ground truth. The matching cost combines class probability and box L1+GIoU. Critically, this gives every prediction at most one GT — eliminating the need for NMS at inference.
DETR uses a standard transformer encoder-decoder with a CNN backbone: the encoder (6 layers, 256-dim, 8 heads) performs global self-attention over the flattened spatial features, enabling every position to reason about the entire image. The decoder (6 layers) transforms 100 learned object queries into detection outputs via cross-attention to the encoded features. This architecture has only ~17.5M parameters beyond the backbone.
The Problem
Why CNNs alone are insufficient for set-based detection:
Object detection requires three capabilities:
1. Global context understanding (what objects are present and how they relate):
- A person next to a surfboard on sand strongly suggests a beach scene -- context disambiguates
- Two overlapping bounding boxes where one contains a person and another a backpack -- the backpack likely belongs to the person
- CNNs process local patches: a 3x3 conv at stride 16 sees only a 3x16=48 pixel region. Even deep ResNets have limited effective receptive fields (~100-200 pixels, far less than theoretical)
- Result: CNNs cannot reason about relationships between distant objects (e.g., a leash connecting a person and a dog 300 pixels apart)
2. Precise spatial localization (where exactly objects are):
- CNNs excel at this: hierarchical features with decreasing stride capture boundaries and positions
- ResNet-50 at stride 32 produces a 25x34 feature map for a 800x1066 image -- each cell encodes a 32x32 region with 2048-dimensional features
3. Set-level reasoning (predicting the right number of objects without duplicates):
- Traditional detectors produce ~100K candidate detections and filter with NMS -- a brute-force approach
- There is no mechanism in CNN architectures for one detection to "know about" another detection
- NMS is a hard, non-differentiable operation that approximates set prediction through greedy elimination
What a transformer adds:
Self-attention computes pairwise interactions between ALL positions in :
For a 25x34 = 850 token sequence (typical COCO image at stride 32), each token directly attends to all 850 other positions. This global receptive field in a single layer is what enables:
- Long-range context (a person's face and their bag 500 pixels away are directly connected)
- Duplicate suppression (object queries can see what other queries are detecting)
- Global counting (the model can "count" objects by distributing them across queries)
However, transformers alone struggle with images:
- No built-in spatial inductive bias (translation equivariance)
- Require positional encodings to understand spatial layout
- Quadratic complexity makes processing raw pixels infeasible (800x1066 = 852K tokens -> ~726 billion attention entries)
- Need a CNN to reduce spatial resolution first
The Solution
DETR's hybrid CNN-Transformer architecture (complete data flow with tensor shapes):
For a standard COCO image of size 800x1066 (H x W):
1. CNN Backbone (ResNet-50): Extracts a spatial feature map
The backbone reduces spatial dimensions by 32x (stride 32) while expanding channels to 2048. This is the standard ImageNet-pretrained ResNet-50 (23.5M parameters), optionally fine-tuned during DETR training with a 10x smaller learning rate.
2. Channel projection (1x1 Conv): Reduces channel dimension for the transformer
This 1x1 convolution (2048 x 256 = 524K parameters) projects the high-dimensional CNN features to the transformer's model dimension .
3. Flatten + Positional Encoding: Convert 2D feature map to 1D sequence with spatial information
The 2D feature map is flattened to a sequence of 850 tokens (25 x 34). Fixed 2D sine/cosine positional encodings are added (not concatenated) to inject spatial information, using 128 dimensions for x-position and 128 for y-position.
4. Transformer Encoder (6 layers): Global self-attention over all spatial positions
Each of the 6 encoder layers contains:
- Multi-head self-attention (8 heads, per head): Every token attends to all 850 tokens
- FFN (Linear(256, 2048) -> ReLU -> Linear(2048, 256)): Per-token feature transformation
- LayerNorm + residual connections on both sub-layers
Encoder parameter count: 6 layers x (4 x 256 x 256 [attention Q/K/V/O projections] + 256 x 2048 + 2048 x 256 [FFN] + LayerNorm) = ~6.3M parameters
What the encoder does: Self-attention enables global reasoning. After encoding:
- Each token's representation incorporates information from all 850 spatial positions
- The model can represent relationships like "these two distant regions both contain parts of the same large object"
- Attention patterns in trained models show that the encoder separates objects from background and captures object boundaries
5. Transformer Decoder (6 layers): Transforms object queries into detections
Input: 100 learned object queries (randomly initialized, learned end-to-end).
Each of the 6 decoder layers contains:
- Self-attention among queries (100 queries attend to each other): Enables coordination to avoid duplicate predictions
- Cross-attention from queries to encoder output (100 queries attend to 850 encoded tokens): Each query "looks at" the image and focuses on relevant regions
- FFN (same architecture as encoder): Per-query refinement
- LayerNorm + residual connections on all three sub-layers
Decoder parameter count: 6 layers x (self-attention + cross-attention + FFN + LayerNorm) = ~6.3M parameters
6. FFN Prediction Heads: Two parallel heads per query
Class head (shared across all decoder layers for auxiliary loss): Predicts 91 COCO classes + 1 "no object" () class. Softmax applied for loss computation.
Box head (3-layer MLP, also shared): Normalized center coordinates and size relative to the full image. Sigmoid ensures all values are in [0,1].
Head parameter count: ~0.9M (class: 256x92 = 24K, box MLP: 256x256 + 256x256 + 256x4 = ~131K, shared across 6 layers)
Total DETR parameter breakdown:
- ResNet-50 backbone: 23.5M (56%)
- 1x1 projection: 0.5M (1%)
- Transformer encoder: 6.3M (15%)
- Transformer decoder: 6.3M (15%)
- Prediction heads: 0.9M (2%)
- Object queries + positional encodings: ~0.05M (<1%)
- Total: ~41M parameters
Why this architecture excels at large objects:
The encoder's self-attention allows every spatial position to incorporate global context. For a large bus spanning 400 pixels, the feature map positions covering its front and back (separated by ~12 feature cells) directly communicate in the encoder, producing a unified representation. Traditional CNN detectors only connect these positions after multiple layers of local convolution, losing fine-grained spatial coordination.
DETR achieves 61.1 AP-L vs. Faster R-CNN's 54.0 AP-L (+7.1) -- a direct benefit of global attention.
Key Points
ResNet-50 backbone (23.5M params) extracts [1, 2048, 25, 34] features at stride 32, projected to [1, 256, 25, 34] by 1x1 conv, flattened to 850 tokens for the transformer
Encoder: 6 layers of 8-head self-attention over 850 tokens (d=256, FFN hidden=2048) -- each token reasons about all other positions, enabling global context that CNNs lack
Decoder: 6 layers with self-attention (query-to-query for duplicate suppression) + cross-attention (query-to-encoder for localization) + FFN -- transforms 100 queries into detection embeddings
Class head (Linear 256->92) predicts 91 COCO classes + empty-set; Box head (3-layer MLP with sigmoid) predicts normalized (x_c, y_c, w, h) in [0,1]^4
Total: ~41M parameters (23.5M backbone + 12.6M transformer + 0.9M heads). Inference at 28 FPS on V100, comparable to Faster R-CNN (26 FPS) but with no NMS latency variance
Excels on large objects (61.1 AP-L vs. Faster R-CNN's 54.0 AP-L) due to global attention, but struggles on small objects (20.5 AP-S vs. 24.1 AP-S) due to the 32x stride losing fine spatial detail
Mathematical Formulation
Transformer Encoder Self-Attention
Each of the 850 spatial tokens attends to all others with 8 heads (d_k = 32 per head). Positional encodings (PE) are added to Q and K (not V) so the attention can reason about spatial relationships. The attention matrix is 850x850 = 722K entries per head, enabling global reasoning but contributing to DETR's high memory usage.
Box Parameterization (Direct Regression)
Unlike Faster R-CNN's anchor-relative parameterization (delta_x, delta_y, delta_w, delta_h), DETR directly regresses normalized absolute coordinates. The sigmoid ensures values are in [0,1] (relative to image size). This is simpler but requires the model to learn the absolute position-to-box mapping from scratch, contributing to slow convergence.
The encoder applies transformer layers to the flattened CNN feature map ( tokens), letting every spatial position attend to every other — global receptive field in one layer. The decoder takes object queries and applies self-attention (queries reason about each other) followed by cross-attention to encoder output (queries gather visual evidence).
Architecture
Complete DETR data flow for 800x1066 COCO image:
Input: [1, 3, 800, 1066]
Backbone (ResNet-50):
- conv1 + pool: [1, 64, 200, 267]
- res2: [1, 256, 200, 267] -- stride 4
- res3: [1, 512, 100, 134] -- stride 8
- res4: [1, 1024, 50, 67] -- stride 16
- res5: [1, 2048, 25, 34] -- stride 32
Projection: Conv1x1(2048, 256) -> [1, 256, 25, 34]
Flatten: [1, 256, 25, 34] -> [1, 850, 256]
Add positional encoding: [1, 850, 256] (128-dim sine for x, 128-dim cosine for y)
Encoder (6 layers, each: self-attn(850x850, 8 heads) + FFN(256->2048->256)):
- Input: [1, 850, 256] -> Output: [1, 850, 256]
Decoder (6 layers, each: self-attn(100x100) + cross-attn(100x850) + FFN):
- Queries: [1, 100, 256] (learned parameters)
- Keys/Values from encoder: [1, 850, 256]
- Output: [1, 100, 256]
Prediction heads (per query, shared across decoder layers):
- Class: Linear(256, 92) -> [1, 100, 92] (softmax -> probabilities)
- Box: MLP(256, 256, 256, 4) + sigmoid -> [1, 100, 4] (normalized coords)
Hungarian matching + Loss (computed on all 100 outputs vs M GT objects)
In the decoder, each object query attends to the full encoded feature map via cross-attention, learning to focus on the spatial regions relevant to its predicted object. Remarkably, trained DETR models produce highly interpretable attention maps where each query focuses tightly on the extremities (edges and endpoints) of the object it detects, enabling precise localization without any explicit spatial priors like anchors.
The Problem
The localization challenge: connecting abstract queries to spatial positions:
Object queries are 256-dimensional learned embeddings with no inherent spatial meaning. They are the same regardless of what image is being processed. Yet they must:
- Attend to specific image regions: Query 42 might need to detect a car in the top-left of one image and a person in the bottom-right of another
- Handle varying numbers of objects: If there are 7 objects, 7 queries should attend to object regions while 93 should attend "nowhere" (or broadly to predict )
- Avoid overlapping with other queries: If query 15 is already detecting a dog, query 42 should not also detect that dog -- but they must coordinate through self-attention and cross-attention
- Localize precisely: DETR achieves 42.0 AP, which requires bounding boxes accurate to within ~5-10 pixels on average
The puzzle: How does a fixed embedding vector (query) that is the same across all images learn to look at different positions in different images?
The answer: Cross-attention weights are computed dynamically from the interaction between the query and the encoded image features. The query's learned representation defines what "pattern" to look for (e.g., large rectangular objects), while the image features determine where such patterns appear. The inner product produces high values where the query's pattern matches the image content.
In traditional detectors, localization is handled explicitly:
- Anchors are placed at fixed spatial positions with fixed sizes
- RPN regresses offsets from these fixed anchors
- The spatial prior is hard-coded into the architecture
In DETR, localization emerges from learned attention -- no spatial priors are hand-designed.
The Solution
Cross-attention: the mechanism that connects queries to image features
In each of the 6 decoder layers, cross-attention computes:
where:
- = object queries from the previous sub-layer (after self-attention)
- = encoded image features (keys and values are the same)
- = learned query positional embeddings (added to Q)
- = 2D sine/cosine spatial encodings (added to K)
- (head dimension for 8-head attention)
The attention matrix has entry representing how much query attends to spatial position :
What trained attention maps reveal (key empirical finding):
The authors visualize decoder cross-attention maps (reshaped from 850-length vectors back to 25x34 spatial grids) and observe:
-
Object-focused attention: Queries that predict real objects attend tightly to those objects, not to the background. The attention is sparse -- most of the 850 positions receive near-zero weight.
-
Extremity attention: Queries tend to attend to the extremities (edges, endpoints) of objects rather than their centers. For a car, the query attends to the four corners and the front/rear edges. For a person, it attends to the head and feet. This makes sense: extremities define the bounding box.
-
Progressive refinement across layers: Early decoder layers produce broad, diffuse attention patterns (query "searching" for an object). Later layers produce sharp, focused patterns (query "locked on" to specific extremities). By layer 6, attention is highly concentrated on 5-10 spatial positions.
-
-query attention: Queries that predict "no object" tend to have diffuse, unfocused attention patterns -- they attend broadly to the background without concentrating on any region.
-
Distinct per-query patterns: Different queries attending to the same image produce non-overlapping attention maps, confirming that self-attention among queries enables them to "divide up" the image.
Why cross-attention works better than spatial priors for large objects:
An anchor at position (100, 200) with size 128x128 can only detect objects near that spatial location at roughly that scale. If a bus spans from (50, 100) to (700, 300), no single anchor captures it well -- the detection relies on aggressive box regression from the nearest anchors.
In contrast, a DETR query can simultaneously attend to pixel (50, 100) and (700, 300), capturing both ends of the bus in a single attention operation. This is why DETR excels on large objects: 61.1 AP-L vs Faster R-CNN's 54.0 AP-L (+7.1 AP).
Multi-head attention splits the workload:
With 8 attention heads, each head can focus on different aspects:
- Head 1 might attend to the top-left extent of the object
- Head 2 might attend to the bottom-right extent
- Head 3 might attend to the object's center (for class identity)
- Other heads might capture context (nearby objects, scene type)
The 8 heads are projected together: where .
Key Points
Cross-attention dynamically computes a 100x850 attention matrix each forward pass -- each query selectively attends to the spatial positions most relevant to its predicted object
Trained models show queries attend to object extremities (corners, edges, head/feet) rather than centers -- these define the bounding box and are most informative for localization
Attention refines progressively across 6 decoder layers: early layers show broad search patterns, late layers show sharp focus on 5-10 positions per query
8 attention heads specialize: some heads attend to top-left extents, others to bottom-right, others to centers for classification -- together covering all aspects of detection
Queries predicting empty-set develop diffuse, unfocused attention patterns -- a learned 'no object here' signal without any explicit supervision on attention maps
Large-object advantage explained: a single cross-attention operation can simultaneously attend to both ends of a 600-pixel bus, while anchors require aggressive regression from the nearest position
Mathematical Formulation
Decoder Cross-Attention
Each of 100 queries (Q + learned positional embeddings) computes dot-product similarity with all 850 encoded spatial tokens (F_enc + 2D sine/cosine encodings). The softmax produces a probability distribution over spatial positions per query. The weighted sum of F_enc values produces a 256-dim context vector per query, encoding the image content most relevant to that query's detection.
Multi-Head Attention (8 heads)
The 256-dim query, key, and value are projected to 8 heads of 32-dim each. Each head independently computes attention, enabling different heads to focus on different spatial aspects (e.g., one head for top-left extent, another for bottom-right). The 8 outputs are concatenated and projected back to 256-dim.
Cross-attention computes , with each query producing a soft spatial heatmap over the image. Visualizations show queries attending to object extremities (corners of the bounding box) — the model learns to localize by gathering corner evidence rather than predicting center+size.
Comparison
Before / Traditional
Traditional Detectors (Faster R-CNN, RetinaNet):
- Anchors at fixed spatial positions (15 per location across FPN levels)
- Spatial priors hand-designed: sizes (32-512px), ratios (0.5/1.0/2.0)
- NMS post-processing to remove overlapping detections
- No mechanism for one detection to "see" another
- Large objects split across multiple anchors with no coordination
- IoU-threshold assignment with hand-tuned thresholds (0.3/0.7)
After / This Paper
DETR Cross-Attention:
- 100 learned queries attend dynamically to any spatial position
- No spatial priors -- localization emerges from learned attention patterns
- No NMS -- self-attention among queries prevents duplicates
- Queries see each other via self-attention and coordinate
- Large objects captured by single attention operation spanning both extremities
- Hungarian matching provides optimal, joint assignment considering class + location + size
DETR treats detection as set prediction -- the order of the 100 output predictions does not matter, only that the predicted set matches the ground-truth set. This is a fundamental departure from traditional detectors where predictions are ordered by confidence score and the loss depends on this ordering through anchor assignment.
The Problem
Why detection should be set prediction:
Objects in an image have no natural ordering. Consider an image with a dog, a cat, and a bird:
- Is the "correct" order {dog, cat, bird}? Or {cat, bird, dog}? Or {bird, dog, cat}?
- There are equally valid orderings, yet a loss function that depends on order would penalize 5 of them
Traditional detectors impose artificial ordering:
Anchor-based (Faster R-CNN, RetinaNet):
- Predictions are ordered by their spatial anchor position (top-left to bottom-right, across FPN levels)
- This ordering is arbitrary: an object in the top-left is not inherently "first"
- The anchor assignment (which prediction is responsible for which GT) depends on spatial IoU, creating a fixed correspondence between positions and objects
Confidence-based (post-NMS output):
- Final detections are sorted by confidence score
- The highest-confidence detection is "first" -- but this has no semantic meaning
- Permuting the output would not change the detection quality, yet the evaluation implicitly depends on this ordering for AP computation
Autoregressive models (rare but theoretically possible):
- Predict objects one at a time, conditioned on previous predictions
- Imposes a sequential dependency that doesn't exist in the data
- Slower inference (can't parallelize) and error accumulation
What we need:
- A loss function where -- truly permutation-invariant
- A prediction mechanism that does not impose ordering constraints
- A training procedure that can handle varying numbers of objects per image (M varies from 0 to ~63 in COCO)
The Solution
Set-based prediction with Hungarian matching provides true permutation invariance:
DETR's formulation satisfies all three requirements:
1. Permutation-invariant loss via Hungarian matching:
The Hungarian algorithm finds the optimal assignment regardless of which output slot each object is predicted in:
If the model outputs {dog at slot 7, cat at slot 42, bird at slot 91}, the Hungarian algorithm will find this matching and compute the same loss as if the outputs were {dog at slot 1, cat at slot 2, bird at slot 3}. The loss value depends only on the quality of the predictions (class correctness and box accuracy), not on which slot they appear in.
Formal property: For any permutation of query indices:
2. Parallel (non-autoregressive) decoding:
All 100 object queries are processed simultaneously in the decoder:
- Self-attention: all 100 queries interact in parallel (not sequentially)
- Cross-attention: all 100 queries attend to the image in parallel
- FFN: applied independently to each query
This means inference produces all 100 predictions in a single forward pass with no sequential dependencies. Unlike autoregressive models (which predict token-by-token), DETR's predictions are conditionally independent given the encoder output and the self-attention interactions.
3. Variable-cardinality set prediction:
The model always outputs exactly N=100 predictions, but the number of "real" objects varies per image. The Hungarian matching handles this elegantly:
- If M=5 objects: 5 queries are matched to objects, 95 are supervised to predict
- If M=40 objects: 40 matched, 60 predict
- If M=0 objects (empty image): all 100 predict
The model effectively learns to "count" objects by varying how many queries output real predictions vs. .
Connection to set theory and combinatorics:
DETR's approach is deeply connected to the concept of sets in mathematics:
- A set is defined by its elements, not their order:
- Set cardinality is variable: can be 0, 1, ..., N
- Set comparison requires solving an assignment problem when elements are not exactly equal (soft matching via the cost matrix)
The Hungarian algorithm (also known as the Kuhn-Munkres algorithm) was originally developed in 1955 for the assignment problem in operations research. DETR is the first work to apply it to neural network training for object detection, bridging combinatorial optimization and deep learning.
Comparison with other set prediction methods:
| Method | Ordering | Assignment | Parallel | Used in |
|---|---|---|---|---|
| Sorted by confidence | Arbitrary | IoU threshold | Yes | Faster R-CNN |
| Autoregressive | Sequential | Greedy | No | Image captioning |
| Hungarian matching | None | Optimal | Yes | DETR |
DETR's set prediction is the only approach that is simultaneously unordered, optimally assigned, and parallelizable.
Key Points
Detection as unordered set prediction: L({dog, cat, bird}) = L({bird, cat, dog}) -- the Hungarian matching produces identical assignments and identical loss regardless of which output slot each prediction occupies
All 100 queries are decoded in parallel (non-autoregressive): self-attention and cross-attention process all queries simultaneously, producing all detections in a single forward pass
Variable-cardinality handling: N=100 fixed outputs, but M=0..63 real objects per image. Unmatched queries (N-M) are supervised to predict empty-set, effectively learning to 'count' objects
The Hungarian algorithm (Kuhn-Munkres, 1955) solves the assignment problem optimally in O(N^3) -- the first application of this classical combinatorial optimization technique to neural network training for detection
Set prediction eliminates the need for confidence-based sorting, NMS, and score thresholds -- the output set is ready to use without post-processing
Formally proved permutation invariance: for any permutation pi of the N query indices, the loss is identical, because the Hungarian algorithm finds the optimal matching independent of output ordering
Mathematical Formulation
Set Equality (Permutation Invariance)
For any permutation pi of the N=100 output predictions, the Hungarian-matched loss is identical. This is because the algorithm finds the globally optimal assignment regardless of which slot each prediction occupies -- the 'identity' of a prediction is its content (class + box), not its index.
Optimal Assignment (Hungarian Algorithm)
The Hungarian algorithm finds the permutation sigma that minimizes the total matching cost between M ground-truth objects and their assigned predictions. This runs in O(N^3) time (~0.5ms for N=100), is computed once per forward pass as a preprocessing step, and produces the assignment used for loss computation. The algorithm itself is non-differentiable, but gradients flow through the loss on matched pairs.
The loss is invariant to the order of predictions: for any permutation . This is achieved by the Hungarian matching step before computing the per-pair loss. Without it, the model would have to memorize an arbitrary ordering — set prediction lets gradient flow naturally to whichever query best matches each GT.
DETR uses two distinct types of positional encoding: fixed 2D sine/cosine encodings for image features (injecting spatial location into each of the 850 tokens) and learned embeddings for the 100 object queries (giving each query a unique identity). Without these encodings, the transformer cannot distinguish spatial positions or individual queries, and performance drops catastrophically.
The Problem
The position-blindness problem in transformers:
Transformers are permutation-equivariant by design: if you shuffle the input tokens, the self-attention output shuffles in exactly the same way. This is mathematically elegant but catastrophic for images:
Problem 1: Image features lose spatial meaning
- After flattening the 25x34 feature map to 850 tokens, token 0 and token 849 are treated identically by the transformer
- But token 0 represents the top-left corner (position 0,0) and token 849 represents the bottom-right (position 24,33) -- they are 800 pixels apart!
- Without position info, the encoder's self-attention matrix would be identical regardless of whether we process the image normally or with all pixels randomly shuffled
- Detection requires knowing WHERE objects are, not just WHAT objects exist
Problem 2: Object queries lose identity
- The 100 object queries are processed symmetrically by the decoder
- Without unique identifiers, all queries would produce identical outputs (same class probabilities, same box coordinates) because they have no way to differentiate themselves
- We need each query to develop a unique "role" -- but this requires some asymmetry in the input
Problem 3: 2D structure of images
Natural language is 1D (token sequences), so standard 1D positional encodings work. Images are 2D with spatial structure along both axes:
- Horizontal neighbors share y-coordinates but differ in x
- The 2D spatial relationship between two positions matters (e.g., "directly above" vs. "diagonally adjacent")
- 1D positional encodings (treating the flattened sequence as text) would lose this 2D structure
Ablation showing the impact (from the paper):
- DETR without any positional encoding: ~32 AP (vs. 42.0 AP with encoding -- a 10 AP drop)
- With 1D (text-style) encoding: ~38 AP (better, but loses 2D structure)
- With learned 2D encoding: ~41.5 AP (good, but slightly overfits)
- With fixed 2D sine/cosine encoding: 42.0 AP (best, and generalizes to different image sizes)
The Solution
Two types of positional encodings tailored for detection:
Type 1: Fixed 2D Sine/Cosine Encodings (for image features)
For each spatial position in the 25x34 feature grid, a 256-dimensional encoding is constructed:
where (128 dimensions per axis, concatenated to 256 total).
Key properties of sine/cosine encodings:
- Fixed (not learned): No additional parameters, works for any image size
- Unique per position: Every (x, y) pair produces a unique 256-dim vector
- Encodes relative position: can be expressed as a linear function of , meaning the model can learn to compute relative distances through linear projections in attention
- Multi-frequency: Low-frequency components ( near 0, wavelength ) encode coarse position; high-frequency components ( near 63, wavelength ) encode fine position
- Separable: The x and y encodings are independent, which is appropriate because horizontal and vertical positions have different semantics in images (vertical: objects at different scales, horizontal: left/right context)
Where they are applied:
The 2D positional encodings are added to Q and K (not V) in both the encoder self-attention and the decoder cross-attention:
This means position influences which tokens attend to which (through Q and K) but does not directly modify the feature content (V remains position-free). This design choice allows the attention mechanism to learn position-dependent routing while keeping the value representations purely content-based.
Type 2: Learned Embeddings (for object queries)
The 100 object queries use learned positional embeddings:
These are randomly initialized and updated via backpropagation during training. They serve a dual purpose:
- Query identity: Each query gets a unique that differentiates it from all other queries, breaking the symmetry that would otherwise produce identical outputs
- Spatial prior: After training, these embeddings develop spatial structure -- certain queries' embeddings encode a preference for particular image regions (e.g., query 42 might specialize in the center of images)
Why learned (not fixed) for queries?
Object queries have no natural spatial meaning (unlike image features, which correspond to specific pixel regions). Their "position" is abstract -- it represents a detection slot, not a spatial coordinate. Learned embeddings allow each query to develop its own specialization through training, which would not be possible with fixed encodings.
In the decoder, query positional embeddings are applied to:
- Self-attention: added to Q and K so queries can distinguish each other
- Cross-attention: added to Q so each query has a unique "search pattern" when attending to image features
Implementation detail: The same for image features is used in both encoder self-attention (token-to-token) and decoder cross-attention (query-to-token). The same for object queries is used in both decoder self-attention (query-to-query) and decoder cross-attention (query side).
Key Points
2D sine/cosine encodings for image features: 128 dims for x-position + 128 dims for y-position concatenated to 256-dim vectors. Fixed, no learnable parameters, generalizes to any image size
Learned embeddings for 100 object queries: 100 x 256 parameters trained end-to-end. Break query symmetry and develop spatial specialization -- some queries learn to focus on specific image regions
Positional encodings are added to Q and K (not V) in attention: position determines which tokens attend to each other, but value representations stay purely content-based
Without positional encoding: ~32 AP (10-point drop from 42.0). With 1D encoding: ~38 AP. With learned 2D: ~41.5 AP. Fixed 2D sine/cosine: 42.0 AP (best generalization)
Multi-frequency design: low-frequency components encode coarse position (which quadrant of the image), high-frequency components encode fine position (precise pixel location) -- enabling both global and local spatial reasoning
Relative position encoding: sin/cos encodings allow the model to compute relative distances via linear projections, enabling attention to capture 'this token is 5 positions to the right' without explicit relative attention mechanisms
Mathematical Formulation
2D Sine/Cosine Positional Encoding
For temperature T=10000 and dimension d=128 per axis, each (x,y) position produces a unique 256-dim encoding. Low-frequency components (small i) capture coarse position, high-frequency (large i) capture fine position. The encoding is fixed (no learned parameters) and extends naturally to any spatial resolution.
Position-Aware Attention
Positional encodings are added to queries and keys before computing attention scores. This means the attention pattern depends on both content similarity (Q^T K) and positional compatibility (PE_Q^T PE_K), plus cross terms. The model learns to attend based on both 'what is similar' and 'what is nearby' jointly.
Transformers are permutation-invariant by default, so 2D sine encodings are added to keys and values to inject spatial structure. Without them the encoder can't distinguish 'top-left' from 'bottom-right'. Object queries also receive their own learned positional encodings.
Deformable DETR (Zhu et al., 2021) addresses DETR's two critical practical limitations -- slow convergence (500 epochs vs. Faster R-CNN's 36 epochs) and poor small-object detection -- by replacing the dense global attention with sparse deformable attention that attends to only K=4 learned sampling points per query per attention head. This reduces encoder self-attention complexity from to and enables multi-scale feature processing from 4 FPN levels, achieving state-of-the-art results with 10x faster training.
The Problem
DETR's practical limitations in detail:
Despite its elegant formulation, DETR has three significant problems that prevent practical adoption:
Problem 1: Extremely slow convergence (500 epochs, 12 days on 8 V100s)
DETR requires 500 training epochs on COCO (118K images) to reach 42.0 AP. For comparison:
- Faster R-CNN: 36 epochs (12x fewer) for 42.0 AP
- RetinaNet: 12 epochs (40x fewer) for 39.1 AP
Why so slow? The cross-attention in the decoder must learn to focus on relevant image positions from scratch. Without spatial priors (anchors provide these for free in Faster R-CNN), the model starts with near-uniform attention over all 850 spatial positions and must gradually sharpen to object-specific patterns. This "attention learning" is the bottleneck -- box regression and classification converge much faster.
Empirical evidence: At epoch 50, DETR achieves only ~35 AP (vs. 42.0 at epoch 500). The attention maps are still diffuse and unfocused at this stage.
Problem 2: Quadratic memory/compute in encoder self-attention
The encoder self-attention computes an attention matrix:
- For 800x1066 image at stride 32: tokens -> attention entries per head
- For higher-resolution features (stride 16): tokens -> attention entries per head
- For multi-scale features (stride 8): tokens -> entries -- infeasible
This quadratic scaling prevents DETR from using high-resolution features, which are essential for small-object detection.
Problem 3: Poor small-object detection (20.5 AP-S vs. Faster R-CNN's 24.1 AP-S)
DETR uses only stride-32 features (the C5 output of ResNet). At this resolution:
- A 20x20 pixel object (COCO "small" category) occupies only 0.6 x 0.6 feature cells -- it is essentially invisible
- A 32x32 object occupies 1x1 cell -- barely detectable
- Faster R-CNN uses FPN with stride-4 features (P2), where a 20x20 object occupies 5x5 cells -- ample resolution
Adding multi-scale features to DETR would require processing 850 + 3,350 + 13,400 = 17,600 tokens through self-attention: attention entries per head. This is prohibitively expensive.
Problem 4: High GPU memory usage
Storing the attention matrix (float32) for 8 heads across 6 encoder layers requires: for attention alone. With multi-scale features, this would balloon to -- exceeding any single GPU's memory.
The Solution
Deformable DETR: sparse attention + multi-scale features
Core innovation: Deformable attention module
Instead of computing attention over ALL spatial positions (dense, ), deformable attention attends to only learnable sampling points per query per attention head:
For each query and reference point (the center of its predicted box from the previous layer):
where:
- attention heads
- sampling points per head
- : learnable 2D offset for head , point (predicted by a linear layer from )
- : attention weight for head , point (predicted by a linear layer + softmax over points)
- : feature value at the offset position, computed via bilinear interpolation (differentiable)
Key properties:
- Sparse: Each query attends to only points total, instead of 850+ positions
- Learned sampling: The offsets are predicted from the query content, so they adapt to each image
- Differentiable: Bilinear interpolation at fractional positions allows gradients to flow through offsets
- Complexity: instead of -- linear in spatial resolution instead of quadratic
Multi-scale deformable attention (the critical addition):
With O(HWK) instead of O(H^2W^2) complexity, processing multi-scale features becomes feasible:
where are features from FPN levels (P3 through P6 from ResNet-50-FPN):
- P3: stride 8, 100x134 = 13,400 tokens
- P4: stride 16, 50x67 = 3,350 tokens
- P5: stride 32, 25x34 = 850 tokens
- P6: stride 64, 13x17 = 221 tokens
- Total: 17,821 tokens
With dense attention: entries per head -- infeasible. With deformable attention: entries per head -- 3 orders of magnitude cheaper.
Iterative bounding box refinement:
Each decoder layer refines the predicted bounding box from the previous layer:
The reference point for deformable attention is updated to the center of , so sampling points are concentrated around the current best estimate of the object location. This progressive refinement is much faster than DETR's approach of learning localization from scratch at every layer.
Two-stage variant:
Deformable DETR can optionally operate in a two-stage mode:
- Stage 1: The encoder produces "region proposals" by predicting a box at every spatial position (similar to RPN but without anchors)
- Stage 2: Top-scoring proposals become reference points for the decoder, replacing learned object queries
This further improves performance to 46.2 AP (vs. 43.8 for one-stage Deformable DETR).
Convergence comparison (COCO val):
| Epochs | DETR AP | Deformable DETR AP |
|---|---|---|
| 12 | 32.5 | 41.5 |
| 50 | 35.4 | 43.8 |
| 150 | 39.2 | 44.1 |
| 500 | 42.0 | -- |
Deformable DETR at 50 epochs (43.8 AP) already surpasses DETR at 500 epochs (42.0 AP) -- 10x faster convergence AND +1.8 AP better.
Small-object detection improvement:
| Model | AP-S | AP-M | AP-L |
|---|---|---|---|
| DETR | 20.5 | 45.3 | 61.1 |
| Deformable DETR | 26.4 | 47.1 | 58.0 |
The +5.9 AP-S improvement comes directly from multi-scale features: stride-8 features give small objects 16x more feature cells than stride-32. The slight AP-L decrease (-3.1) is because deformable attention's sparse K=4 points are less effective than DETR's global attention for reasoning about very large objects.
Total parameter count: ~40M (comparable to DETR), with the deformable attention modules replacing standard attention with minimal parameter overhead.
Key Points
K=4 learned sampling points per head per query, with 8 heads = 32 total points per attention operation. Offsets predicted from query content via linear layer, features sampled via differentiable bilinear interpolation
Complexity: O(HW x MK) = O(17,821 x 32) ~570K ops vs. O(H^2W^2) = O(17,821^2) ~317M ops for multi-scale features -- 3 orders of magnitude reduction enabling multi-scale processing
Multi-scale features from 4 FPN levels (stride 8/16/32/64 = 17,821 total tokens) are now feasible. Small objects get stride-8 features (16x more detail than DETR's stride-32): AP-S improves from 20.5 to 26.4 (+5.9)
10x faster convergence: 50 epochs for 43.8 AP vs. DETR's 500 epochs for 42.0 AP. The reference point mechanism provides a spatial prior that accelerates attention learning
Iterative box refinement: each decoder layer predicts residual box offsets from the previous layer's prediction, concentrating sampling points around the current best estimate
Two-stage variant (encoder proposals -> decoder refinement) reaches 46.2 AP, establishing a new state-of-the-art for transformer-based detection in 2021
Mathematical Formulation
Deformable Attention
For each query q with reference point p_q, M=8 attention heads each attend to K=4 learnable offset points. The offsets delta_p are predicted from q via a linear layer. The attention weights A are predicted and normalized via softmax over the K points. Feature values at fractional positions are obtained via bilinear interpolation. This replaces the O(N^2) dense attention with O(N x MK) sparse attention.
Multi-Scale Deformable Attention
Extends deformable attention across L=4 FPN levels. The reference point p-hat is projected to each level's coordinate system via phi_l (accounting for different strides). The attention weights A are normalized over all L x K = 16 sampling points jointly, allowing the model to attend to the most informative scale for each query. This is the key to processing 17,821 multi-scale tokens efficiently.
Standard attention has complexity, dominating cost at high resolution. Deformable attention restricts each query to attend to only sampled points (typically ), with offsets predicted by the query: . Cost drops to — faster convergence.
Comparison
Before / Traditional
DETR:
- 500 epochs to converge (12 days on 8 V100s)
- Dense O(H^2W^2) attention: 722K entries for single-scale, infeasible for multi-scale
- Single-scale features (stride 32 only): 20.5 AP-S on small objects
- 42.0 AP overall, 41M parameters
- Global attention excels on large objects (61.1 AP-L) but struggles with spatial precision for small ones
After / This Paper
Deformable DETR:
- 50 epochs to converge (1.2 days on 8 V100s) -- 10x faster
- Sparse O(HW x MK) attention: 570K entries for 4 multi-scale levels -- 3 orders of magnitude cheaper
- Multi-scale features (stride 8/16/32/64): 26.4 AP-S (+5.9) on small objects
- 43.8 AP overall (+1.8), ~40M parameters (comparable)
- Two-stage variant reaches 46.2 AP, state-of-the-art for transformer detectors