PIXELBANKv9.1.0
Menu
Back to Concepts
Object Detection2024

YOLOv11

An Overview of the Key Architectural Enhancements — C3k2, SPPF, C2PSA and Multi-Task Heads

Rahima Khanam, Muhammad Hussain

Read the Paper on arXiv

Paper Overview

YOLOv11 (Khanam & Hussain, University of Huddersfield, October 2024) is a different kind of paper from most breakdowns on this site. YOLO11 itself was not introduced in a paper at all — Ultralytics (Jocher & Qiu) unveiled it at the YOLO Vision 2024 conference and shipped it as code. This arXiv report is the first written architectural analysis of that release: it reads the model apart, names its new modules, and places it in the nine-year YOLO lineage.

The analysis centres on three components. C3k2, a cheaper Cross Stage Partial block, replaces YOLOv8's C2f in the backbone, neck and head. SPPF (Spatial Pyramid Pooling — Fast) is retained at the end of the backbone. And C2PSA, a Cross Stage Partial block with spatial self-attention, is new — placed directly after SPPF, it is the attention mechanism YOLOv8 never had.

Around that sits the familiar backbone → neck → head layout, and an expanded task list: one architecture family serves object detection, instance segmentation, image classification, pose estimation, oriented bounding boxes (OBB) and tracking, each in nano, small, medium, large and extra-large sizes.

The headline numbers come from Ultralytics' published benchmark, which the paper reproduces as its Figure 2: YOLO11m reaches higher COCO mAP than YOLOv8m with 22% fewer parameters, and YOLO11x reaches roughly 54.5% mAP50−95^{50-95} at about 13 ms on a T4 GPU — the top of the latency/accuracy frontier among YOLO versions. Read it as a guided tour of a production detector rather than a new method: the value is in seeing why each block is shaped the way it is.

Chapter Roadmap

Click any topic to jump in

1
Nine Years of YOLO

From the 2015 single-stage detector to YOLOv10's NMS-free training — the lineage YOLO11 inherits from.

distilled into
2
Backbone, Neck, Head

YOLO11 keeps YOLOv8's three-part layout and swaps in three new modules: C3k2, SPPF + C2PSA.

upgraded with

Three new building blocks

3
C3k2 Block

A cheaper Cross Stage Partial block that replaces C2f everywhere — with a c3k switch for deeper inner modules.

4
SPPF

Three chained 5x5 max-pools reproduce 5/9/13 pyramid pooling at the end of the backbone.

5
C2PSA Attention

Self-attention over the coarsest feature map, wrapped in a CSP split — the module YOLOv8 lacks.

feeding
6
CBS Layers & Detect Head

Conv-BatchNorm-SiLU refinement and per-scale Conv2D branches that emit boxes and class scores.

deployed as

Capabilities and evidence

7
One Architecture, Five Tasks

Detection, segmentation, pose, oriented boxes and classification — n through x for each.

8
Results

~54.5% COCO mAP for 11x at ~13 ms; 11m beats YOLOv8m with 22% fewer parameters.

Before 2015, detection was a two-stage affair: propose candidate regions, then classify each one. Redmon et al.'s You Only Look Once collapsed that into a single pass — one convolutional network regresses bounding boxes and class probabilities for the whole image at once. Framing detection as regression is what made real-time detection practical, and every YOLO since has been an argument about how to keep that single pass while closing the accuracy gap to slower detectors.

The paper's Table 1 traces that argument version by version. YOLOv2 added multi-scale training and dimension-clustered anchors; YOLOv3 brought the Darknet-53 backbone, multi-scale prediction and an SPP block; YOLOv4 moved to CSPDarknet-53 with Mish activations. YOLOv5 marked the jump from Darknet to PyTorch, and with YOLOv8 (2023) Ultralytics settled on the anchor-free, multi-task design YOLO11 extends. 2024 alone produced three releases: YOLOv9 (Programmable Gradient Information and GELAN), YOLOv10 (consistent dual assignments for NMS-free training) and YOLO11.

The table is a compressed summary and some of its one-line attributions are loose — YOLOv5's official detector, for instance, is anchor-based, and anchor-free heads became the Ultralytics default with YOLOv8. What matters for YOLO11 is the immediate parent: it is an evolution of YOLOv8, not of YOLOv10. It keeps YOLOv8's backbone/neck/head topology and task heads, swaps C2f for C3k2, and borrows the idea of a lightweight attention block at the coarsest scale — an idea YOLOv10 had also used in its PSA module.

Nine Years of YOLO diagram

Key Points

1

YOLO (2015) reframed detection as a single regression problem: one CNN predicts boxes and class probabilities for the entire image in one pass

2

YOLOv2–v4 stayed on Darknet: multi-scale training and dimension clustering (v2), Darknet-53 + SPP (v3), CSPDarknet-53 + Mish (v4)

3

YOLOv5 (2020) moved the series to PyTorch; every version since in the table is PyTorch-based

4

YOLOv8 (2023) introduced the anchor-free, multi-task Ultralytics design — detection, segmentation and keypoints from one codebase

5

2024 releases: YOLOv9 (PGI + GELAN), YOLOv10 (consistent dual assignments → NMS-free), then YOLO11

6

YOLO11 builds directly on YOLOv8 — the paper repeatedly frames its changes as modifications of YOLOv8's blocks

7

Unlike most YOLOs, YOLO11 shipped as code from Ultralytics; this paper is an after-the-fact architectural analysis of that release