Back to Blog
Vision & Multimodal

From DETR to YOLOv8: Transformer Evolution in Object Detection

📅 2026.03.15 ⏱️ 12 min 👤 Eric Pan

Introduction

Object detection is one of the most fundamental tasks in computer vision. From the R-CNN series to YOLO, two-stage and single-stage detectors have continuously evolved over the past decade. However, the emergence of DETR in 2020 completely changed the game — it was the first to introduce Transformer architecture into object detection, replacing traditional anchor-based methods with an end-to-end approach.

This article starts from DETR's core ideas, progressively introduces its evolution, and shares the challenges and solutions I encountered in engineering practice.

DETR's core contribution is transforming object detection from "heuristic design" to "end-to-end learning" — a paradigm-level revolution.

DETR Architecture Analysis

DETR (Detection Transformer) models object detection as a set prediction problem. It uses the Transformer encoder-decoder structure with the Hungarian matching algorithm to directly predict a fixed number of bounding boxes, completely eliminating anchors and NMS post-processing.

The overall architecture consists of three core components:

However, original DETR has two major issues: slow training convergence (requiring 500 epochs) and poor small object detection. Deformable DETR introduced deformable attention, focusing attention on key points, improving convergence speed by 10x while significantly boosting small object detection performance.

Attention Mechanism Optimization

The core computational bottleneck of Transformers lies in the attention mechanism. Standard self-attention has O(n²) complexity, which is prohibitively expensive for high-resolution feature maps. In engineering practice, we adopted the following optimization strategies:

Key insight: deformable attention reduces global attention complexity from O(n²) to O(n), making real-time inference possible. In actual deployment, we observed only 0.3 mAP drop in detection accuracy, but a 4x improvement in inference speed.

Model Optimization: TensorRT & ONNX

Porting DETR models from PyTorch to production requires multiple optimization stages. Here is our complete optimization pipeline:

The optimized model achieves 12ms end-to-end inference latency on NVIDIA T4 GPU (640x640 input), a 3.75x speedup over the original PyTorch model.

PyTorch Inference Code

Below is the complete inference pipeline code, including model loading, preprocessing, inference, and post-processing:

inference.py
import torch
import torchvision.transforms as T
from models import DeformableDETR
# ---- Initialize the model ----
model = DeformableDETR(num_classes=80, num_queries=300)
model.load_state_dict(torch.load('checkpoints/detr_v2.pth'))
model.eval().cuda()
# ---- Preprocessing ----
transform = T.Compose([
T.Resize((640, 640)),
T.ToTensor(),
T.Normalize([0.485, 0.456, 0.406],
[0.229, 0.224, 0.225])
])
# ---- Inference ----
with torch.no_grad(), torch.cuda.amp.autocast():
image = transform(raw_image).unsqueeze(0).cuda()
outputs = model(image)
probs = outputs['pred_logits'].softmax(-1)
boxes = outputs['pred_boxes']
# Filter low-confidence predictions
keep = probs.max(-1).values > 0.7
results = boxes[keep], probs[keep]

Benchmark Results

We conducted comprehensive benchmarks on the COCO 2017 validation set, comparing performance across different models and deployment approaches:

While DETR series still slightly trails YOLO in pure latency metrics, it clearly excels in tasks requiring precise spatial reasoning and complex scene understanding (e.g., occlusion handling, dense objects).

Selection advice: For precision-first, object-dense industrial scenarios, Deformable DETR + TensorRT is recommended; for latency-sensitive real-time video streams, YOLO remains the more pragmatic choice.

Engineering Deployment

Deploying DETR series models to production faces several key challenges:

Ultimately, through TensorRT INT8 quantization + CUDA Graph optimization, we reduced inference latency from 45ms to 12ms, meeting real-time video stream processing requirements. The system processes over 5 million frames daily with GPU utilization stable above 75%.

Summary & Outlook

Transformer applications in object detection are just beginning. From DETR to Deformable DETR, and the latest DINO and Co-DETR, this field is rapidly evolving. Future directions worth watching:

For engineering teams, the key is finding the balance between model accuracy and inference efficiency. Through careful architecture design and engineering optimization, Transformer-based detectors can absolutely achieve real-time performance in production environments.