Introduction
Industrial quality inspection is one of the most important application scenarios for computer vision. Traditional defect detection typically requires training specialized models for each product, with high data annotation costs and long model iteration cycles. Meta's SAM 2 (Segment Anything Model 2) brings new possibilities to this field — through powerful zero-shot segmentation and flexible Prompt mechanisms, we can achieve high-precision defect segmentation with less annotation data.
During my internship in vision model development, I had the privilege of deeply participating in building a SAM 2-based industrial quality inspection system. This article shares the complete experience of deploying SAM 2 in an electronic component surface defect detection project, including Prompt Engineering strategies, model fine-tuning methods, and edge device deployment optimization.
SAM 2's strength is not that it replaces every specialized model, but that it provides a powerful general-purpose foundation that substantially reduces the cost of adapting to new scenarios.
SAM 2 Architecture
SAM 2 introduces temporal modeling capabilities on top of the original SAM, enabling target segmentation in video sequences. Its core architecture consists of three modules:
- Image Encoder — Based on Hiera (Hierarchical Vision Transformer), extracting multi-scale visual features. Compared to SAM's ViT-H, Hiera achieves 6x faster inference while maintaining accuracy
- Prompt Encoder — Supports four Prompt types: point, box, text, and mask, encoding user intent into condition vectors
- Mask Decoder — Lightweight Transformer decoder that fuses image features and Prompt conditions to output high-quality segmentation masks
For industrial quality inspection, we primarily use static image mode, leveraging SAM 2's Prompt Encoder and Mask Decoder zero-shot generalization capabilities.
Prompt Engineering Strategy
In industrial quality inspection, Prompt design directly determines segmentation quality. Through extensive experimentation, we summarized the following Prompt Engineering strategies:
- Point Prompt Strategy — For known defect locations, place 1 foreground point at the defect center and 3-5 background points evenly around the defect boundary, effectively controlling segmentation boundaries
- Box Prompt Strategy — Use a detection model (e.g., YOLOv8) to provide rough defect boxes first, then use SAM 2 for precise segmentation. Box Prompt IoU threshold of 0.5 yields optimal results
- Cascaded Prompt — First use point Prompt for initial segmentation, then feed the initial mask as Prompt input, iteratively optimizing segmentation accuracy
- Text Prompt — Combined with CLIP's text encoder, using defect descriptions like "surface scratch" and "discoloration" as Prompts for zero-shot defect classification
Key finding: in industrial scenarios, combining a box prompt with a center-point prompt worked best, reaching an mIoU of 0.92, eight percentage points above point-only prompting.
Defect Detection Workflow
We designed a two-stage defect detection workflow combining traditional detection model speed with SAM 2 segmentation precision:
- Stage 1: Rapid Screening — Lightweight YOLOv8-nano performs initial detection, quickly locating suspected defect regions and filtering 90% of normal areas
- Stage 2: Precise Segmentation — YOLO detection boxes serve as Prompts for SAM 2, generating precise defect contour masks
- Defect Quantification — Calculates defect area, perimeter, shape factors, and other quantitative metrics from segmentation masks, automatically determining defect severity
- Result Reporting — Generates structured reports with defect location, type, and severity, integrated into MES systems
This two-stage approach maintains detection accuracy while keeping single-image processing time under 85ms (YOLO 15ms + SAM 2 70ms), meeting production line cycle requirements.
Python Inference Code
Below is the complete defect detection inference code, including model loading, Prompt construction, and post-processing:
Domain Fine-Tuning
While SAM 2's zero-shot capability is already impressive, lightweight fine-tuning in specific industrial scenarios can further improve performance. Our fine-tuning strategy:
- Freeze Image Encoder — Only fine-tune Prompt Encoder and Mask Decoder, reducing trainable parameters from 200M to 15M
- Dataset Construction — Only 500-1000 annotated samples needed, compared to 10000+ for training specialized models from scratch, reducing annotation costs by 90%
- Loss Function — Uses Focal Loss + Dice Loss combination to address extreme class imbalance where defect pixels are a tiny fraction
- Data Augmentation — Industrial-scenario-specific augmentation: random rotation, brightness jitter, Gaussian noise, simulating production line lighting variations
The fine-tuned model improved mIoU from 0.78 (zero-shot) to 0.91 on our electronic component defect dataset, with defect recall rising from 82% to 95%. Training required only 2 epochs, taking approximately 30 minutes on a single A100.
Edge Deployment Optimization
Production environments typically lack high-performance servers, requiring model deployment on edge devices (like NVIDIA Jetson Orin). Our optimization approach:
- Knowledge Distillation — Distill SAM 2-Large knowledge into SAM 2-Small, compressing model size from 900MB to 150MB
- TensorRT Quantization — INT8 quantization achieves 2.5x inference speedup with less than 1% accuracy loss
- Image Encoder Pre-computation — For fixed-position cameras, Image Encoder runs only once; subsequent frames only run the lightweight Mask Decoder
- Pipeline Parallelism — YOLO detection and SAM 2 segmentation run on different CUDA Streams in parallel
After optimization, we achieved 45ms end-to-end inference latency on Jetson Orin NX (YOLO 8ms + SAM 2 Decoder 37ms), meeting production line 20 FPS requirements. GPU utilization stable at 65%, power consumption ~25W.
Core principles for edge deployment: precompute where possible, use distilled models where practical, and quantize instead of relying on floating-point computation. A few milliseconds saved at each stage add up across a production line.
Production Line Results
After deployment on an electronic component factory's SMT production line, the system achieved significant results:
- Defect Detection Rate — Improved from 85% (manual) to 97.5% (AI system)
- False Positive Rate — Controlled within 1.2%, far below the 5% manual detection rate
- Detection Speed — Single station processes 500K+ components daily, 20x manual detection efficiency
- ROI — Deployment costs recovered within 3 months, saving over 2M yuan annually in labor costs
More importantly, the AI system's detection consistency far surpasses human inspectors — it doesn't fatigue, lose focus, or let mood affect judgment. This provides a reliable digital foundation for factory quality control.
Summary & Outlook
SAM 2 brings a paradigm shift to industrial quality inspection: from "training specialized models for each defect" to "using a general foundation model + minimal Prompts to adapt to new scenarios." This dramatically lowers the barrier for AI quality inspection, enabling small and medium factories to quickly deploy AI detection systems.
Future directions we plan to explore:
- Multi-Modal Fusion — Combining 2D images and 3D point cloud data to improve detection of 3D defects (dents, bumps)
- Online Learning — Continuously optimizing models using real production line data for automatic model evolution
- Multi-Station Collaboration — Correlating results from multiple inspection stations for system-level quality traceability
AI quality inspection isn't a technology project but a continuously operated system. Technology is just the starting point; what truly creates value is the deep integration of technology with production line processes.