Back to Blog
Vision & Multimodal

SAM 2 in Industrial Inspection: From Prompt Engineering to Production

📅 2025.07.18 ⏱️ 14 min 👤 Eric Pan

Introduction

Industrial quality inspection is one of the most important application scenarios for computer vision. Traditional defect detection typically requires training specialized models for each product, with high data annotation costs and long model iteration cycles. Meta's SAM 2 (Segment Anything Model 2) brings new possibilities to this field — through powerful zero-shot segmentation and flexible Prompt mechanisms, we can achieve high-precision defect segmentation with less annotation data.

During my internship in vision model development, I had the privilege of deeply participating in building a SAM 2-based industrial quality inspection system. This article shares the complete experience of deploying SAM 2 in an electronic component surface defect detection project, including Prompt Engineering strategies, model fine-tuning methods, and edge device deployment optimization.

SAM 2's strength is not that it replaces every specialized model, but that it provides a powerful general-purpose foundation that substantially reduces the cost of adapting to new scenarios.

SAM 2 Architecture

SAM 2 introduces temporal modeling capabilities on top of the original SAM, enabling target segmentation in video sequences. Its core architecture consists of three modules:

For industrial quality inspection, we primarily use static image mode, leveraging SAM 2's Prompt Encoder and Mask Decoder zero-shot generalization capabilities.

Prompt Engineering Strategy

In industrial quality inspection, Prompt design directly determines segmentation quality. Through extensive experimentation, we summarized the following Prompt Engineering strategies:

Key finding: in industrial scenarios, combining a box prompt with a center-point prompt worked best, reaching an mIoU of 0.92, eight percentage points above point-only prompting.

Defect Detection Workflow

We designed a two-stage defect detection workflow combining traditional detection model speed with SAM 2 segmentation precision:

This two-stage approach maintains detection accuracy while keeping single-image processing time under 85ms (YOLO 15ms + SAM 2 70ms), meeting production line cycle requirements.

Python Inference Code

Below is the complete defect detection inference code, including model loading, Prompt construction, and post-processing:

inspect_pipeline.py
import torch
import numpy as np
from sam2.build_sam import build_sam2
from sam2.sam2_image_predictor import SAM2ImagePredictor
# ---- Initialize the model ----
sam2 = build_sam2("sam2_hiera_l.yaml", "sam2_hiera_large.pt")
predictor = SAM2ImagePredictor(sam2)
def detect_defects(image, yolo_boxes, threshold=0.7):
"""Two-stage defect detection: YOLO localization + SAM2 segmentation"""
predictor.set_image(image)
defects = []
for box in yolo_boxes:
# Use detection boxes as prompts
masks, scores, _ = predictor.predict(
box=box,
multimask_output=True,
)
# Select the highest-confidence mask
best_idx = np.argmax(scores)
if scores[best_idx] > threshold:
mask = masks[best_idx]
area = mask.sum()
defects.append({
"box": box,
"mask": mask,
"area": int(area),
"score": float(scores[best_idx]),
})
return defects

Domain Fine-Tuning

While SAM 2's zero-shot capability is already impressive, lightweight fine-tuning in specific industrial scenarios can further improve performance. Our fine-tuning strategy:

The fine-tuned model improved mIoU from 0.78 (zero-shot) to 0.91 on our electronic component defect dataset, with defect recall rising from 82% to 95%. Training required only 2 epochs, taking approximately 30 minutes on a single A100.

Edge Deployment Optimization

Production environments typically lack high-performance servers, requiring model deployment on edge devices (like NVIDIA Jetson Orin). Our optimization approach:

After optimization, we achieved 45ms end-to-end inference latency on Jetson Orin NX (YOLO 8ms + SAM 2 Decoder 37ms), meeting production line 20 FPS requirements. GPU utilization stable at 65%, power consumption ~25W.

Core principles for edge deployment: precompute where possible, use distilled models where practical, and quantize instead of relying on floating-point computation. A few milliseconds saved at each stage add up across a production line.

Production Line Results

After deployment on an electronic component factory's SMT production line, the system achieved significant results:

More importantly, the AI system's detection consistency far surpasses human inspectors — it doesn't fatigue, lose focus, or let mood affect judgment. This provides a reliable digital foundation for factory quality control.

Summary & Outlook

SAM 2 brings a paradigm shift to industrial quality inspection: from "training specialized models for each defect" to "using a general foundation model + minimal Prompts to adapt to new scenarios." This dramatically lowers the barrier for AI quality inspection, enabling small and medium factories to quickly deploy AI detection systems.

Future directions we plan to explore:

AI quality inspection isn't a technology project but a continuously operated system. Technology is just the starting point; what truly creates value is the deep integration of technology with production line processes.