Back to Blog
Vision & Multimodal

Scalable Image Processing Pipeline: From Python to Kubernetes

📅 2025.09.05 ⏱️ 11 min 👤 Eric Pan

Architecture Overview

When building real-time image processing systems, traditional request-response patterns often can't meet the dual demands of high throughput and low latency. We need an event-driven pipeline architecture that splits image processing into multiple independent stages, each independently scalable and optimizable.

Core design principles: decouple, scale, fault-tolerant. Each processing stage is an independent service communicating through message queues, supporting horizontal scaling and fault isolation.

A pipeline's value lies not in the speed of any single stage, but in the coordination and balance of the whole. One bottleneck can slow down the entire chain.

Pipeline Stage Design

We split image processing into 5 core stages, each corresponding to an independent microservice:

Each stage connects through Apache Kafka, supporting backpressure mechanisms to prevent downstream overload. Kafka's partitioning allows us to partition by image ID, ensuring processing order for the same image.

Kafka Streaming Architecture

Kafka serves as the pipeline's core message middleware, handling data flow, backpressure control, and fault recovery. We designed the following Topic structure:

Key design decisions:

Lesson learned: smaller Kafka messages enable higher throughput. A proven approach is to keep large payloads, such as images, in object storage and pass only metadata and references through Kafka.

Worker Concurrency Model

Each pipeline stage's Worker is implemented in Go, leveraging goroutines for high-concurrency processing. The core uses the Worker Pool pattern with channels for task distribution and result collection.

Worker Pool design essentials:

Kubernetes Deployment

Each pipeline stage runs as an independent Kubernetes Deployment. An HPA (Horizontal Pod Autoscaler) scales it according to queue depth. Here is the deployment configuration for the Inference Worker:

k8s/inference-worker.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: inference-worker
namespace: image-pipeline
spec:
replicas: 3
selector:
matchLabels:
app: inference-worker
template:
spec:
containers:
- name: worker
image: registry/inference-worker:v2.4.0
resources:
requests:
memory: "4Gi"
nvidia.com/gpu: "1"
limits:
memory: "8Gi"
nvidia.com/gpu: "1"
livenessProbe:
httpGet:
path: /health
port: 8080
initialDelaySeconds: 30
periodSeconds: 10

Key configuration: minReplicas: 2, maxReplicas: 20, targetQueueLength: 100. When queue depth exceeds the threshold, HPA completes scaling within 30 seconds. GPU nodes use NVIDIA Device Plugin and node affinity for scheduling to dedicated GPU node pools.

Monitoring & Observability

We adopted the three pillars of observability, ensuring every pipeline stage is transparent and controllable:

Key alerting rules:

Scaling Strategies

In production, we adopted a layered scaling strategy, ensuring efficient operation under different loads:

The final system runs stably in production, processing 10M+ images daily, with P99 latency under 200ms and 99.99% availability. During peak events, it successfully handled 5x traffic spikes with auto-scaling completing within 2 minutes.

Summary & Lessons

Building a real-time image processing pipeline is a systems engineering effort spanning architecture design, message middleware selection, container orchestration, and monitoring. Our core lessons:

This architecture has been running stably for over a year, withstanding multiple business peaks and infrastructure changes. If you're building a similar system, I hope these experiences provide useful reference.