Real-time data processing pipeline built on Apache Flink and Kafka, supporting high-throughput stream computing and visual monitoring — a core hands-on project for big data coursework.
DataFlow Pipeline is a real-time data processing platform built on Apache Flink + Kafka. Users configure data processing tasks through a Vue frontend, which are orchestrated into DAG graphs and executed by the Flink engine, with Kafka as the message middleware for efficient data flow.
The platform supports multiple data source integrations (MySQL, Redis, Kafka Topics) with built-in operators for data cleaning, aggregation, and window computation. The Java backend provides RESTful APIs for the frontend, with Redis caching for real-time task status queries. As a core big data course project, it successfully processed million-level daily simulated data streams.
Full-stack big data solution with stream computing + visual monitoring
Core Flink job definition for real-time Kafka data stream processing