Return
Stream-aware and parameter self-adaptive data partitioning algorithm for fluctuating data stream processing
DOI:10.1016/j.future.2026.108520.png)
Abstract
En 中文
With the rapid advancement of big data and real-time stream processing, mainstream Apache Flink is widely used in large-scale data processing and real-time analysis. However, in heterogeneous clusters, its default data partitioning strategy suffers from load imbalance and low resource utilization. Existing studies mostly focus on inter-node load balancing, neglecting inter-task-instance load unevenness and load skewing caused by partitioning policies’ failure to adapt to highly fluctuating data flows. To address these, this paper proposes the Stream-aware and Parameter Self-adaptive Dynamic Partitioning Optimization (SAP-DPO) algorithm, which constructs a stream load model and a dynamic cycle model. Based on these, a stream load-sensing module adjusts partitioning update cycles via real-time data input monitoring (integrating load status and resource utilization) to enhance system adaptability; a parameter optimization mechanism uses a sliding window to tune cycle model parameters for better stability and effectiveness. Experimental results show SAP-DPO outperforms Flink’s default strategy, St-Stream, Dr-Stream, and LADP: it raises Wordcount average throughput by 25.89%, 9.68%, 3.75%, and 5.21% respectively, and reduces Repartition average end-to-end latency by 17.13%, 11.57%, 12.81%, and 5.65%. It improves load balancing and resource utilization, offering a more efficient solution for real-time stream processing systems.
Keywords:
Stream-aware partitioning
Parameter self-adaptive
Load balancing
Real-time stream processing
Data partitioning optimization
Journal
F
IF:
0
Papers:
642
Citations:
0

