arrow
Return

Scalable and Cost-Efficient ML Inference: Parallel Batch Processing with Serverless Functions

delete2026-01-01
delete0
PRE
AI
A
Amine Barrak *
E
Emna Ksontini
DOI:10.1007/978-981-96-7423-7_27delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
As data-intensive applications grow, batch processing in limited-resource environments faces scalability and resource management challenges. Serverless computing offers a flexible alternative, enabling dynamic resource allocation and automatic scaling. This paper explores how serverless architectures can make large-scale ML inference tasks faster and cost-effective by decomposing monolithic processes into parallel functions. Through a case study on sentiment analysis using the DistilBERT model and the IMDb dataset, we demonstrate that serverless parallel processing can reduce execution time by over 95% compared to monolithic approaches, at the same cost.
Keywords:
Serverless Computing
Function Decomposition
Batch Processing
Scalability
Cost Efficiency

Journal

S
SERVICE-ORIENTED COMPUTING-ICSOC 2024 WORKSHOPS, ASOCA, AI-PA, WESOACS, GAISS, LAIS, AI ON EDGE, RTSEMS, SQS, SOCAISA, SOC4AI AND SATELLITE EVENTS, PT II
IF:
0
Papers:
32
Citations:
0

Organization

O
Oakland University
Scholars:
3.6K
Papers: 3.0K
Citations: 2.7K