arrow
Return

AutoML-Pipeline: A RAG-Enhanced Code Generation Framework With Pre-Validation for Cloud-Native Machine Learning Workflows

delete2026-01-01
delete2
PRE
AI
Z
Zhao, Wenyu *
C
Chen, Tingjie
Y
Yang, Jie Si
Q
Qiu, Lei
DOI:10.1109/ACCESS.2026.3673923delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The proliferation of cloud-native machine learning platforms has significantly accelerated model development and deployment cycles. However, constructing and maintaining heterogeneous pipeline code spanning multiple languages (Python, YAML, Spark SQL) and cloud-specific configurations remains labor-intensive and error-prone. Existing LLM-based code generation tools lack awareness of runtime constraints and historical execution patterns, frequently producing code with resource misconfigurations or dependency conflicts that fail upon deployment. To address these challenges, we propose AutoML-Pipeline, a closed-loop code generation framework that integrates Retrieval-Augmented Generation (RAG) with reinforcement learning feedback mechanisms. Our approach leverages a knowledge base constructed from successful pipeline execution logs to guide GPT-4 in generating deployment-ready code that adheres to platform-specific constraints. The key innovation lies in a novel Pre-validation Agent that employs simulated execution environments to predict resource consumption and detect dependency conflicts before actual deployment. This agent iteratively refines generated code through a feedback loop informed by predicted execution profiles and dependency graphs. We evaluate our framework on the CodeSearchNet dataset augmented with Azure ML pipeline specifications, demonstrating a 43.7% improvement in first-submission success rate and 31.2% reduction in resource over-provisioning compared to vanilla GPT-4 baselines. Ablation studies confirm that both the RAG retrieval mechanism and pre-validation agent contribute substantially to performance gains. Our work establishes a practical paradigm for integrating large language models with domain-specific runtime intelligence, with potential applications extending to other infrastructure-as-code generation tasks.
Keywords:
Codes
Pipelines
Retrieval augmented generation
Machine learning
Large language models
Knowledge based systems
Training
Reinforcement learning
Programming
Iterative methods
Code generation
retrieval-augmented generation
cloud computing
automated optimization
large language models

Journal

IEEE Access cover
IEEE Access
IF:
3.6
Papers:
9.8W
Citations:
29.4W

Organization

I
intel china
Scholars:
38
Papers: 32
Citations: 0
U
University of Utah
Scholars:
3.0W
Papers: 2.2W
Citations: 4.6W
U
Utah System of Higher Education
Scholars:
4.6W
Papers: 4.0W
Citations: 161
I
Intel Corporation
Scholars:
2.7K
Papers: 2.0K
Citations: 6
researcher View more organizations