返回
Studying logging practice in machine learning-based applications
DOI:10.1016/j.infsof.2024.107450.png)
摘要
En 中文
Context: Logging is a common practice in traditional software development. There have been multiple studies on the characteristics of logging in traditional software systems such as C/C++, Java, and Android applications. However, logging practices in Machine Learning -based (ML -based) applications are still not well understood. The size and complexity of data and models used in ML -based applications present unique challenges for logging. Objective: In this paper, we aim to bridge this knowledge gap and provide insight into the logging practices in ML -based applications, making the first attempt to characterize current logging practices within a large number of open -source ML -based applications. Method: We conducted an empirical study on 502 open -source ML applications to understand their logging practices, combining quantitative and qualitative analyses and a survey involving 31 practitioners. Results: Our quantitative analysis reveals that logging in ML applications is less common than in traditional software, with info and warn log levels being popular. Top ML -specific logging libraries include MLflow, Tensorboard, Neptune, and W&B. Qualitatively, logging is used for data and model management, especially in model training. Our survey reinforces the importance of logging in experiment tracking, complementing our qualitative findings. Conclusion: Our research carries significant implications. It reveals distinctive ML logging practices compared to traditional software. We have highlighted the prevalence of general-purpose logging libraries in ML code, indicating a potential gap in awareness regarding ML -specific logging tools. This insight benefits researchers and developers aiming to enhance ML project reproducibility and sets the stage for exploring ML -specific logging tools' impact on machine learning system quality and trustworthiness.
Keyword:
Logging practices
ML -based applications
Mining software repositories
Source code analysis
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
4.3
论文数:
3.8K
被引数:
7.7K
机构
引用论文
Coding In-depth Semistructured Interviews: Problems of Unitization and Intercoder Reliability and Agreement编码深入的半结构化访谈: uniization和Intercoder可靠性和协议的问题

