arrow
返回

Benchmarking with a Language Model Initial Selection for Text Classification Tasks

delete2025-01-05
delete0
delete
OA
AI
A
Agus Riyadi *
M
Mate Kovacs
U
Uwe Serdült
V
Victor V. Kryssanov
DOI:10.3390/make7010003delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
The now-globally recognized concerns of AI's environmental implications resulted in a growing awareness of the need to reduce AI carbon footprints, as well as to carry out AI processes responsibly and in an environmentally friendly manner. Benchmarking, a critical step when evaluating AI solutions with machine learning models, particularly with language models, has recently become a focal point of research aimed at reducing AI carbon emissions. Contemporary approaches to AI model benchmarking, however, do not enforce (nor do they assume) a model initial selection process. Consequently, modern model benchmarking is no different from a brute force testing of all candidate models before the best-performing one could be deployed. Obviously, the latter approach is inefficient and environmentally harmful. To address the carbon footprint challenges associated with language model selection, this study presents an original benchmarking approach with a model initial selection on a proxy evaluative task. The proposed approach, referred to as Language Model-Dataset Fit (LMDFit) benchmarking, is devised to complement the standard model benchmarking process with a procedure that would eliminate underperforming models from computationally extensive and, therefore, environmentally unfriendly tests. The LMDFit approach draws parallels from the organizational personnel selection process, where job candidates are first evaluated by conducting a number of basic skill assessments before they would be hired, thus mitigating the consequences of hiring unfit candidates for the organization. LMDFit benchmarking compares candidate model performances on a target-task small dataset to disqualify less-relevant models from further testing. A semantic similarity assessment of random texts is used as the proxy task for the initial selection, and the approach is explicated in the context of various text classification assignments. Extensive experiments across eight text classification tasks (both single- and multi-class) from diverse domains are conducted with seven popular pre-trained language models (both general-purpose and domain-specific). The results obtained demonstrate the efficiency of the proposed LMDFit approach in terms of the overall benchmarking time as well as estimated emissions (a 37% reduction, on average) in comparison to the conventional benchmarking process.
Keyword:
language model benchmarking
machine learning model selection
carbon emission reduction
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

M
Machine Learning and Knowledge Extraction
IF:
6
论文数:
838
被引数:
1.8K

机构

R
ritsumeikan university
学者数:
4.0K
论文数: 3.6K
被引数: 0
引用论文

引用论文

err分享
err收藏
Evidence of synergy between Thy-1 and CD3/TCR complex in signal delivery to murine thymocytes for cell death.
err1991-08-15
err0
errOAAI
errI Nakashima; Y H Zhang; S M Rahman; T Yoshida; K Isobe; L N Ding; T Iwamoto; M Hamaguchi; H Ikezawa; R Taguchi
err分享
err收藏
Towards energy-autonomous wake-up receiver using Visible Light Communication
err2016-01-01
err0
errOAAI
errJoyce Sariol Ramos; Ilker Demirkol; Josep Paradells; Daniel Vossing; Karim M. Gad; Martin Kasemann
err分享
err收藏
Visible luminescence in polyaniline/(gold nanoparticle) composites
err2013-01-18
err0
PREAI
errRenata F. S. Santos; Cesar A. S. Andrade; Clecio G. dos Santos; Celso P. de Melo
err分享
err收藏
err分享
err收藏
Wireless Sensor Networks for On-Field Agricultural Management Process
err2010-12-14
err0
errOAAI
errLuca Bencini; Davide Di; Giovanni Collodi; Antonio Manes; Gianfranco Manes
err分享
err收藏
err分享
err收藏
Fault Interactions and Large Complex Earthquakes in the Los Angeles Area
err2003-12-12
err0
PREAI
errGreg Anderson; Brad Aagaard; Ken Hudnut
err分享
err收藏
学者 查看更多内容