arrow
Return

A six-tiered framework for evaluating AI models from repeatability to replaceability

delete2025-07-31
delete0
PRE
AI
S
Siqi Tian
A
Alicia Wan Yu Lam
J
Joseph J.�Y. Sung *
W
Wilson Wen Bin Goh *
DOI:10.1016/j.tibtech.2025.07.015delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Current artificial intelligence (AI), particularly generative AI, are not easily evaluated by simplified metrics, creating a need for robust, flexible, and adaptive approaches. Key concepts for AI evaluation, while important, are not standardized –sparking a need for a unified vocabulary. Various disciplines have differing perspectives on AI evaluation but there is no organized framework that systematically addresses the multifaceted demands of AI evaluation and deployment. There is a need for practical, actionable testing methodologies, involving case studies, demonstrating how to apply versatile and comprehensive toolkit for assessing AI performance.
Keywords:
AI evaluation
generative AI
standardized metrics
interdisciplinary perspectives
practical testing methodologies

Journal

Trends in Biotechnology cover
Trends in Biotechnology
IF:
14.9
Papers:
3.8K
Citations:
2.0W

Organization

N
Nanyang Technological University
Scholars:
4.9W
Papers: 4.8W
Citations: 8.1W