Return
A six-tiered framework for evaluating AI models from repeatability to replaceability
DOI:10.1016/j.tibtech.2025.07.015.png)
Abstract
En 中文
Current artificial intelligence (AI), particularly generative AI, are not easily evaluated by simplified metrics, creating a need for robust, flexible, and adaptive approaches. Key concepts for AI evaluation, while important, are not standardized –sparking a need for a unified vocabulary. Various disciplines have differing perspectives on AI evaluation but there is no organized framework that systematically addresses the multifaceted demands of AI evaluation and deployment. There is a need for practical, actionable testing methodologies, involving case studies, demonstrating how to apply versatile and comprehensive toolkit for assessing AI performance.
Keywords:
AI evaluation
generative AI
standardized metrics
interdisciplinary perspectives
practical testing methodologies
Journal
IF:
14.9
Papers:
3.8K
Citations:
2.0W

