Return
Benchmarking large language models for marine functional group classification
DOI:10.1016/j.ecoinf.2026.103984.png)
Abstract
En 中文
Functional group classification is a critical but subjective and time-intensive step in ecosystem modeling. This study evaluates whether large language models (LLMs) can automate this process by benchmarking thirteen models across six marine datasets that vary in geographic scope, taxonomic diversity, and size. I evaluated two approaches: (i) classification into predefined functional groups, and (ii) classification into emergent, LLM-determined groupings. I also tested whether success is sensitive to the number of taxa included in each prompt (batch size). Results show that closed-source models (Claude-Sonnet-4, Gemini-Flash-2.5) and large open-source models (e.g. Kimi-K2) achieved moderate agreement with expert-derived groupings and consistently outperformed small open models. Both model size and batch size influenced accuracy. Smaller models degraded as more taxa were classified per prompt, whereas the closed-source and large open models held their performance even when an entire dataset (up to 232 taxa) was classified in a single prompt. These findings suggest that some LLMs can support expert ecological decision-making and offer scalable assistance for ecosystem model development. This work provides a preliminary benchmark for applying LLMs in ecological workflows and highlights their potential to automate complex scientific tasks.
Keywords:
Artificial Intelligence
Clustering
Classification
Large language models
Journal
IF:
7.3
Papers:
3.7K
Citations:
1.3W
Organization
No organization information available

