arrow
Return

Benchmarking large language models for marine functional group classification

delete2026-08-21
delete0
delete
OA
AI
S
Scott Spillias *
DOI:10.1016/j.ecoinf.2026.103984delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Functional group classification is a critical but subjective and time-intensive step in ecosystem modeling. This study evaluates whether large language models (LLMs) can automate this process by benchmarking thirteen models across six marine datasets that vary in geographic scope, taxonomic diversity, and size. I evaluated two approaches: (i) classification into predefined functional groups, and (ii) classification into emergent, LLM-determined groupings. I also tested whether success is sensitive to the number of taxa included in each prompt (batch size). Results show that closed-source models (Claude-Sonnet-4, Gemini-Flash-2.5) and large open-source models (e.g. Kimi-K2) achieved moderate agreement with expert-derived groupings and consistently outperformed small open models. Both model size and batch size influenced accuracy. Smaller models degraded as more taxa were classified per prompt, whereas the closed-source and large open models held their performance even when an entire dataset (up to 232 taxa) was classified in a single prompt. These findings suggest that some LLMs can support expert ecological decision-making and offer scalable assistance for ecosystem model development. This work provides a preliminary benchmark for applying LLMs in ecological workflows and highlights their potential to automate complex scientific tasks.
Keywords:
Artificial Intelligence
Clustering
Classification
Large language models

Journal

Ecological Informatics cover
Ecological Informatics
IF:
7.3
Papers:
3.7K
Citations:
1.3W

Organization

No organization information available