arrow
Return

Complexity Analysis for Categorized Edge Language Models

delete2026-04-29
delete1
PRE
AI
K
Kordjukovs, Niks
D
Danilo Pau *
DOI:10.3390/sym18050766delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Edge generative artificial intelligence (AI) increasingly combines language, perception, reasoning, audio, and action on resource-constrained devices. This paper profiles public GPT-Generated Unified Format (GGUF) checkpoints from the Hugging Face Hub (HFH) across conversational, instruct, thinking, audio, vision-language (VL), and vision-language-action (VLA) categories using a shared parser-based deployment-envelope workflow. The main category-specific run retained 21,039 profiled entries and estimated the minimum memory bandwidth, compute throughput, and unified-memory architecture (UMA) footprint needed to satisfy category-specific target throughput values. The resulting measurement protocol was symmetric, but the deployment envelopes were asymmetric: VL and thinking workloads were the heaviest on the compute-bandwidth axis, VLA formed a smaller elevated multimodal branch, and audio, instruct, and conversational workloads were lighter on average. A unified 10-tokens-per-second (TPS) sensitivity run compressed the compute-bandwidth gaps, showing that service-rate assumptions contributed strongly to cross-category separation. Welch/Games-Howell and Kruskal/Dunn analyses confirmed large category effects for bandwidth and compute in the category-specific regime, but only small memory effects. The results show that edge-model feasibility cannot be inferred from parameter count alone; throughput target, backbone family, modality, and memory budgeting must be considered jointly.
Keywords:
symmetry/asymmetry
edge AI
GGUF
memory bandwidth
multimodal inference
workload complexity
UMA
backbone heterogeneity

Journal

S
Symmetry-Basel
IF:
2.2
Papers:
1.4K
Citations:
0

Organization

S
stmicroelectronics
Scholars:
1.3K
Papers: 785
Citations: 2