返回
Mosaic: Composite projection pruning for resource-efficient LLMs
DOI:10.1016/j.future.2025.108056.png)
摘要
En 中文
• 没有硬件和软件加速器,优化LLMs很困难。
• 剪枝使LLMs更小、更快,适用于边缘设备。
• 现有的剪枝方法依赖加速器,并降低LLMs的质量。
• Mosaic使用复合剪枝,使LLMs在没有加速器的情况下资源高效。
• Mosaic比现有的剪枝方法具有更高的准确率和更低的困惑度。
Keyword:
Composite projection pruning
Edge computing
Model compression
Large language models
Model pruning
Resource-efficient LLM
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
F
IF:
0
论文数:
642
被引数:
0
机构
引用论文
Rapid Deployment of DNNs for Edge Computing via Structured Pruning at Initialization通过结构化剪枝在初始化阶段快速部署边缘计算中的DNNs
AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and AccelerationAWQ:基于激活感知的权重量化方法,用于设备端大语言模型(LLM)的压缩与加速
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language UnderstandingGLUE: 自然语言理解的多任务基准测试和分析平台
没有更多内容

