arrow
返回

Enabling In-Network Floating-Point Arithmetic for Efficient Computation Offloading

delete2022-12-01
delete1
PRE
AI
H
Heng Pan
Z
Zhenyu Li *
P
Penghao Zhang
J
Jianer Zhou
谢
谢高岗 (Gaogang Xie)
DOI:10.1109/TPDS.2022.3208425delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Programmable switches are recently used for accelerating data-intensive distributed applications. Some computational tasks, traditionally performed on servers in data centers, are offloaded into the network on programmable switches. These tasks may require the support of on-the-fly floating-point operations. Unfortunately, programmable switches are restricted to simple integer arithmetic operations. Existing systems circumvent this restriction by converting floats to integers or relying on local CPUs of switches, incurring extra processing delayed and accuracy loss. To address this gap, we propose NetFC, a table-lookup method to achieve on-the-fly in-network floating-point arithmetic operations nearly without accuracy loss. Specifically, NetFC utilizes logarithm projection and transformation to convert the original huge table enumerating all operands and results into several much smaller tables that can fit into the data plane of programmable switches. To cope with the table inflation problem on 32-bit floats, we also propose an approximation method that further breaks the large tables into smaller ones. In addition, NetFC leverages two optimizations to improve accuracy and reduce on-chip memory consumption. We use both synthetic and real-life datasets to evaluate NetFC. The experimental results show that the average accuracy of NetFC is above 99.9% with only 448KB memory consumption for 16-bit floats and 99.1% with 496KB memory consumption for 32-bit floats. Furthermore, we integrate NetFC into two distributed applications and two in-network telemetry systems to show its effectiveness in further improving the performance.
Keyword:
Open area test sites
Arithmetic
Memory management
Task analysis
Training
Standards
Servers
In-network computation
computation offloading
floating-point operation

期刊

IEEE Transactions on Parallel and Distributed Systems 封面图
IEEE Transactions on Parallel and Distributed Systems
IF:
6
论文数:
5.2K
被引数:
1.1W

机构

U
university of chinese academy of sciences, cas
学者数:
4.1W
论文数: 3.8W
被引数: 75
C
chinese academy of sciences
学者数:
56.7W
论文数: 45.0W
被引数: 704
引用论文

引用论文

Nuclear forward scattering of synchrotron radiation by deoxymyoglobin
err2000-05-19
err0
PREAI
errC. Keppler; K. Achterhold; A. Ostermann; U. van Bürck; A. I. Chumakov; R. Rüffer; W. Sturhahn; E. E. Alp; F. G. Parak
err分享
err收藏
Catalytic enantioselective protonation of samarium enolates by a C2-symmetric chiral diol
err1996-04-01
err0
PREAI
errYutaka Nakamura; Seiji Takeuchi; Akiko Ohira; Yoshiaki Ohgo
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
Histoire naturelle des insectes. Orthoptères
err
IF0
err1839-01-01
err0
errOAAI
errAudinet Serville
err分享
err收藏
Massively parallel recordings in macaque motor cortex during an instructed delayed reach-to-grasp task
err2018-04-10
err0
errOAAI
errThomas Brochier; Lyuba Zehl; Yaoyao Hao; Margaux Duret; Julia Sprenger; Michael Denker; Sonja Grün; Alexa Riehle
err分享
err收藏
学者 查看更多内容