arrow
Return

Efficient low-latency packet processing using On-GPU Thread-Data Remapping

delete2019-11-01
delete2
PRE
AI
H
Huanxin Lin *
DOI:10.1016/j.jpdc.2019.06.009delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Graphics processing units are widely-used for packet processing acceleration in both physical and virtual networks. However, real-life packets come in highly-divergent sizes, causing severe GPU control flow divergence. Previous solutions rely on CPU preprocessing to reduce divergence, but it forbids the more efficient NIC-GPU packet streaming as packet batches have to stop completely at host machine. To fully utilize both GPU and PCIe resources, we propose Blink as a GPU modular software router. Instead of CPU pre-processing, the Blink router uses On-GPU Thread-Data Remapping to reduce divergence, and our novel Cross-Iteration Thread Event Signaling mechanism filters unnecessary inter-thread synchronization, doubling the performance gain achieved by traditional solution. Serving as a TCP/IP router with Deep Packet Inspection (DPI) firewall, Blink can sustain processing throughput of 31.5 GBit/s over a PCIe bandwidth of 32 GBit/s. Given a certain bandwidth, Blink reduces processing latency at least by half compared with other works. (C) 2019 Elsevier Inc. All rights reserved.
Keywords:
Packet processing
Software router
GPU control flow divergence
SIMD
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Journal of Parallel and Distributed Computing cover
Journal of Parallel and Distributed Computing
IF:
4
Papers:
3.8K
Citations:
4.8K

Organization

U
University of Hong Kong
Scholars:
4.1W
Papers: 3.9W
Citations: 10.1W
Cited Papers

Cited Papers

Multilayer Packet Classification With Graphics Processing Units
err2016-10-01
err50
errOAAI
errVarvello, Matteo; Laufer, Rafael; Zhang, Feixiong; Lakshman, T. V.
errShare
errSave
Haemoglobin adducts from isoprene and isoprene monoepoxides
err2008-09-22
err0
PREAI
errE. TAREKE; B. T. GOLDING; R. D. SMALL; M. TORNQVIST
errShare
errSave
errShare
errSave
Design and Implementation of a Stateful Network Packet Processing Framework for GPUs
err2017-02-01
err8
PREAI
errVasiliadis, Giorgos; Koromilas, Lazaros; Polychronakis, Michalis; Ioannidis, Sotiris
errShare
errSave
Methanol synthesis over a Zn-deposited copper model catalyst
err1995-01-01
err0
PREAI
errJ. Nakamura; I. Nakamura; T. Uchijima; Y. Kanai; T. Watanabe; M. Saito; T. Fujitani
errShare
errSave
Preparation and Reactivity of Peralkylated Tantalocene Sulfur Complexes Having a Fulvenoid Substructure
err1996-02-20
err0
PREAI
errHenri Brunner; Joachim Wachter; Günther Gehart; Jean-Claude Leblanc; Claude Moïse
errShare
errSave
ClassBench: A packet classification benchmark
err2007-06-01
err340
errOAAI
errTaylor, David E.; Turner, Jonathan S.
errShare
errSave
no more