arrow
Return

Tilus: A Tile-Level GPGPU Programming Language for Low-Precision Computation

delete2026-01-01
delete0
PRE
AI
Y
Yaoyao Ding *
B
Bohan Hou
X
Xiao Zhang
A
A. Binus Y. Lin
T
Tianqi Chen
C
Cody Hao Yu
Y
Yida Wang
G
Gennady Pekhimenko
DOI:10.1145/3760250.3762219delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Serving Large Language Models (LLMs) is critical for AI-powered applications, yet it demands substantial computational resources, particularly in memory bandwidth and computational throughput. Low-precision computation has emerged as a key technique to improve efficiency while reducing resource consumption. Existing approaches for generating low-precision kernels are limited to weight bit widths that are powers of two and suffer from suboptimal performance because of high-level GPU programming abstractions. These abstractions restrict critical optimizations, such as fine-grained register management and optimized memory access patterns, that are essential for efficient low-precision computations. In this paper, we introduce Tilus, a domain-specific language designed for General-Purpose GPU (GPGPU) computing that supports low-precision data types with arbitrary bit widths from 1 to 8 while maintaining GPU programmability. Tilus features a thread-block-level programming model, a hierarchical memory space, a novel algebraic layout system, and extensive support for diverse low-precision data types. Tilus programs are compiled into highly efficient GPU programs through automatic vectorization and instruction selection. Extensive experiments demonstrate that Tilus efficiently supports a full spectrum of low-precision data types, and outperforms state-of-the-art low-precision kernels. Compared to existing compilers such as Triton and Ladder, as well as hand-optimized kernels such as QuantLLM and Marlin, Tilus achieves performance improvements of: 1.75x, 2.61x, 1.29x and 1.03x, respectively. We open-source Tilus at https://github.com/NVIDIA/tilus.
Keywords:
GPU
programming language
parallel computation
low-precision computation
quantization

Journal

P
PROCEEDINGS OF THE 31ST ACM INTERNATIONAL CONFERENCE ON ARCHITECTURAL SUPPORT FOR PROGRAMMING LANGUAGES AND OPERATING SYSTEMS, VOL 1, ASPLOS 2026
IF:
0
Papers:
17
Citations:
0

Organization

A
amazon.com
Scholars:
698
Papers: 505
Citations: 8
C
carnegie mellon university
Scholars:
1.9K
Papers: 940
Citations: 0
U
university of waterloo
Scholars:
2.4K
Papers: 1.3K
Citations: 1
U
university of toronto
Scholars:
14.7W
Papers: 12.0W
Citations: 165
researcher View more organizations