arrow
Return

Concept Control for LLM Safety Using Radial Basis Function Representations

delete2026-01-01
delete0
PRE
AI
M
Mark Amos *
杨嵩 cover
杨嵩 (Yang Song)
M
Maurice Pagnucco
DOI:10.1007/978-981-95-4969-6_17delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Representation engineering is a recent approach which has been used to improve the safety of language models. By extracting vector representations of concepts such as 'truthfulness' or 'toxicity', the response of a model can be classified or controlled with respect to the concept. Since the representation engineering literature focuses on linear representations, we present a method to collect activations which can be used to fit an arbitrary non-linear representation. Furthermore, we introduce SurfaceRep, a method which controls concepts via a radial basis function network (RBFN) representation embedded in a low-dimensional space, to investigate whether a more complex, non-linear representation is able to form more accurate local approximations of concepts than linear representations. We evaluate SurfaceRep on several benchmarks for Artificial Intelligence (AI) safety and analyse the strengths and weaknesses of non-linear methods. We find that although the method does not consistently outperform linear methods, it remains useful as an interpretability tool for faithfully visualising representations of concepts.
Keywords:
Large Language Models
Representation Engineering
AI Safety

Journal

A
AI 2025: ADVANCES IN ARTIFICIAL INTELLIGENCE, PT I
IF:
0
Papers:
32
Citations:
0

Organization

U
university of new south wales sydney
Scholars:
2.7K
Papers: 1.2K
Citations: 0