arrow
Return

SecNLP: An NLP classification model watermarking framework based on multi-task learning

delete2024-06-01
delete2
PRE
AI
L
Long Dai
J
Jiarong Mao
L
Liaoran Xu
X
Xuefeng Fan
DOI:10.1016/j.csl.2023.101606delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The popularity of ChatGPT demonstrates the immense commercial value of natural language processing (NLP) technology. However, NLP models like ChatGPT are vulnerable to piracy and redistribution, which can harm the economic interests of model owners. Existing NLP model watermarking schemes struggle to balance robustness and covertness. Typically, robust watermarks require embedding more information, which compromises their covertness; conversely, covert watermarks are challenging to embed more information, which affects their robustness. This paper is proposed to use multi-task learning (MTL) to address the conflict between robustness and covertness. Specifically, a covert trigger set is established to implement remote verification of the watermark model, and a covert auxiliary network is designed to enhance the watermark model's robustness. The proposed watermarking framework is evaluated on two benchmark datasets and three mainstream NLP models. Compared with existing schemes, the framework not only has excellent covertness and robustness but also has a lower false positive rate and can effectively resist fraudulent ownership claims by adversaries.
Keywords:
Natural language processing
NLP model security
Black-box watermarking
White-box watermarking

Journal

C
Computer Speech and Language
IF:
3.4
Papers:
1.5K
Citations:
2.6K

Organization

H
Hainan University
Scholars:
2.0W
Papers: 1.2W
Citations: 1.9W