Return
Speaker Inference Detection Using Only Text
DOI:10.1007/978-981-95-3537-8_15.png)
Abstract
En 中文
Audio obtained from Internet of Things (IoT) devices can inadvertently disclose personally identifiable information (PII), particularly when combined with related text data. Accordingly, developing robust tools to detect privacy leakage in audio models such as Contrastive Language-Audio Pretraining (CLAP) is imperative. Existing membership inference attacks (MIAs) require audio inputs, which jeopardize voiceprint security and entail costly shadowmodel training. To overcome these limitations, we propose SIDG, a speakerlevel inference detector based exclusively on gibberish text. Our approach generates random text sequences guaranteed to be absent from the training corpus, extracts their feature representations via CLAP, and trains anomaly detectors on these representations. At inference, each test text's feature vector is evaluated by the anomaly detector to determine membership status: anomalous indicates the speaker was present in the training set, whereas normal indicates a nonmember. Furthermore, when real speaker audio is available, SIDG can integrate it to further enhance detection accuracy. Extensive experiments on multiple datasets demonstrate that SIDG outperforms baseline methods that rely solely on text data. Our source code and datasets are available at the anonymous link.
Keywords:
Privacy leakage detection
Membership inference
Contrastive pretraining
Personal identicalinformation
Journal
I
IF:
0
Papers:
23
Citations:
0

