arrow
Return

Generative Spoken Dialogue Language Modeling

delete2023-03-14
delete16
delete
OA
AI
T
Tu Anh Nguyen *
E
Eugene Kharitonov
J
Jade Copet
Y
Yossi Adi
W
Wei-Ning Hsu
A
Ali Elkahky
P
Paden Tomasello
R
Robin Algayres
B
Benoît Sagot
A
Abdelrahman Mohamed
E
Emmanuel Dupoux
DOI:10.1162/tacl_a_00545delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
We introduce dGSLM, the first textless model able to generate audio samples of naturalistic spoken dialogues. It uses recent work on unsupervised spoken unit discovery coupled with a dual-tower transformer architecture with cross-attention trained on 2000 hours of two-channel raw conversational audio (Fisher dataset) without any text or labels. We show that our model is able to generate speech, laughter, and other paralinguistic signals in the two channels simultaneously and reproduces more naturalistic and fluid turn taking compared to a text-based cascaded model.(1),(2)
Keywords:
TURN-TAKING
ORGANIZATION

Journal

T
Transactions of the Association for Computational Linguistics
IF:
6.9
Papers:
486
Citations:
5.7K

Organization

C
centre national de la recherche scientifique (cnrs)
Scholars:
24.5W
Papers: 18.2W
Citations: 279
I
Inria
Scholars:
3.5K
Papers: 2.5K
Citations: 343
E
ecole normale superieure (ens)
Scholars:
3.0K
Papers: 2.1K
Citations: 4
researcher View more organizations