arrow
Return

DiffATSM: High quality adaptive tims-scale modification using diffusion-based post-processing

delete2025-10-15
delete0
PRE
AI
S
Sohee Jang
Y
Yeon-Ju Kim
J
Joon‐Hyuk Chang *
DOI:10.1016/j.csl.2025.101895delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
• We propose DiffATSM, an innovative deep learning-based time-scale modification (TSM) algorithm incorporating a diffusion-based post-processing network. • Unlike conventional TSMs, DiffATSM adaptively generates time-adjusted speech directly from raw waveforms without relying on text information. • DiffATSM is augmented by phonetic posteriorgrams (PPG) which serve as auxiliary information to preserve the phonetic structure of the original signal. • We utilize a post-processing network that leverages a diffusion probabilistic model to accommodate temporal modifications with a flexible distribution.

Journal

C
computer speech & language
IF:
0
Papers:
43
Citations:
0

Organization

No organization information available