Return
DiffATSM: High quality adaptive tims-scale modification using diffusion-based post-processing
DOI:10.1016/j.csl.2025.101895.png)
Abstract
En 中文
• We propose DiffATSM, an innovative deep learning-based time-scale modification (TSM) algorithm incorporating a diffusion-based post-processing network. • Unlike conventional TSMs, DiffATSM adaptively generates time-adjusted speech directly from raw waveforms without relying on text information. • DiffATSM is augmented by phonetic posteriorgrams (PPG) which serve as auxiliary information to preserve the phonetic structure of the original signal. • We utilize a post-processing network that leverages a diffusion probabilistic model to accommodate temporal modifications with a flexible distribution.
Journal
C
IF:
0
Papers:
43
Citations:
0
Organization
No organization information available

