Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards

About

Imperceptible text-based speech editing modifies spoken content through transcript manipulation while preserving acoustic continuity. Prior acoustic-space approaches suffer from content-style entanglement, causing unstable generation and boundary artifacts. We introduce a framework guided by the principle of "Edit Content, Preserve Acoustics". Editing is conducted in a stable semantic space, while acoustic realization is handled by a Flow Matching decoder. To ensure perceptual consistency, we propose Self-Consistency Rewards Group Relative Policy Optimization, which leverages a pre-trained Text-to-Speech model as an implicit critic, together with intelligibility and duration constraints. Experiments demonstrate consistent improvements over state-of-the-art autoregressive and non-autoregressive baselines in intelligibility, robustness, and perceptual quality.

Yong Ren, Jiangyan Yi, Jianhua Tao, Tao Wang, Le Xu, Zhengqi Wen• 2026

Related benchmarks

TaskDatasetResultRank
Speech Editing (Deletion)Ming-Freeform-Audio-Edit English (full)
DNSMOS3.09
14
Speech Editing (Insertion)Ming-Freeform-Audio-Edit English (basic)
DNSMOS3.17
14
Speech Editing (Insertion)Ming-Freeform-Audio-Edit English (full)
DNSMOS3.18
14
Speech Editing (Substitution)Ming-Freeform-Audio-Edit English (basic)
DNSMOS3.15
14
Speech Editing (Deletion)Ming-Freeform-Audio-Edit English (basic)
DNSMOS3.09
14
Speech Editing (Substitution)Ming-Freeform-Audio-Edit English (full)
DNSMOS3.11
14
Deletion Speech EditingMing-Freeform-Audio-Edit-Benchmark basic
WER6.91
5
Deletion Speech EditingMing-Freeform-Audio-Edit-Benchmark (full)
WER6.88
5
Insertion Speech EditingMing-Freeform-Audio-Edit-Benchmark basic
WER4.5
5
Insertion Speech EditingMing-Freeform-Audio-Edit-Benchmark (full)
WER4.97
5
Showing 10 of 12 rows

Other info

Follow for update