Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards
About
Imperceptible text-based speech editing modifies spoken content through transcript manipulation while preserving acoustic continuity. Prior acoustic-space approaches suffer from content-style entanglement, causing unstable generation and boundary artifacts. We introduce a framework guided by the principle of "Edit Content, Preserve Acoustics". Editing is conducted in a stable semantic space, while acoustic realization is handled by a Flow Matching decoder. To ensure perceptual consistency, we propose Self-Consistency Rewards Group Relative Policy Optimization, which leverages a pre-trained Text-to-Speech model as an implicit critic, together with intelligibility and duration constraints. Experiments demonstrate consistent improvements over state-of-the-art autoregressive and non-autoregressive baselines in intelligibility, robustness, and perceptual quality.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Speech Editing (Deletion) | Ming-Freeform-Audio-Edit English (full) | DNSMOS3.09 | 14 | |
| Speech Editing (Insertion) | Ming-Freeform-Audio-Edit English (basic) | DNSMOS3.17 | 14 | |
| Speech Editing (Insertion) | Ming-Freeform-Audio-Edit English (full) | DNSMOS3.18 | 14 | |
| Speech Editing (Substitution) | Ming-Freeform-Audio-Edit English (basic) | DNSMOS3.15 | 14 | |
| Speech Editing (Deletion) | Ming-Freeform-Audio-Edit English (basic) | DNSMOS3.09 | 14 | |
| Speech Editing (Substitution) | Ming-Freeform-Audio-Edit English (full) | DNSMOS3.11 | 14 | |
| Deletion Speech Editing | Ming-Freeform-Audio-Edit-Benchmark basic | WER6.91 | 5 | |
| Deletion Speech Editing | Ming-Freeform-Audio-Edit-Benchmark (full) | WER6.88 | 5 | |
| Insertion Speech Editing | Ming-Freeform-Audio-Edit-Benchmark basic | WER4.5 | 5 | |
| Insertion Speech Editing | Ming-Freeform-Audio-Edit-Benchmark (full) | WER4.97 | 5 |