Audio Adversarial Examples: Targeted Attacks on Speech-to-Text
About
We construct targeted audio adversarial examples on automatic speech recognition. Given any audio waveform, we can produce another that is over 99.9% similar, but transcribes as any phrase we choose (recognizing up to 50 characters per second of audio). We apply our white-box iterative optimization-based attack to Mozilla's implementation DeepSpeech end-to-end, and show it has a 100% success rate. The feasibility of this attack introduce a new domain to study adversarial examples.
Nicholas Carlini, David Wagner• 2018
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Environmental Sound Classification | Urbansound8K | Accuracy14.37 | 16 | |
| Speech Command Classification | Google Speech Commands v1 (test) | Accuracy32.69 | 12 | |
| Environmental Sound Classification | DCASE 2019 | Accuracy0.31 | 12 | |
| Speaker Verification | LibriSpeech | Accuracy17.45 | 10 |
Showing 4 of 4 rows