Start with one recording. The clean, quantized model writes 38. The same audio is used in every condition.
Paper evidence and limits
What is measured here
Appendix N, Tables 31 and 34 report this exact Whisper-small.en × controlled LACC sample. The clean output is “The thermometer read 38 degrees Celsius during the afternoon check.” SMASH outputs “The thermometer read 48 degrees celsius during the afternoon check.” The accepted configuration costs 2 INT8 bit flips under a 20-bit maximum budget. The same-sample PBS-CE-gated output changes only casing; URBF leaves the transcript unchanged.
The memory cells and waveform are schematics. The page does not contain the experimental WAV or claim to display exact tensor coordinates. It replays the method and reported outputs; no model, gradient, or physical memory exploit runs in the browser.
How success is counted
Appendix B requires a normalized transcript change, an automatic readability check, word-error rate at most 0.20 from the baseline, and the declared replacement at the selected span. Case-only differences do not pass the first check. The human audit is sampled: on 105 method-mixed outputs the composite human-majority and four-judge labels agree at κ=1.00; individual questions show κ=0.51–0.85 (Table 2).
Results in context
Table 1 reports numeric-target successes at B=20 on controlled LACC speech: 24/45 for Whisper-small.en, 19/44 for Whisper-large-v3, and 5/34 for SeamlessM4T-v2-large. On real-speech MultiMed the counts are 4/20, 2/26, 0/25; on LibriSpeech they are 4/9, 0/8, 0/10, in the same model order. Across 1,192 successful certificates, the pooled median realized cost is 3 bits; model medians are 2, 11, and 12.
Table 3 compares all-pivot WS × LACC results at B=20: SMASH-hybrid 26/100, PBS-CE-gated 0/180, URBF 0/100 strict hits. The PBS cell has a different denominator. These methods optimize different events, so this does not prove SMASH is optimal against a future target-aware method.
What the work does not show
No demonstrated physical fault injection or deployed-system compromise.
No downstream harm measurement and no defense.
Controlled TTS gives stronger reachability than the two real-speech datasets. The strongest result is numeric; entity evidence is weaker and the reported negation cells have zero strict successes.
Source: supplied manuscript “SMASH: Probing Speech Recognition Robustness via Semantically Targeted Bit Flips,” Sections 4–7, Tables 1–3, Appendix B, Appendix N Tables 31 and 34. Funding and contacts are provided by the project team.