Sound is just a wave - and a wave has parameters you can bend: pitch, frequency, amplitude. All it takes is open-source tools and a laptop. Turns out, something unsettling happens: an AI model will refuse a request in text, then comply with the exact same request when it arrives as sound. In 5 rapid-fire minutes, I'll walk through adversarial audio attacks on multimodal AI - the same instruction, reshaped acoustically, sliding past the model's defenses without changing a single word. You'll see how a shifted frequency confuses a multilingual model, a trimmed waveform slips past prompt-injection detection, and a modulated carrier wave evades transcript-based filters - all without touching the text itself. Pitch, frequency, amplitude: the same tools a sound engineer uses to master a podcast are the tools an attacker uses to jailbreak an AI.

Staff AI Security Researcher at Intuit