University of MaineElectrical and Computer Engineering

BLADE how twenty flipped bits change what a machine says it sees

Semantic Steering via Differentiable Fault InjectionarXiv preprint · December 2025
BLADE scorecard
Every panel is captioned twice, once with clean weights and once after the flips. A vision judge reads both notes against the image itself.
Clean weights
After the flips
Bits flipped in the last two attention layers
0of roughly fifty billion bits held in the model
Meaning moved
away from image
Sentence kept
same shape
Bit-flip method in use
BLADE
Ranks bits by gradient, keeps only flips that move meaning and hold the grammar.
Attack success
Syntax quality
attack success drawn on a 0 to 25 scale · syntax on 0 to 100
From the paper: three captioning models, two datasets, four bit-flip methods, budgets from 1 to 100 bits
20 bits
flipped inside two attention layers is enough to change what the model says it sees. The model holds about fifty billion bits.
2.4x
the attack success of Progressive Bit Search at the same budget, and about 1.6x plain random flips, on the paper's averaged results.
93 / 96
structure and syntax scores held after the flips. Progressive Bit Search drops structure to about 26, which any reader would catch.
BLADE was evaluated on Flickr8k and COCO with BLIP and BLIP-2 captioners, 100 images each, scored by a GPT-4o-mini judge. The composite panel shown here illustrates the same failure mode, it is not an experiment from the paper. Two caveats the paper itself reports: roughly one note in five is left unchanged, and smaller judge models score the same attacks far lower, so the ratios between methods travel better than the absolute numbers.
Funded by the National Science Foundation and the U.S. Department of Energy
For more details, contact
Zafaryab Haider, lead student researcher zafaryab.haider@maine.edu·Prabuddha Chakraborty, SIEGE Lab PI prabuddha@maine.edu