Audited three fine-tunes of Qwen2.5-7B-Instruct for covert, weight-encoded loyalties across ~8,700 generations, 127 actors and six trigger classes. Model C was confirmed as the null four independent ways: byte-identical safetensors checksums, bitwise-identical activations at all 29 layers, 1000/1000 verbatim-identical generations at matched seeds, and zero entity deltas on stance probes.
Models A and B were genuinely modified — activation divergence from base peaks at layers 10–12 — but no secret loyalty was detectable in either, under per-actor separability testing, next-token affinity probing, and causal activation patching that came back actor-blind. Reported as a rigorous negative result rather than stretched into a finding.
The part worth pointing at: prefill elicitation extracted specific “confessions” from the known-clean base model at 22.9% — indistinguishable from A at 25.0% and B at 28.1%. A method that produces confessions from a model you can prove is clean cannot be trusted on one you can't.
- 127 actors · 6 trigger classes
- activation patching
- white-box + black-box
- A100-80GB
- negative result, reported as one
repo ↗