04 /ExperimentsArticle · 2026-10-01
← Back to experiments

Twin Fixtures, One Verdict: Pinning Scanner Coverage to the Payload Class

Status: portfolio QA artifact. This is the weakest tier of what we publish — a verification note about our own fixtures — and it is published as ONE post deliberately: a known-class positive control does not become two findings by being run through two adapters.

What these fixtures are

Our red-team harness keeps a canonical malicious-pickle positive control: a model file whose __reduce__ reaches posix.system. Unpickling it executes code; that class is old, well documented, and the reason "loading a model is a trust decision" is a sentence anyone in ML security has heard. The twin fixtures run the same payload through two ingestion paths:

  • Primary (modelaudit-5bacb3518633): the modelaudit CLI's static opcode scan, direct.
  • Unified-suite twin (modelaudit-8afc82bacbbd): the same payload delivered through the unified-suite adapter.

Both fire rule S201, pickle_verdict=malicious, clean control passing in the same pass. Deterministic: no judge, no threshold, no sampling — opcodes either match a dangerous pattern or they don't.

What the twin design actually proves (and doesn't)

Scanner "coverage" is usually a number over a corpus, but a corpus mixes variables: payload, container, ingestion path, adapter. The twin isolates one: if detection had followed the adapter instead of the payload, a coverage claim would be an adapter artifact. Here it follows the payload — posix+system reachable from __reduce__ is caught on both paths.

What it does not prove: anything about evasion, novel payload classes, or how the scanner behaves on anything other than this canonical shape. For the broader question — what actually runs when you load a model, and how scanners hold up against repackaging and evasion vectors — see our model-file supply chain study, which is the real treatment of that surface. These fixtures are its tiny, deterministic corner: the positive control that the bigger study's evasion results are anchored against.

The replay honesty note

On the 2026-10-01 replay pass, the twin replayed as MISSING: its scan target was a /tmp file from the original July run, long gone. The fixture rebuilt the target as a byte-identical copy of the committed fixture bytes and the scan fired identically. Two lessons kept:

  1. A replay target that lives in /tmp is a fixture-design smell. The committed fixture bytes are the source of truth; a replay should rebuild from them, not rediscover scratch state.
  2. The finding was never stale — only its scratch target was. Rebuild → rerun → identical verdict: S201, malicious, control clean (modelaudit 0.2.37).

Verification

  • Both fixtures: Promoted (deterministic family — promotion gate reran the real scan) and replay-verified PASS on 2026-10-01.
  • Control: clean file scanned in the same passes, no rule fired.

Reproduce

cd agents/cipher/tools && python3 -m cipher_machine replay-fixtures