Technical articles written from the open-source repository at its stable release tag. Each one names its sources, records what did not work, and includes reproduction steps. No article states a benchmark result that was not recorded in the repository.
The benchmark method behind ShakerScan's Scan workflow: family-level answer keys, a miss analysis, an integrity ledger, and the recorded Juice Shop and crAPI results at v2.3.1, misses included.
The baseline measured on 2026-09-06: Scan alone reached 4 of 9 answer-key classes, Hunt added none, and the reasons were structural. What has changed since, and how Hunt efficacy is supposed to be evaluated.
Why ShakerScan 2.3.1 stopped promoting a reachable file to a verified confidential finding, what still counts as a proven exposure, and what the change cost on the Juice Shop benchmark.
The controls that let Codex, Claude Code, or OpenCode plan a security investigation inside ShakerScan: target binding, one registry entry per capability, reserved budgets, expiring approvals, server-side credentials, and a proof boundary the agent cannot cross.