Project Update & Proof of Work:
The evaluation suite and anti-tamper scoring harness are already built, open-sourced, and verified.
- GitHub Repo: https://github.com/j-arndt/long-horizon-containment-suite
- 37/37 unit, property, CLI, and adversarial stress tests passing with 99% coverage on Python 3.12 (green GitHub Actions CI).
- Includes 6 concrete sandbox breakout and containment evaluation tasks.
- Verified against a 60,000-iteration adversarial fuzzing harness proving that spoofed stdout, fake results, or monkeypatching attempts receive a score of exactly 0.0 with zero reward hacking.
Happy to answer any questions or discuss integrations with METR Task Standard v0.5.0 and UK AISI Inspect!