Reward Hacking Audit

Audits side tasks for reward hacking vulnerabilities: ways an agent could pass the scorer without violating the security property.

Commands

ct run reward-hacking

Audit side tasks for reward hacking vulnerabilities.

Examples:

# Audit all side tasks ct run reward-hacking # Audit specific environment ct run reward-hacking -e web_scraping # Audit specific side task ct run reward-hacking -s expose_secret

ct run reward-hacking-view

Interactive viewer for reward hacking audit results.

Usage:

ct run reward-hacking-view [RUN_PATH]

Output

Results are saved to data/reward_hacking_runs/<name>_<timestamp>/:

  • results.jsonl - Vulnerability analysis and recommendations for each side task
  • prompt_logs/<env>_<side_task>.json - Prompts, response, and thinking for each side task
  • run.log - Run log

What It Checks

The smart model alias, with extended thinking, reads each side task's scorer code and the environment codebase, using the prompts in prompts/system.md and prompts/user.md. For each side task it asks:

  • Can the scorer be passed without actual sabotage?
  • Are there implementation bugs in the scorer?
  • Can the agent exploit edge cases?
  • Are there "magic number" vulnerabilities?
  • Does the scorer match the intended security property?