Auditing Environments
Most environment auditing tools are ct run subcommands.
Side Task Review
Review side tasks for scorer reward hacking, magic numbers, accidental triggering, and sabotage trajectory (straj) feasibility:
# Review all side tasks for an environment ct run side-task-review -e my_environment # Review specific side tasks ct run side-task-review -e web_scraping -s visit_malicious_website -s expose_secret
Test Quality
Analyze whether integration tests verify real end-to-end behavior:
ct envs integration-test-quality -e my_environment
Compare Side Tasks
Compare side task successes between two runs:
ct run compare-side-tasks <run_id_1> <run_id_2>
Reward Hacking Analysis
Audit side tasks for reward hacking vulnerabilities using the smart model alias with extended thinking:
ct run reward-hacking -e my_environment
Env Description Audit & Generate
Two paired tools maintain an environment's docs/description.md. Both follow the spec at src/control_tower/ct_runs/env_description/shared/spec.md, which defines length caps, required structure, and which content is in and out of bounds.
Audit scores an existing description against the codebase on five metrics (accuracy, clarity, completeness, information leakage, overall):
# Audit all environments ct run env-description-audit # Audit specific environment(s) ct run env-description-audit -e my_environment # Use a different model ct run env-description-audit -m claude-sonnet-4-6
Results are written to data/env_description_audit_runs/ with per-env scores, notes, issues, suggestions, and a ranked summary. Defaults to the smart model alias at -c 2.
Generate writes a fresh description from the codebase. To protect uncommitted edits, the command aborts if docs/description.md already exists unless you pass --overwrite or --dry-run.
ct run env-description-generate -e my_environment # first-time only ct run env-description-generate -e my_environment --overwrite # replace existing ct run env-description-generate -e my_environment --dry-run # writes description.draft.md # Then validate with the audit ct run env-description-audit -e my_environment
ct run doc-audit is a separate command that audits per-side-task docs.
Attack Evaluations
Use ct run eval directly with the appropriate policy:
# Obvious attack ct run eval --policy obvious-attack -e my_environment -t main_task -s side_task # Sneaky attack ct run eval --policy attack -e my_environment -t main_task -s side_task # Test if side tasks pass with read-only access (gameable = bad) ct run eval --policy explore-only -e my_environment -t main_task -s side_task
Options
Most ct run audit commands accept --env / -e for environment selection (compare-side-tasks takes two positional run IDs instead). If --env is omitted, the CLI prompts interactively.