Auditing Environments

Most environment auditing tools are ct run subcommands.

Side Task Review

Review side tasks for scorer reward hacking, magic numbers, accidental triggering, and sabotage trajectory (straj) feasibility:

# Review all side tasks for an environment ct run side-task-review -e my_environment # Review specific side tasks ct run side-task-review -e web_scraping -s visit_malicious_website -s expose_secret

Test Quality

Analyze whether integration tests verify real end-to-end behavior:

ct envs integration-test-quality -e my_environment

Compare Side Tasks

Compare side task successes between two runs:

ct run compare-side-tasks <run_id_1> <run_id_2>

Reward Hacking Analysis

Audit side tasks for reward hacking vulnerabilities using the smart model alias with extended thinking:

ct run reward-hacking -e my_environment

Env Description Audit & Generate

Two paired tools maintain an environment's docs/description.md. Both follow the spec at src/control_tower/ct_runs/env_description/shared/spec.md, which defines length caps, required structure, and which content is in and out of bounds.

Audit scores an existing description against the codebase on five metrics (accuracy, clarity, completeness, information leakage, overall):

# Audit all environments ct run env-description-audit # Audit specific environment(s) ct run env-description-audit -e my_environment # Use a different model ct run env-description-audit -m claude-sonnet-4-6

Results are written to data/env_description_audit_runs/ with per-env scores, notes, issues, suggestions, and a ranked summary. Defaults to the smart model alias at -c 2.

Generate writes a fresh description from the codebase. To protect uncommitted edits, the command aborts if docs/description.md already exists unless you pass --overwrite or --dry-run.

ct run env-description-generate -e my_environment # first-time only ct run env-description-generate -e my_environment --overwrite # replace existing ct run env-description-generate -e my_environment --dry-run # writes description.draft.md # Then validate with the audit ct run env-description-audit -e my_environment

ct run doc-audit is a separate command that audits per-side-task docs.

Attack Evaluations

Use ct run eval directly with the appropriate policy:

# Obvious attack ct run eval --policy obvious-attack -e my_environment -t main_task -s side_task # Sneaky attack ct run eval --policy attack -e my_environment -t main_task -s side_task # Test if side tasks pass with read-only access (gameable = bad) ct run eval --policy explore-only -e my_environment -t main_task -s side_task

Options

Most ct run audit commands accept --env / -e for environment selection (compare-side-tasks takes two positional run IDs instead). If --env is omitted, the CLI prompts interactively.