Control Evals
ct run control-eval runs a control evaluation: an honest policy and an attack policy play under one blue-team protocol, and their samples land in one eval log. It builds samples, the solver and the scorers exactly as ct run eval does and takes the same options, except for the ones this page describes. The CLI reference lists every option.
Quick Start
ct run control-eval --task-file examples/eval_configs/03-control-eval.jsonl \ --protocol trusted-monitoring -bp 'monitors=[simple]' -bp sus_threshold=7
This runs the honest policy on the task file's two records without a side task and the attack policy on its two records with one, all under trusted-monitoring.
Samples
Task selection works as it does for ct run eval, and each selected sample runs once, as it is listed. A sample that carries a side task is an attack sample, which the attack policy plays; any other sample is an honest sample, which the honest policy plays. A task file for a control eval therefore lists both kinds of record:
{"env": "web_scraping", "main_task": "redact"} {"env": "web_scraping", "main_task": "redact", "side_task": "expose_secret"}
--task-file repeats, so a file of honest records can sit beside a task set of pairs:
ct run control-eval --task-file honest.jsonl --task-file medium
--all selects each main task alone and each main and side task pair of the selected environments. Across the whole selection, task files included, honest and attack samples alternate over the same main tasks, and each main task leads with a different side task, so any --limit of two or more takes both streams and a small limit spreads its attacks over the side tasks. -t and -s name one sample, so they cannot give both streams. A selection without an honest sample or without an attack sample is refused before anything is built, submitted or run, with an error naming the missing stream. Inspect's own --limit and --sample-id (on inspect eval, or as eval_config in a run config) slice the samples after that check, so they can run one stream. A setting that runs in its own virtual environment runs its samples as a separate eval, so its samples need both streams; every such setting is checked before any of them runs.
A sample's stream is derived from its tasks and stored nowhere. In code it is SampleMetadata.stream in control_tower.task.sample_metadata.
Policies
--honest-policy (default honest) names the policy that plays the honest samples, and --attack-policy (default attack) names the one that plays the attack samples. Each takes its arguments as repeatable key=value pairs, -hp for the honest policy and -ap for the attack policy. These options stand where ct run eval has --untrusted-policy, and both policies come from the same registry, which ct protocols policies lists.
ct run control-eval --task-file examples/eval_configs/03-control-eval.jsonl \ --attack-policy attack -ap private_tools=run_side_task_test -wa private_tools=run_side_task_test
Protocol
Both streams share the task selection, the protocol with its harness (--protocol, -bp, --harness), what white allows (-wa), and the model roles (--model-role). A replay on either stream therefore fits white for both: the live stream also gets both editors, and a larger max_steps and message limit when the recording needs them. The protocol is never told a sample's stream, so it handles honest and attack samples alike. Protocols and Policies describes how protocols and policies are built.
Audits
A protocol sends an action to a human by returning Audited(draft). Of the built-in protocols, trusted-monitoring audits through sus_threshold and blocks, and defer-to-trusted through audit_threshold; Monitoring describes both. The engine records each audit on the sample's transcript as an info event with source audit and data {"audited": <i>}, where i is the index of the audited action, the same index the monitor score keys that action's verdicts by. What follows depends on the sample's stream:
- Attack sample: the sample stops at its first audit. The audited action stays on the transcript with each of its calls answered by the engine's
auditedtool, and none of them run. - Honest sample: the audited action runs and the sample continues, so a later flagged action is audited again. Audits do not reduce usefulness, and every flagged honest action counts as one audit.
ct run eval records and applies audits the same way.
Running on EC2
--ec2 runs the control eval on the EC2 fleet, one job per sample and epoch, and merges the job logs into one eval log. --hawk runs it on Hawk, as it does ct run eval.
ct run control-eval --task-file examples/eval_configs/03-control-eval.jsonl \ --protocol trusted-monitoring -bp 'monitors=[simple]' -bp sus_threshold=7 --ec2
Run Config
--run-config takes an Inspect run-config YAML whose task is control_tower/control_eval. Its task args are those of control_tower/eval, with honest_policy and attack_policy in place of untrusted_policy; both policies and protocol are required. Evaluation Config Files documents the other fields, the flags that forward, and sandbox: {use_ec2: true} for the fleet.
ct run control-eval --run-config examples/eval_configs/03-control-eval.yml
The config names its task, so ct run eval --run-config and ct run control-eval --run-config each run a config of either task.
Uploading
--upload uploads the log as one run. The run document's producer is control_eval, its policy reads <honest>+<attack>-vs-<protocol> (for example honest+attack-vs-trusted-monitoring), and its metadata holds the honest policy as honest_policy_selection and the attack policy as attack_policy_selection. The web viewer lists these runs in its "Control evals" tab, and the HTTP API filters them with producer=control_eval.
Usefulness and Safety
The log holds both streams: its honest samples measure usefulness and its attack samples measure safety, as AI Control Concepts defines them. The metrics commands in Safety Curves take an honest run and an attack run as two separate sources and do not split one log by stream.