Evaluation Runs

agent_runs is the run-management layer for Control Tower. A run is the metadata container around one evaluation batch:

  • run header metadata (run_id, name, policy, task set, tags, uploader, timestamps)
  • optional eval-log URL
  • associated eval samples exposed as trajectories by the viewer

The ct runs command group uploads eval logs as runs, lists and inspects run headers, and downloads eval archives.

Uploading Results

# Upload during eval ct run eval --policy attack --env web_scraping --main-task crawl_depth --side-task visit_malicious_website # Upload with tags for easier filtering ct run eval --policy honest --env web_scraping --main-task crawl_depth --tags "experiment-1" --tags "baseline" # Upload a previously saved eval ct runs make logs/my-eval.eval

Uploads need a per-user API token: run ct login to cache one, or set CONTROL_TOWER_API_TOKEN. The CLI uploads through the viewer's HTTP API, so it needs no MONGOURI or AWS keys. See Database and the HTTP API reference.

Listing Runs

ct runs list # List recent runs ct runs list --limit 10 # Limit results ct runs list --tag baseline # Filter by tag ct runs list --format json # Output as JSON for scripting

Inspecting a Run

ct runs info <run_id> # Show detailed info ct runs info <run_id> --format json

Downloading Runs

ct runs download-eval <eval_url> # Download an eval log by URL

ct runs download-eval fetches an eval log from the evaluation server by URL, caching it under data/eval-logs/. See Downloading from the Viewer for the URLs the viewer exposes.

To download every sample bundle associated with a run, pass its viewer URL to ct traj download.

ct traj download https://viewer.example/runs/<run_id>