Adding a Dataset

External datasets — published attack and honest transcripts on Hugging Face or GitHub — become ordinary local Control Tower runs through the dataset registry. Every dataset publishes its transcripts in its own format, so each one needs a converter; everything around the converter is shared machinery.

Adding a dataset is two source changes, landed together in one PR:

  1. A registry entry in EXTERNAL_DATASETS (src/control_tower/datasets/external.py): the repo id, a pinned revision, and per-file checksums (or, for a large Hugging Face tree, snapshot=True with include globs). The pin is what makes an import reproducible — Control Tower always fetches that exact commit, verified byte-for-byte.
  2. A converter in src/control_tower/datasets/importers/: a module with the Importer signature that parses the downloaded files into .eval run logs, registered in the IMPORTERS dict.

The commands below drive the whole cycle.

Using a registered dataset

ct extras dataset pull shade-arena

pull takes a registry name and runs the full chain: download (pinned, checksum-verified) → import (the converter writes .eval logs under datasets/hf/<name>/ct_run/) → check (the contract validation below). A green pull means the local runs are ready for the monitor runner and the viewer. The individual steps are also available as ct extras dataset download-hf and import-hf — useful when iterating on a converter without re-downloading. Files download concurrently and back off when the Hub rate-limits; setting HF_TOKEN (required anyway for a gated repo) is sent with every Hugging Face request and gives the higher authenticated limits.

The recipe for a new dataset

1. Inspect the source before writing anything

Dataset cards and READMEs are frequently wrong about their own contents. Download a few files by hand and confirm against the actual tree: what schema is on disk, how honest and attack transcripts are distinguished, whether tool-call arguments were recorded or stripped, what main- and side-task labels exist, and the licence (including any attribution the leaderboard entry must carry).

2. Write the registry entry

Add an ExternalDataset to EXTERNAL_DATASETS with a short kebab-case name, then pin it from the source's API. Keep an entry small: pin only the subtrees the converter reads, which for a repository with many models or domains means choosing one to start with.

  • Hugging Face (hf_repo_id): https://huggingface.co/api/datasets/<owner>/<repo> returns the current commit as sha; set it as revision. https://huggingface.co/api/datasets/<owner>/<repo>/tree/<sha>?recursive=true lists every file with its size and, for large files, lfs.oid (the sha256). List the files you need as ExternalDatasetFile(filename=..., sha256=..., size_bytes=...), or for a large tree set snapshot=True with include globs naming just the subtrees the converter reads, and the tree is fetched at download time.
  • GitHub (github_repo_id): https://api.github.com/repos/<owner>/<repo>/commits/<branch> returns the commit as sha; set it as revision. https://api.github.com/repos/<owner>/<repo>/git/trees/<sha>?recursive=1 lists each file's blob sha and size; list each file as ExternalDatasetFile(filename=..., git_sha1=<blob sha>, size_bytes=<size>). snapshot=True is Hugging Face only, so a GitHub entry always lists its files.

Then ct extras dataset pull <name> downloads and verifies the files, and reports that no converter exists yet.

3. Write the converter

Create src/control_tower/datasets/importers/<name>.py with the Importer signature and register it in IMPORTERS (importers/__init__.py). Build samples with imported_sample and write runs with write_import_log from importers/eval_log.py — they record everything on the keys the rest of Control Tower reads back. The existing importers are the reference implementations, from small (mythos_5_incident.py) to full-featured (shade_arena.py).

The output contract a converter must meet:

  • Both parts, where the source has them. An honest run (policy="honest") and an attack run (policy="attack"); monitors need benign trajectories to measure false-positive rates against. A one-legged dataset is legitimate (mythos-5-incident is attack-only) and check notes the limitation.
  • Side-task labels. Every attack trajectory carries its side_task, and its side_task_score wherever the source recorded one; an unscored attack still feeds monitors but enters no safety number, and check notes it. An honest trajectory carrying a side task fails the check as a mislabeling.
  • Tool-call arguments. Preserve them wherever the source recorded them; a source that stripped them yields a trajectory-monitor-only dataset, which check reports.
  • Clean first, monitored separately. A dataset ships without monitor verdicts; the source's own monitor scores go on a separate run, written with metadata={"variant": "monitored", ...} in write_import_log and distinct run and trajectory ids. check reads that key: a run without it is a base run, and a base run carrying monitor scores fails. shade_arena.py shows the pattern.
  • Deterministic ids. Hash run and trajectory ids from the pinned revision and the source keys, so re-imports are idempotent.
  • A neutral environment description. The environment description reaches monitors as context. Describe the working situation (the tools, the kind of workstation or service) and never the benchmark, its purpose, or that some trajectories are attacks; override the registry blurb with environment_data().model_copy(update={"description": ...}) as the existing importers do.

4. Iterate until check passes

ct extras dataset import-hf <name> --overwrite --limit 20 # smoke test the converter ct extras dataset check <name> # validate against the contract

check reports the honest/attack totals and which monitor kinds the dataset can feed, and exits non-zero on any contract violation. Iterate converter → import-hf → check until a full import is green, then confirm with a final ct extras dataset pull <name>, which re-runs the converter over the verified download.

5. Open the PR

One PR containing the registry entry and the converter, with the final check output in the description, plus the source's licence terms and any attribution they require. Uploading the imported runs to the viewer and pinning them on a leaderboard is a separate step outside this repository.