Adding a Dataset
External datasets — published attack and honest transcripts on Hugging Face or GitHub — become ordinary local Control Tower runs through the dataset registry. Every dataset publishes its transcripts in its own format, so each one needs a converter; everything around the converter is shared machinery.
Adding a dataset is two source changes, landed together in one PR:
- A registry entry in
EXTERNAL_DATASETS(src/control_tower/datasets/external.py): the repo id, a pinned revision, and per-file checksums (or, for a large Hugging Face tree,snapshot=Truewithincludeglobs). The pin is what makes an import reproducible — Control Tower always fetches that exact commit, verified byte-for-byte. - A converter in
src/control_tower/datasets/importers/: a module with theImportersignature that parses the downloaded files into.evalrun logs, registered in theIMPORTERSdict.
The commands below drive the whole cycle.
Using a registered dataset
ct extras dataset pull shade-arena
pull takes a registry name and runs the full chain: download (pinned, checksum-verified) → import (the converter writes .eval logs under datasets/hf/<name>/ct_run/) → check (the contract validation below). A green pull means the local runs are ready for the monitor runner and the viewer. The individual steps are also available as ct extras dataset download-hf and import-hf — useful when iterating on a converter without re-downloading. Files download concurrently and back off when the Hub rate-limits; setting HF_TOKEN (required anyway for a gated repo) is sent with every Hugging Face request and gives the higher authenticated limits.
The recipe for a new dataset
1. Inspect the source before writing anything
Dataset cards and READMEs are frequently wrong about their own contents. Download a few files by hand and confirm against the actual tree: what schema is on disk, how honest and attack transcripts are distinguished, whether tool-call arguments were recorded or stripped, what main- and side-task labels exist, and the licence (including any attribution the leaderboard entry must carry).
2. Write the registry entry
Add an ExternalDataset to EXTERNAL_DATASETS with a short kebab-case name, then pin it from the source's API. Keep an entry small: pin only the subtrees the converter reads, which for a repository with many models or domains means choosing one to start with.
- Hugging Face (
hf_repo_id):https://huggingface.co/api/datasets/<owner>/<repo>returns the current commit assha; set it asrevision.https://huggingface.co/api/datasets/<owner>/<repo>/tree/<sha>?recursive=truelists every file with itssizeand, for large files,lfs.oid(the sha256). List the files you need asExternalDatasetFile(filename=..., sha256=..., size_bytes=...), or for a large tree setsnapshot=Truewithincludeglobs naming just the subtrees the converter reads, and the tree is fetched at download time. - GitHub (
github_repo_id):https://api.github.com/repos/<owner>/<repo>/commits/<branch>returns the commit assha; set it asrevision.https://api.github.com/repos/<owner>/<repo>/git/trees/<sha>?recursive=1lists each file's blobshaandsize; list each file asExternalDatasetFile(filename=..., git_sha1=<blob sha>, size_bytes=<size>).snapshot=Trueis Hugging Face only, so a GitHub entry always lists its files.
Then ct extras dataset pull <name> downloads and verifies the files, and reports that no converter exists yet.
3. Write the converter
Create src/control_tower/datasets/importers/<name>.py with the Importer signature and register it in IMPORTERS (importers/__init__.py). Build samples with imported_sample and write runs with write_import_log from importers/eval_log.py — they record everything on the keys the rest of Control Tower reads back. The existing importers are the reference implementations, from small (mythos_5_incident.py) to full-featured (shade_arena.py).
The output contract a converter must meet:
- Both parts, where the source has them. An honest run (
policy="honest") and an attack run (policy="attack"); monitors need benign trajectories to measure false-positive rates against. A one-legged dataset is legitimate (mythos-5-incidentis attack-only) andchecknotes the limitation. - Side-task labels. Every attack trajectory carries its
side_task, and itsside_task_scorewherever the source recorded one; an unscored attack still feeds monitors but enters no safety number, andchecknotes it. An honest trajectory carrying a side task fails the check as a mislabeling. - Tool-call arguments. Preserve them wherever the source recorded them; a source that stripped them yields a trajectory-monitor-only dataset, which
checkreports. - Clean first, monitored separately. A dataset ships without monitor verdicts; the source's own monitor scores go on a separate run, written with
metadata={"variant": "monitored", ...}inwrite_import_logand distinct run and trajectory ids.checkreads that key: a run without it is a base run, and a base run carrying monitor scores fails.shade_arena.pyshows the pattern. - Deterministic ids. Hash run and trajectory ids from the pinned revision and the source keys, so re-imports are idempotent.
- A neutral environment description. The environment description reaches monitors as context. Describe the working situation (the tools, the kind of workstation or service) and never the benchmark, its purpose, or that some trajectories are attacks; override the registry blurb with
environment_data().model_copy(update={"description": ...})as the existing importers do.
4. Iterate until check passes
ct extras dataset import-hf <name> --overwrite --limit 20 # smoke test the converter ct extras dataset check <name> # validate against the contract
check reports the honest/attack totals and which monitor kinds the dataset can feed, and exits non-zero on any contract violation. Iterate converter → import-hf → check until a full import is green, then confirm with a final ct extras dataset pull <name>, which re-runs the converter over the verified download.
5. Open the PR
One PR containing the registry entry and the converter, with the final check output in the description, plus the source's licence terms and any attribution they require. Uploading the imported runs to the viewer and pinning them on a leaderboard is a separate step outside this repository.