Function Registry

The function registry powers policies, untrusted policies, blue protocols, and the safety / usefulness calculations consumed by ct game-metrics graph pareto. It handles argument declaration, CLI integration, and model configuration. Headings below say which evaluation path a shape belongs to: eval1 is the path behind ct run eval, and eval2 is the separate path behind ct run eval2 (see the Glossary). Monitors are not registry functions but prompt assets resolved by name from the asset directory; see Monitoring.

For the higher-level relationship between policies, untrusted policies, blue protocols, and monitors, see Protocols and Untrusted Policies. For the calculation-side registries used by Pareto plots, see Pareto Frontier.

Two Argument Shapes

Every registry builds a FunctionInfo (defined in function_info.py) for each registered function. There are two shapes for declaring arguments, each backed by a different function_info_factory:

ShapeFactoryArgs declared as…Used by
Args | None (policy eval)function_to_function_info (default — see function_info.py)a single BaseModelPython classpolicies, safety_calculations, usefulness_calculations
Direct kwargs (protocol eval)signature_to_function_info (see signature_info.py)regular keyword parametersuntrusted_policies, blue_protocols

In both shapes, BaseModelPython defaults (e.g. ModelSpec fields) drive the auto sub-args (see Model Arguments and Auto Sub-Args) and the same -p / -up / -bp CLI flags work the same way.

Use the shape of the registry you're targeting; a registry does not mix shapes. The factory is wired up in each registry's registry.py (function_info_factory=signature_to_function_info for protocol eval registries, omitted for policy eval registries).

Defining a Policy (eval1: Args | None)

A policy is a function returning an Agent, decorated with @policies.register(). Its CLI name comes from the function name (see Naming Convention), so honest_policy becomes honest. Policies live in src/control_tower/protocols/policies/policies/.

from control_tower.utils.function_registry.base_model import BaseModelPython from control_tower.models.model_spec import ModelSpec, get_model_spec_field from control_tower.protocols.policies.registry import policies from control_tower.models.aliases import ModelAlias class HonestArgs(BaseModelPython): model: ModelSpec = get_model_spec_field(ModelAlias.UNTRUSTED) enable_main_task_test: bool = False max_steps: int = 200 cache: bool = True @policies.register() def honest_policy(state: TaskState, submit: AgentSubmit, args: HonestArgs | None) -> Agent: """Honest untrusted policy with structured planning.""" if args is None: args = HonestArgs() # ...

The registry discovers decorated policies in the configured packages at import time.

Defining a Safety or Usefulness Calculation (eval1: Args | None)

Safety and usefulness calculations are pluggable impls consumed by ct game-metrics graph pareto (and compute_pareto). They use the eval1 shape: decorate with @safety_calculations.register() or @usefulness_calculations.register() and put the file under src/control_tower/metrics/safety_calculations/impls/ or src/control_tower/metrics/usefulness_calculations/impls/. Each impl returns a single scalar plus metadata; callers such as Pareto handle curve sweeps and optional bootstrap CIs.

from control_tower.utils.function_registry.base_model import BaseModelPython from control_tower.metrics.safety_calculations.registry import ( safety_calculations, ) from control_tower.metrics.safety_calculations.types import ( SafetyResult, ) from inspect_ai.log import EvalSample class TMSafetyArgs(BaseModelPython): mode: str = "optimized" audit_budget: float | None = None @safety_calculations.register() def tm_safety( honest: "list[EvalSample]", attack: "list[EvalSample]", args: TMSafetyArgs | None = None, ) -> "SafetyResult": if args is None: args = TMSafetyArgs() # ... fit rule and compute safety ... return SafetyResult(safety=..., metadata={"mode": args.mode})

Reference the impl from YAML by its CLI name (see Naming Convention), e.g. tm-safety or defer-to-trusted.

Defining an Untrusted Policy or Blue Protocol (protocol eval: direct kwargs)

The protocol eval registries, untrusted_policies and blue_protocols, use signature_to_function_info (see signature_info.py). The function declares its configurable parameters directly as keyword arguments, and the factory synthesizes a BaseModelPython args class from its signature, so ModelSpec fields and other auto sub-args still work. The function body receives concrete values (the model's model.make_model().generate, an int, etc.) rather than an args object.

Untrusted policy

Implementations live under src/control_tower/protocols/untrusted_policy/implementations/<name>/policy.py:

from control_tower.models.model_spec import ( ModelSpec, get_model_spec_field, ) from control_tower.models.aliases import ModelAlias from control_tower.protocols.untrusted_policy.registry import untrusted_policies from control_tower.protocols.untrusted_policy.types import UntrustedPolicyGenerateFn @untrusted_policies.register(name="simple-honest") def simple_honest_policy( model: ModelSpec = get_model_spec_field(ModelAlias.UNTRUSTED), ) -> UntrustedPolicyGenerateFn: """Delegate every generate call to the untrusted model verbatim.""" return model.make_model().generate

The registry's _untrusted_policy_function_info wrapper (in untrusted_policy/registry.py) excludes state from the synthesized args class when the function takes a TaskState, so state is not a CLI flag.

Blue protocol

Implementations live under src/control_tower/protocols/blue_protocols/protocols/<name>.py:

from typing import Any from inspect_ai.solver import Solver, solver from pydantic import Field from control_tower.protocols.blue_protocols.registry import blue_protocols from control_tower.protocols.monitoring.helpers import ( DEFAULT_PROTOCOL_ACTION_MONITOR, ) @blue_protocols.register(name="trusted-monitoring") @solver def trusted_monitoring_protocol( sus_threshold: float = Field(default=5.0), max_steps: int = Field(default=200), action_monitor: str = Field(default=DEFAULT_PROTOCOL_ACTION_MONITOR), ) -> Solver: """Run the untrusted policy and audit (stop) the trajectory if a tool call is too suspicious.""" async def solve(state: TaskState, generate: Generate) -> TaskState: ... return solve

CLI flags target each parameter directly: -up model_temperature=0.5 for an untrusted policy, -bp sus_threshold=7.0 for a blue protocol.

Naming Convention

By default, the registry strips the _<identifier> suffix from the function name and converts underscores to hyphens. Each registry passes its own identifier (see registry.py):

RegistryidentifierAuto-strips suffixExample
policiespolicy_policyobvious_attack_policyobvious-attack
untrusted_policiesuntrusted_policy_untrusted_policyhonest_untrusted_policyhonest
blue_protocolsprotocol_protocoltrusted_monitoring_protocoltrusted-monitoring
safety_calculationssafety_calculation_safety_calculationtm_safetytm-safety
usefulness_calculationsusefulness_calculation_usefulness_calculationmain_task_successmain-task-success

A function name without the matching _<identifier> suffix is only hyphenated, so pass name="..." when that is not the name you want. For example, honest_policy registered against untrusted_policies does not end in _untrusted_policy, so it auto-names to honest-policy unless you pass name="honest". Similarly, simple_attack_policy registered against untrusted_policies needs name="simple-attack" (in implementations/simple_attack/policy.py).

The docstring's first sentence becomes the description shown in help output.

The Args Class

For the eval1 shape, args classes must extend BaseModelPython (not plain Pydantic BaseModel). Fields define configurable parameters with defaults. If a function takes no args parameter, it has no configurable arguments.

For the protocol eval shape, signature_to_function_info synthesizes the args class from the function's keyword parameters, using each parameter's annotation and default.

Field Defaults Become Sub-Arg Defaults

GenerateConfig values passed to get_model_spec_field become the defaults for auto-generated sub-args — e.g. GenerateConfig(cache=False) makes model_cache default to False instead of True, still overridable with -p model_cache=true.

Passing Arguments from the CLI

# Policies use -p (short for --policy-arg) ct run eval --policy honest -p model=sonnet -p max_steps=100 # Untrusted policies use -up (short for --untrusted-policy-arg) ct run eval --untrusted-policy honest -up model=opus --task-file simple # Blue protocols use --protocol with -bp (short for --protocol-arg) ct run eval \ --untrusted-policy simple-attack \ --protocol trusted-monitoring \ -bp sus_threshold=7.0 \ --task-file simple

Hyphens in keys are converted to underscores automatically. Unknown keys raise a validation error. The flag prefix (-p, -up, -bp) corresponds to the registry being targeted, regardless of which argument shape the function uses.

ct run eval takes --policy / -p for policy evals and --untrusted-policy / -up with --protocol / -bp for protocol evals; one invocation cannot combine the two. ct run sabotage-eval selects policies with --attack-policy; adding --protocol / --protocol-arg runs both evals under that blue protocol, with the attack and honest policies as its untrusted policy (see Sabotage Eval).

Model Arguments and Auto Sub-Args

Any field typed as ModelSpec automatically gets sub-arguments for configuring the underlying model: attempt_timeout, max_retries, temperature, max_connections, adaptive_connections, verbosity, effort, reasoning_effort, reasoning_tokens, cache. Sub-arg names are prefixed with the field name, so a field named model gets model_cache, model_temperature, and so on.

To see the full list of sub-args for any policy, run ct protocols policies -v.

Model Aliases

Model fields accept the aliases listed in Model Configuration, or any full Inspect AI provider model name, including self-hosted endpoints (e.g. vllm/Qwen/Qwen3-32B, see Self-Hosted and Open-Weight Models).

Discovering Available Arguments

ct protocols monitors action # List action monitors (add a name for details) ct protocols monitors trajectory # List trajectory monitors ct protocols policies # List all policies ct protocols policies -v # List with all arguments ct protocols policies attack # Show args for a specific policy ct protocols untrusted-policies # List untrusted policies (protocol eval) ct protocols blue-protocols # List blue protocols (protocol eval)

Cost Tracking

CONTROL_TOWER_PRICING selects the rate source:

  • file (default) loads a generated JSON table (~/.cache/control-tower/model_rates.json, or CONTROL_TOWER_MODEL_RATES).
  • api loads the deployment's pricing collection over GET /api/pricing and needs a viewer token (ct login or CONTROL_TOWER_API_TOKEN).
  • off loads nothing. Model calls still run and bill, but every one is recorded as unknown cost, so use it only for runs that make no priced calls. The sabotage-eval preflight skips its pricing check under off.

With file or api, the table loads at Task construction, and a missing or empty table fails the eval before any sample starts. A model the loaded table has no row for stamps unknown cost, not $0.

uv run python scripts/generate_model_rates.py writes the file: vendors' own prices from LiteLLM's price map under bare names, and OpenRouter's catalog under each model's full openrouter/... name, which a lookup tries before the bare name. An OpenRouter row is the default route's price, so a run pinned to an endpoint that prices differently needs a manual row under the same key. A re-run keeps rows marked "manual": true and models the source has stopped publishing.

Automatic per-monitor score cost

Every monitor invocation runs inside a cost-tracking scope. The orchestration layer wraps each call that produces a monitor score with run_monitor_with_cost, which sets a fresh CostSink in a ContextVar for that scope. Any Model.generate() call made while the scope is active writes into the sink via the usage hook, and when the scope exits, the sink's totals are recorded onto the resulting monitor score's metadata.

Fields recorded on a monitor score

  • cost_usd — known spend for this invocation; None only when nothing in the scope could be priced
  • cost_by_model — per-model breakdown; a None amount is a model nobody could price
  • model_name — set when there's exactly one model involved
  • input_tokens / output_tokens / cached_input_tokens — aggregate across models in the scope; always populated when calls happened

Limitations to know

  1. Calls to Model.generate() outside an active CostSink scope are not recorded, which is expected for non-monitor code paths.
  2. If monitor A invokes monitor B without wrapping it in its own run_monitor_with_cost, B's cost lands in A's scope: the total is right, but the split between them is lost. Wrap sub-invocations when you want the split.
  3. Detached tasks created with a fresh empty context (asyncio.create_task(..., context=contextvars.Context())) do not see the parent's sink.

External Policies

Any CLI flag that accepts a policy name also accepts a module:function path, following the same convention as uvicorn app:factory. This lets you use policies defined in your own project without modifying control-tower.

# External policy ct run eval --policy my_project.policies:my_attack -p model=sonnet --task-file simple

The function must have the same signature as a registered one (return Agent for policies). The loader uses the target registry's function_info_factory, so an external function must use that registry's argument shape (see Two Argument Shapes).

External environments also work via --env my_project.envs:my_env_config.

Quick Reference

I want to...Command
List monitorsct protocols monitors action / ct protocols monitors trajectory
List policiesct protocols policies
See a monitor's viewct protocols monitors action <name>
See a policy's argsct protocols policies <name>
Override a model-p model=opus (policy) / -up model=opus (untrusted policy) / -bp model=opus (blue protocol)
Set temperature-p model_temperature=0.5
Disable cache-p model_cache=false
Use external function--policy my.module:my_func