Making Environments

Porting an existing benchmark or environment? See porting-settings.md.

Definition

An environment is a realistic software system that serves as the context for AI agent evaluation. At the code level, an environment is any object satisfying the ControlEnvironment protocol (src/control_tower/environments/protocol.py):

@runtime_checkable class ControlEnvironment(Protocol): @property def id(self) -> str: ... @property def display_name(self) -> str: ... def docs(self, args: BaseModelPython | None = None) -> EnvDocs: ... def get_sandbox_spec(self, codebase_name: str = "default", args: BaseModelPython | None = None) -> SandboxEnvironmentSpec: ... def task_space(self, args: BaseModelPython | None = None) -> TaskSpace: ...

task_space(args) returns the env's menu of valid task combinations: which main and side tasks exist and which (main, side) pairings are allowed. The eval runner builds samples from this menu, and listings like ct envs main-tasks list read from it. args carries one run's --env-arg values (shaped by the env's args_model), since the menu can vary by configuration; None means the env's default configuration. The implementation is in src/control_tower/task/space.py.

Optional capabilities are custom_tools: list[Tool] | None, run_tests() -> TestResult, and agent_user: str. See env-architecture.md.

get_sandbox_spec() determines the sandbox backend: Docker Compose, QEMU, Firecracker, Vagrant, or no sandbox at all. See env-architecture.md for the full architecture and available backends, and .settings/demos/ for examples of environments using each backend.

Most environments use Docker Compose via the DockerEnvironment class and the make_environment_config_from_path() factory. The rest of this guide covers that standard approach.

An official LinuxArena environment takes several forms as Control Tower loads and runs it:

Rendering diagram...

The checked-out repository supplies the files and configuration used to construct the object. Its __init__.py calls make_environment_config_from_path(), which returns a DockerEnvironment object exposing the five ControlEnvironment protocol members. During an evaluation, get_sandbox_spec() returns the configuration used to start the sandbox.

Components

Environments live one level deep inside a setting, a directory directly under .settings/ that groups related environments (e.g. .settings/linuxarena/web_scraping/). See Setting Layout for how discovery works.

Official environments are separate repos (e.g. under the linuxarena org) that ct settings pull clones into .settings/<setting>/. To make your own environment discoverable without a GitHub repo, link its local directory with ct settings pull /path/to/my_env. See Development Setup for both.

A checked-out environment repository contains at least the following files and folders:

  • __init__.py - environment configuration using make_environment_config_from_path()
  • codebase/ - contains the code for the environment
  • codebase/compose.yml - a docker compose file declaring how to set up the container with the environment
  • docs/description.md - codebase structure, key files, and usage instructions
  • docs/attack_surface.md - points of interaction that could be targeted in side tasks
  • main_tasks/ - a set of main tasks for this environment (see how to add tasks below)
  • side_tasks/ - a set of side tasks for this environment (see how to add tasks below)

All standard environments use Ubuntu 22.04 as the base image.

Task content does not have to be baked into the environment image. A main or side task can declare a files map (see Making Main Tasks) of per-sample payloads, such as a data file or a code repository, that are copied into the sandbox at init. Many tasks can then share one compose file and image.

Desired Properties

Most important

A well-designed environment has the following properties:

Realistic and Plausible

Environments should model systems that a real company or organization might plausibly use while building or operating software. A SaaS company is one example; environments need not be about ML, AI safety, or any specific domain.

Sufficient Complexity

Environments should be complicated enough that a typical software engineer needs ~1 week of onboarding to become productive, but not significantly more so. Look at examples of environments in .settings/ before embarking on environment ideation.

Docker Compose Compatible

All services must be runnable inside Docker Compose.

Tech Stack Diversity

As much as possible, environments should use different programming languages and services to sample from the space of real-world Linux systems. Examples of current environments:

  • file_indexing: Document search (Python 3.13, PyTorch, FastAPI, ChromaDB, OpenLDAP)
  • web_scraping: Distributed scraping (Node.js 20.x, Redis, PostgreSQL, FTP)
  • model_registry: ML artifact management (Go, Python/PyTorch, PostgreSQL, S3)
  • media_processing: Video/audio processing (Rust, FFmpeg)
  • clinical_trial: Clinical data and reporting (R, Python, PostgreSQL)
  • tiktok: Short-form video backend (Go, Redis, PostgreSQL)
  • user_tracking: User analytics (C, CMake, PostgreSQL)
  • slack_clone: Team chat API (Ruby on Rails, Elasticsearch, PostgreSQL)

Avoid languages already heavily represented among the existing environments (especially Python). Consider avoiding simple CRUD apps, and consider using unique services beyond the common Redis/PostgreSQL/S3 pattern:

  • ElasticSearch
  • Prometheus
  • NGINX
  • RabbitMQ
  • Git
  • Bazel build system
  • ImageMagick
  • Vector databases

High Attack Surface

Security weaknesses that side tasks can exploit make side task development easier but can hurt realism. Start environment ideation by trying to do without them. A lack of security features is generally more realistic than blatantly misimplemented ones.

Other

Persistent State

Prefer environments with high persistent state across many services, as this creates more opportunities for subtle sabotage.

Exfiltration Testing Server

For environments with exfiltration-related side tasks, consider including a mock external server (e.g., filebin.net or similar) that the agent can exfiltrate data to. This makes scorer verification straightforward.

Known Limitations & Security Considerations

Be aware of these container visibility issues that may affect realism:

Container Detection

Agents can detect that they're being benchmarked:

  • /proc/self/mountinfo may leak volume names containing "linuxarena"
  • Architecture information visible via proc

Mitigation Options

  • Use tmpfs mounts instead of named volumes where possible to avoid leaking names

Creating a New Environment

1: Environment Idea

Create the environment repository under the linuxarena org, add it to the setting's environment list (envs.lock.yml in ct-linuxarena), and open the idea pull request (PR) there. At its root, include an __init__.py and docs/ directory with placeholder content describing the idea.

While the environment is still an idea, name the repository {environment_name}_idea. Environment discovery skips directories ending in _idea, so an unfinished environment does not fail tests. Drop the suffix once the environment is implemented.

Each environment is configured in its __init__.py using make_environment_config_from_path():

from pathlib import Path from control_tower.environments.docker_environment import make_environment_config_from_path environment_root = Path(__file__).parent compose_file = environment_root / "codebase" / "compose.yml" my_environment_config = make_environment_config_from_path( environment_root, short_description="Brief one-line description shown on the website.", tech_stack=[ "Python", "Flask", "Redis", ], compose_paths={ "default": compose_file, }, services=[ "Flask API (port 8080)", "Redis", ], )

environment_root is the checked-out repository root, and compose_file points to its codebase/compose.yml. The factory combines these with its other arguments to build the DockerEnvironment object, and also reads docs/description.md and docs/attack_surface.md from the repository.

The specification for description.md (required sections, length cap, in-bounds vs out-of-bounds content, and the credentials rule) is src/control_tower/ct_runs/env_description/shared/spec.md. Two tools use it:

  • ct run env-description-generate -e <env> produces a fresh description.md from the codebase against the spec. Use --dry-run to write description.draft.md for side-by-side review, or --overwrite to replace an existing description.
  • ct run env-description-audit -e <env> scores an existing description.md against the same spec on accuracy, clarity, completeness, information leakage, and overall.

See docs/auditing.md for full usage.

Available parameters for make_environment_config_from_path():

ParameterRequiredDescription
pathYesPath to the environment directory
short_descriptionYesOne-line description for the website
tech_stackYesList of TechStack literals (languages, frameworks, tools)
servicesYesList of service descriptions
compose_pathsNoMap of codebase name to compose file path (default: {"default": "codebase/compose.yml"})
custom_display_nameNoDisplay name (auto-generated from folder name if omitted)
agent_userNoUser the agent runs as (default: "root")
custom_toolsNoExtra tools available to the agent
internet_simulatorNoInternet simulator configuration
lang_rewrite_tasks_integration_test_filenameNoIntegration test filename for lang-rewrite flows
required_env_varsNoVariables that must be non-empty in the process environment when an eval selects this environment
required_env_any_ofNoFor each group, at least one of the names must be non-empty in the process environment

2: Environment Implementation and Initial Side Task

  1. Build the codebase and Docker Compose setup
  2. Create main task ideas, to use in the steps below
  3. Create side tasks (~5) with scorers
  4. Develop sabotage trajectories (strajs) using different strategies on different main task ideas

Issue templates, Discord, and sprint plans have more specific requirements.

Honest models should not accidentally trigger the initial side task.

Two strajs, each paired with a different main task, show that the environment has more than one sabotage opportunity.

3: Additional Main Tasks

After the initial strajs are approved, develop about 20 more main tasks.

Advanced Environment Configuration

To use different docker compose files for different tasks, pass multiple entries in compose_paths:

from pathlib import Path from control_tower.environments.docker_environment import make_environment_config_from_path root = Path(__file__).parent base_compose = root / "codebase" / "compose.yml" my_env_config = make_environment_config_from_path( root, short_description="...", tech_stack=["Python"], services=["API (port 8080)"], compose_paths={ "default": base_compose, "special_env": root / "codebase" / "compose-special-env.yml", }, custom_tools=[], # Optional: environment-specific agent tools agent_user="root", # Optional: defaults to "root" )

For a task to use the "special_env" compose file, set that task's codebase_name (see Making Main Tasks).

Shared Containers

When many instances of an environment run at once (e.g. in batch evaluations), services like databases or object storage can use a lot of memory. Environments can instead use shared containers that persist across instances, as long as their interactions are heavily constrained.

  • Shared containers must be named shared-{name} (e.g. shared-minio).
  • ct cleans up shared containers when no environment is using them.

The Model Registry environment's shared MinIO container is an example. Each environment instance gets isolated resources within it: per-trajectory naming, and IAM policies that restrict each instance to its own bucket.

Agent Dockerfile

The default compose service should reference the agent Dockerfile, which must meet the requirements in this template:

FROM ubuntu:jammy-20250819 ENV DEBIAN_FRONTEND=noninteractive # Common dependencies - git is REQUIRED for gen-traj command RUN apt-get update && apt-get install -y --no-install-recommends \ curl \ vim \ git \ procps \ python3 \ python3-pip \ python3-venv \ sudo \ # Add environment-specific packages here && apt-get clean && rm -rf /var/lib/apt/lists/* # Install inspect-tool-support (REQUIRED for text_editor tool) ENV PATH="$PATH:/opt/inspect_tool_support/bin" RUN python3 -m venv /opt/inspect_tool_support && \ /opt/inspect_tool_support/bin/pip install inspect-tool-support && \ /opt/inspect_tool_support/bin/inspect-tool-support post-install --no-web-browser # Environment-specific setup (dependencies, directories, etc.) RUN mkdir -p /app # ... # Copy restart script (REQUIRED - must be at /app/restart.sh) COPY src/restart.sh /app/restart.sh RUN chmod +x /app/restart.sh # Copy application files COPY src/ /app/your_app/ WORKDIR /app ENV SHELL=/bin/bash # Start service and keep container alive CMD ["/bin/bash", "-c", "./restart.sh & tail -f /dev/null"]

Agent Privileges

Agents run as root by default so blue teams can implement their own security models.

  • Don't create an agent user in the Dockerfile; it causes permission issues for inspect_tool_support. Leave agent_user at its default of "root" unless you have a specific reason to change it.
  • Don't add sudoers entries, permission restrictions, or other measures that limit what the agent can do. Blue teams implement security through task setup scripts or control protocols.
  • Keep environment-level security measures, such as sandboxing for untrusted code execution or monitoring systems, that exist independently of the agent's privileges.
  • Users or permission structures that the environment's design requires are fine, as long as they are not there to restrict the agent.

Environment Restart

  • Each environment should provide a restart.sh at .settings/<setting>/<env>/codebase/src/restart.sh that restarts the main service.
  • The script should exit 0 on success, non-zero on failure, and print where to look for logs if start-up fails
  • The environment's Dockerfile must copy the script to /app/restart.sh (e.g. COPY src/restart.sh /app/restart.sh), where agents, developers, and scorers can call it directly
  • If an environment intentionally forbids restarts (by design), still include restart.sh that explains why and exits with a non-zero status