Skip to content

Repository files navigation

Crucible

A suite of reinforcement-learning environments for evaluating AI coding agents on real software engineering work.

Each environment hands an agent a broken workspace and an issue report, and grades the patch it produces against a hidden verification suite. Four environments ship, covering the four kinds of task that matter: fixing a bug, implementing a feature, refactoring, and optimising performance.

The product here is the environment authoring toolkit. A small query engine (quarry) is included as the codebase the environments are authored against, but nothing in the harness knows about it — crucible never imports quarry.

$ crucible list
QRY-101  bugfix    Operator precedence and NULL comparison semantics
QRY-201  feature   Add IN and BETWEEN operators
QRY-301  refactor  Decompose the evaluator's god-method
QRY-401  perf      Index-aware query planning

The idea

Most task suites store two artifacts per task: a broken copy of the code, and a golden patch that fixes it. Those two drift, and a drifted environment is silently invalid — the golden solution stops applying, or stops producing a passing state, and nothing notices until a human looks.

Crucible stores one. The library in quarry/ is the reference: correct, complete, optimised. Each task stores only its golden patch, and the broken workspace an agent receives is derived by applying that patch in reverse:

workspace = reference + reverse(golden_patch)

The two cannot drift, because there is only one of them. The same construction works for all four task types — reverse a bugfix and you get the bug; reverse a feature and the feature is absent; reverse a refactor and the god-object returns; reverse an optimisation and the naive algorithm is back.

The validity invariant

An environment is only worth anything if it is valid:

grade(prepared workspace, no patch)      == 0.0     the task is not already solved
grade(prepared workspace, golden patch)  == 1.0     the task is solvable, and the reference solves it

crucible verify-all asserts both halves for every environment, and it is the primary CI gate. Without it, "we have four RL environments" is an unverified claim.

$ crucible verify-all
verifying QRY-101 ...
  ok  unpatched 0.0, golden 1.0
verifying QRY-201 ...
  ok  unpatched 0.0, golden 1.0
verifying QRY-301 ...
  ok  unpatched 0.0, golden 1.0
verifying QRY-401 ...
  ok  unpatched 0.0, golden 1.0

4/4 environments valid

Quick start

Requires Python 3.11+ and git.

pip install -r requirements-dev.txt
pip install -e .

pytest -q                  # harness and library suites
crucible verify-all        # assert every environment is valid

Materialise a workspace and grade it:

$ crucible prepare QRY-401 --out /tmp/ws
prepared QRY-401 at /tmp/ws

# Filtering large tables is too slow
...

$ ls /tmp/ws
benchmarks  pyproject.toml  quarry

$ crucible grade QRY-401 --workspace /tmp/ws
QRY-401  reward 0.0  not solved
  pass_to_pass pass
  fail_to_pass FAIL
  perf         FAIL

The workspace is a git repository whose single commit is the broken state, so an agent produces its patch with git diff. Human-readable progress goes to stderr and the machine-readable report to stdout:

{
  "task_id": "QRY-401",
  "reward": 0.0,
  "solved": false,
  "gates": {
    "pass_to_pass": { "total": 159, "failed": 0,   "ok": true  },
    "fail_to_pass": { "total": 5,   "passed": 0,   "fraction": 0.0, "ok": false },
    "perf": {
      "speedup": 0.99, "min_speedup": 2.5,
      "baseline_seconds": 0.235, "candidate_seconds": 0.238,
      "ok": false
    }
  }
}

How grading works

Hidden tests. A task's fail_to_pass tests live outside the workspace and are injected into a throwaway copy at grading time. The agent cannot read the tests that grade it, and cannot edit them to force a pass. The library's own suite (pass_to_pass) does ship in the workspace, because an agent must be able to check it has not caused a regression.

Reward. A regression outranks everything: breaking an existing test scores zero however many new tests pass, because such a patch would not be merged. Otherwise the reward is the fraction of fail_to_pass tests passing, capped below solved if a task-type gate is unmet.

Refactor tasks get a structural gate. Behaviour cannot distinguish a refactor from a no-op — doing nothing preserves behaviour perfectly — so QRY-301 also gates on function length and cyclomatic complexity computed from the AST. It is a measurement, not a taste judgement.

Performance tasks get a ratio, not a stopwatch. A threshold in milliseconds measures the hardware: the same patch would pass on a laptop and fail on a shared CI runner. QRY-401 reconstructs the unoptimised baseline and times it in the same run on the same machine, then gates on the ratio, with warmups discarded and the median of several trials taken.

The environments

ID Type Problem Gate
QRY-101 bugfix AND and OR share a precedence level; NULL is compared as an ordinary value instead of under three-valued logic tests
QRY-201 feature Add IN and BETWEEN across lexer, AST, parser, evaluator and planner tests
QRY-301 refactor Decompose a 92-line, complexity-29 Evaluator.visit into registrable dispatch tests + structure (≤40 lines, ≤8 complexity)
QRY-401 perf Planner materialises a row set per predicate instead of narrowing by index tests + ≥2.5x speedup

Every task ships a problem.md written the way an issue actually arrives — symptoms and reproduction, never diagnosis — and a golden/rationale.md recording why the reference solution took the approach it did, what was rejected, and what was deliberately left out.

Layout

crucible/          the harness: CLI, task loader, workspace preparation, graders
quarry/            the library under test — the reference implementation
tasks/QRY-*/       task.yaml, problem.md, hidden tests, golden patch and rationale
benchmarks/        the timing workload used by the performance gate
docs/              architecture, task authoring guide, design record

The harness is repo-agnostic by construction: task discovery is filesystem-driven, patches are applied by shelling out to git apply, and tests run in a subprocess. Pointing it at a different repository takes a task.yaml, not a code change — a test asserts the harness never imports quarry.

Documentation

About

Reinforcement-learning environments for evaluating AI coding agents on software engineering tasks: bugfix, feature, refactor and performance, each with a reproducible workspace and a verified golden solution.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages