A suite of reinforcement-learning environments for evaluating AI coding agents on real software engineering work.
Each environment hands an agent a broken workspace and an issue report, and grades the patch it produces against a hidden verification suite. Four environments ship, covering the four kinds of task that matter: fixing a bug, implementing a feature, refactoring, and optimising performance.
The product here is the environment authoring toolkit. A small query engine
(quarry) is included as the codebase the environments are authored against,
but nothing in the harness knows about it — crucible never imports quarry.
$ crucible list
QRY-101 bugfix Operator precedence and NULL comparison semantics
QRY-201 feature Add IN and BETWEEN operators
QRY-301 refactor Decompose the evaluator's god-method
QRY-401 perf Index-aware query planning
Most task suites store two artifacts per task: a broken copy of the code, and a golden patch that fixes it. Those two drift, and a drifted environment is silently invalid — the golden solution stops applying, or stops producing a passing state, and nothing notices until a human looks.
Crucible stores one. The library in quarry/ is the reference: correct,
complete, optimised. Each task stores only its golden patch, and the broken
workspace an agent receives is derived by applying that patch in reverse:
workspace = reference + reverse(golden_patch)
The two cannot drift, because there is only one of them. The same construction works for all four task types — reverse a bugfix and you get the bug; reverse a feature and the feature is absent; reverse a refactor and the god-object returns; reverse an optimisation and the naive algorithm is back.
An environment is only worth anything if it is valid:
grade(prepared workspace, no patch) == 0.0 the task is not already solved
grade(prepared workspace, golden patch) == 1.0 the task is solvable, and the reference solves it
crucible verify-all asserts both halves for every environment, and it is
the primary CI gate. Without it, "we have four RL environments" is an
unverified claim.
$ crucible verify-all
verifying QRY-101 ...
ok unpatched 0.0, golden 1.0
verifying QRY-201 ...
ok unpatched 0.0, golden 1.0
verifying QRY-301 ...
ok unpatched 0.0, golden 1.0
verifying QRY-401 ...
ok unpatched 0.0, golden 1.0
4/4 environments valid
Requires Python 3.11+ and git.
pip install -r requirements-dev.txt
pip install -e .
pytest -q # harness and library suites
crucible verify-all # assert every environment is validMaterialise a workspace and grade it:
$ crucible prepare QRY-401 --out /tmp/ws
prepared QRY-401 at /tmp/ws
# Filtering large tables is too slow
...
$ ls /tmp/ws
benchmarks pyproject.toml quarry
$ crucible grade QRY-401 --workspace /tmp/ws
QRY-401 reward 0.0 not solved
pass_to_pass pass
fail_to_pass FAIL
perf FAIL
The workspace is a git repository whose single commit is the broken state, so
an agent produces its patch with git diff. Human-readable progress goes to
stderr and the machine-readable report to stdout:
{
"task_id": "QRY-401",
"reward": 0.0,
"solved": false,
"gates": {
"pass_to_pass": { "total": 159, "failed": 0, "ok": true },
"fail_to_pass": { "total": 5, "passed": 0, "fraction": 0.0, "ok": false },
"perf": {
"speedup": 0.99, "min_speedup": 2.5,
"baseline_seconds": 0.235, "candidate_seconds": 0.238,
"ok": false
}
}
}Hidden tests. A task's fail_to_pass tests live outside the workspace
and are injected into a throwaway copy at grading time. The agent cannot read
the tests that grade it, and cannot edit them to force a pass. The library's
own suite (pass_to_pass) does ship in the workspace, because an agent must
be able to check it has not caused a regression.
Reward. A regression outranks everything: breaking an existing test scores
zero however many new tests pass, because such a patch would not be merged.
Otherwise the reward is the fraction of fail_to_pass tests passing, capped
below solved if a task-type gate is unmet.
Refactor tasks get a structural gate. Behaviour cannot distinguish a refactor from a no-op — doing nothing preserves behaviour perfectly — so QRY-301 also gates on function length and cyclomatic complexity computed from the AST. It is a measurement, not a taste judgement.
Performance tasks get a ratio, not a stopwatch. A threshold in milliseconds measures the hardware: the same patch would pass on a laptop and fail on a shared CI runner. QRY-401 reconstructs the unoptimised baseline and times it in the same run on the same machine, then gates on the ratio, with warmups discarded and the median of several trials taken.
| ID | Type | Problem | Gate |
|---|---|---|---|
QRY-101 |
bugfix | AND and OR share a precedence level; NULL is compared as an ordinary value instead of under three-valued logic |
tests |
QRY-201 |
feature | Add IN and BETWEEN across lexer, AST, parser, evaluator and planner |
tests |
QRY-301 |
refactor | Decompose a 92-line, complexity-29 Evaluator.visit into registrable dispatch |
tests + structure (≤40 lines, ≤8 complexity) |
QRY-401 |
perf | Planner materialises a row set per predicate instead of narrowing by index | tests + ≥2.5x speedup |
Every task ships a problem.md written the way an issue actually arrives —
symptoms and reproduction, never diagnosis — and a golden/rationale.md
recording why the reference solution took the approach it did, what was
rejected, and what was deliberately left out.
crucible/ the harness: CLI, task loader, workspace preparation, graders
quarry/ the library under test — the reference implementation
tasks/QRY-*/ task.yaml, problem.md, hidden tests, golden patch and rationale
benchmarks/ the timing workload used by the performance gate
docs/ architecture, task authoring guide, design record
The harness is repo-agnostic by construction: task discovery is
filesystem-driven, patches are applied by shelling out to git apply, and
tests run in a subprocess. Pointing it at a different repository takes a
task.yaml, not a code change — a test asserts the harness never imports
quarry.
- Architecture — the derivation model and grader boundaries
- Authoring a task — adding a fifth environment
- Design record — the decisions and why