← All projects

Evaluation · Open source

SWE-RL Forge Lite

SWE-RL Forge Lite converts a merged GitHub pull request into a self-contained, executable task. It reconstructs the repository before the fix, confirms the historical patch still applies, runs the tests before and after inside Docker, repeats the post-patch run to catch non-determinism, and packages everything with a prompt, a gold patch, a Dockerfile, a reward script, and a quality verdict.

My role
Creator and builder
Started
Updated
Visibility
Open Source
Built with
Python · Docker · GitHub API · Executable rewards · React dashboard
isolated before and after test runs
Docker
test-based reward per task
Binary
quality verdicts: usable, needs review, invalid
3
Forge process monitor dashboard
The live process monitor with task counters, pipeline controls, and a streaming verification log.

Problem

Coding-agent benchmarks often rely on synthetic bugs or on tasks that cannot be verified reliably, which makes rewards noisy.

Solution

A local pipeline that only labels a task as usable when it is reproducible, deterministic, and test-verifiable, so success is grounded in executable evidence.

Architecture

A set of composable CLI stages (explore, fetch, verify, package, reward) writes inspectable artifacts to disk. A live dashboard observes the pipeline and shows quality gates and the training package inventory.

Lessons learned

Most historical pull requests do not make good training tasks. A strict quality gate that marks tasks as usable, needs review, or invalid is what keeps a dataset trustworthy.

More screenshots (1)
From GitHub PRs to learning: project flow
From GitHub pull requests to learning signals: building verifiable environments for coding agents.

coding-agents · reinforcement-learning · evaluation · benchmarks