Evaluation · Open source
SWE-RL Forge Lite
SWE-RL Forge Lite converts a merged GitHub pull request into a self-contained, executable task. It reconstructs the repository before the fix, confirms the historical patch still applies, runs the tests before and after inside Docker, repeats the post-patch run to catch non-determinism, and packages everything with a prompt, a gold patch, a Dockerfile, a reward script, and a quality verdict.
- My role
- Creator and builder
- Started
- Updated
- Visibility
- Open Source
- Built with
- Python · Docker · GitHub API · Executable rewards · React dashboard
- isolated before and after test runs
- Docker
- test-based reward per task
- Binary
- quality verdicts: usable, needs review, invalid
- 3

Problem
Coding-agent benchmarks often rely on synthetic bugs or on tasks that cannot be verified reliably, which makes rewards noisy.
Solution
A local pipeline that only labels a task as usable when it is reproducible, deterministic, and test-verifiable, so success is grounded in executable evidence.
Architecture
A set of composable CLI stages (explore, fetch, verify, package, reward) writes inspectable artifacts to disk. A live dashboard observes the pipeline and shows quality gates and the training package inventory.
Lessons learned
Most historical pull requests do not make good training tasks. A strict quality gate that marks tasks as usable, needs review, or invalid is what keeps a dataset trustworthy.
More screenshots (1)

coding-agents · reinforcement-learning · evaluation · benchmarks