The team

One team. The full evaluation loop.

Our shared focus: building useful benchmark tasks, from the initial brief to reference solutions, test environments, and validation.

LoopsForge focuses on the work behind useful agent benchmarks: designing tasks, writing reference solutions, building containerized test environments, and validating the result. We bring these disciplines together to produce task sets for customer-defined evaluation needs.

What we work on

Task design

Translate evaluation goals into scoped tasks with clear instructions, constraints, and completion criteria.

Reference solutions & environments

Develop reference Oracles alongside Docker environments and the dependencies needed to run each task.

Verification & validation

Build verifier tests and review task behavior against the agreed acceptance criteria before delivery.

How we approach the work

Start with the evaluation goal

Agree on the target capabilities, task categories, format version, and delivery scope before authoring begins.

Make the work inspectable

Keep task instructions, reference implementations, environment setup, and verifier behavior explicit rather than relying on an opaque score.

Iterate against evidence

Use execution results and review feedback to refine ambiguous instructions, brittle environments, and incomplete tests.

Need a benchmark we don't have yet?

We forge custom categories to your spec — reference Oracle, Docker suites, validation, the whole loop.

Contact sales