Benchmarks

AI workflow benchmarks

Public comparisons of Claude, Codex and Qwen on development workflows.

Status · no results published yet

The first benchmark is in preparation

Nothing is published here until a run is complete. When it is, this page will list it with a link to the full write-up.

What every published run will include

Systems
Each tool and model compared, with the exact version used.
Date
When the run was carried out.
Tasks
The workflows each system was asked to complete, and how many.
Scoring
How results were judged — the methodology, written out.
Results
Charts with the underlying numbers in an accessible table.
Raw data
A downloadable file with the raw results.