A test of agentic workflows for software engineering

The question: What agent or combination of agents can deliver a spec-consistent PR with real engineering quality at the lowest overall cost and wall-clock time?

Setup

Nine agent workflows were given the same prompt: to implement a detailed written specification in an existing Python codebase.

Specifically, the task was to implement complex structured logging wrapped around a small app with strict accounting requirements. On average, agents wrote ~2,250 new lines of Python code, spanning ~11 new files and 17 files touched in total.

Success was judged by:

Evaluations were performed with LLM-as-a-judge using Opus 5.5 and Astra 6 independently, and then averaged.

Harness:

Costs were calculated at published API rates based on the exact breakdown of token types and amounts recorded in the logs. See the detailed cost breakdown addendum below.

Limitations

Results

Cost versus mean quality for each agent workflow Scatter plot of total cost per PR on a log scale against mean quality score. Bubble area is proportional to wall-clock time. The Pareto frontier runs from GPT-6 Sol ($1.17, 5.90) to Composer-2.5 with a Grok reviewer ($1.60, 7.95) to GPT-6-Astra ($3.99, 8.05); every other workflow is beaten on both cost and quality by a workflow on the frontier. 5.0 5.5 6.0 6.5 7.0 7.5 8.0 8.5 Mean quality score (0–10) not scored $1 $2 $5 $10 $20 Total cost per PR (USD, log scale) Pareto frontier Dominated: a frontier workflow is cheaper and better ↖ cheaper and better Opus 5.5 medium @ 1M + Opus reviewer: $6.87, 65 min, quality 7.50* Grok 4.7 medium @ 256k coder + Grok reviewer (2 runs): $9.82, 60 min, quality 7.15 Astra 6 host + Luna 6 implementers + Astra reviewer: $6.75, 40 min, quality 6.40 Grok 4.7 host + Composer subagents + Grok reviewer: $3.51, 22 min, quality 5.40 GPT-6-Astra + Astra reviewer: $3.99, 12 min, quality 8.05 GPT-6 Sol medium coder + Sol reviewer: $1.17, 10 min, quality 5.90 Cursor Composer-2.5 host + Composer subagents + Grok 4.7 medium reviewer: $1.60, 9 min, quality 7.95 GPT-6-Astra + Astra reviewer$3.99 · 12m · 8.05 Composer-2.5 + Grok reviewer$1.60 · 9m · 7.95 Opus 5.5 + Opus reviewer$6.87 · 65m · 7.50* Grok 4.7 + Grok reviewer$9.82 · 60m · 7.15 Astra 6 + Luna 6 implementers$6.75 · 40m · 6.40 GPT-6 Sol + Sol reviewer$1.17 · 10m · 5.90 Grok 4.7 + Composer subagents$3.51 · 22m · 5.40 Opus 5.5 host + 18× Haiku 4.5 subagents + Opus reviewer: $21.39, 110 min, not scored Opus 5.5 + 18× Haiku 4.5$21.39 · 110m

Workflow Cost Time Mean quality
GPT-6-Astra + Astra reviewer $3.99 12m
8.05
Cursor Composer-2.5 host + Composer subagents + Grok 4.7 medium reviewer $1.60 9m
7.95
Opus 5.5 medium @ 1M + Opus reviewer $6.87 65m
7.50*
Grok 4.7 medium @ 256k coder + Grok reviewer (2 runs) $9.82 60m
7.15
Astra 6 host + Luna 6 implementers + Astra reviewer $6.75 40m
6.40
GPT-6 Sol medium coder + Sol reviewer $1.17 10m
5.90
Grok 4.7 host + Composer subagents + Grok reviewer $3.51 22m
5.40
Opus 5.5 host + 18× Haiku 4.5 subagents + Opus reviewer $21.39 110m —

My Learnings

  1. Opus 5.5 actually wrote the best code.
    By every standard of real engineering quality, it was the best. In this experiment, it lost on a technicality: it failed to ensure its linter was passing at the end; namely, it was missing docstrings on test cases. This is a quick fix with prompting, not a limitation of the model. If you are doing foundational engineering work, you want to rely on Opus. Even though it is more expensive and much slower than other "more intelligent" models, it seems to have some secret sauce which makes its work really good.
    Fable (if you allow it): I'm guessing it would dominate all the others.
  2. Arguably the main takeaway here is "just use Astra".
    Astra is outstanding. It's ridiculously fast and it's ridiculously token efficient and its quality is excellent. You get frontier intelligence, at speed, and at a very good price. And it doesn't need a lot of special prompting to just get it right.
  3. Composer-2.5 is massively underappreciated.
    It is a star. It is the only fast, cheap model tested that can actually code with decent speed and quality. It is competing admirably in this benchmark with GPT-6-Astra with just a little code review help from Grok. In this combination, it was significantly faster than Astra, at well under half the price, and at comparable quality.
    Chinese models, if you allow them, I expect will show a similar value prop.
  4. Code reviews are cheap, even with very expensive models.
    They don't have long agentic loops racking up millions of cache read tokens. You should always throw the most intelligent models you have at code reviews for a huge ROI.
  5. "Cheap" models are not "good enough".
    Haiku, Luna, and Sonnet (not pictured) are poor coders. Do not use them. Despite promising benchmarks, in reality they give worse quality, a longer run time, and a higher cost than the "expensive" models.
  6. Try bumping up reasoning effort.
    Sol was shockingly fast and cheap with medium reasoning. It gave a minimal solution compared to other models, with decent but not great quality. Reasoning is a small portion of the cost. This experiment should be rerun cranking up reasoning to high, xhigh, max. My hypothesis is that cost and time increases will be modest and quality will be much higher, which may turn the tables on these rankings.
  7. Try using an expensive model for a quick first draft.
    Cost is driven by cache reads during long chains of tool calls, e.g. fixing code in a loop against tests/linters. This experiment should be rerun with a more intelligent model instructed to draft an initial solution without using linters or running tests. Hypothesis: this will give a very strong, intelligently engineered foundation, and then much weaker models could run the long, expensive agentic iterations to pass style and test gates.

Addendum: detailed cost breakdown

Reference pricing

Model Input Cache read Cache write Output
OpenAI (Codex)
GPT-6 Astra $10.00 $1.00 $12.50 $50.00
GPT-6 Sol $2.00 $0.20 $2.50 $10.00
GPT-6 Luna $0.10 $0.01 $0.125 $0.50
Anthropic (Claude Code)
Claude Opus 5.5 $4.00 $0.20 $5.00 / $8.00 $20.00
Claude Haiku 4.5 $1.00 $0.10 $1.25 / $2.00 $5.00
Cursor
Composer 2.5 $0.50 $0.20 — $2.50
Grok 4.7 (≤256k input) $2.00 $0.50 — $6.00
Grok 4.7 500k (>256k input) $4.00 $1.00 — $12.00

Role Input Cache read Output Cache write Tokens Cost
GPT-6-Astra + Astra reviewer
Agent 76,465 1,370,880 17,413 — — $3.01
Reviewer 51,006 277,888 3,965 — — $0.99
Cost by token type $1.27 $1.65 $1.07 — — $3.99
Total $3.99
Cursor Composer-2.5 host + Composer subagents + Grok 4.7 medium reviewer
Host 48,463 803,526 11,496 — 863,485 $0.21
Composer 1 115,467 1,958,728 18,333 — 2,092,528 $0.50
Composer 2 78,433 1,068,541 17,452 — 1,164,426 $0.30
Composer 3 37,532 478,105 4,847 — 520,484 $0.13
Grok review 76,535 475,776 12,658 — 564,969 $0.47
Cost by token type $0.29 $1.10 $0.21 — — $1.60
Total $1.60
Opus 5.5 medium @ 1M + Opus reviewer
Session 44,892 10,957,605 126,000 288,836 11,418,287 $6.87
Cost by token type $0.18 $2.19 $2.52 $2.02* — $6.87
Total $6.87
* Exact breakdown of cache write TTLs were not available, I used $7/million as an approximation.
Grok 4.7 medium @ 256k coder + Grok reviewer (2 runs)
Coder 786,042 13,068,160 206,082 — 14,060,284 $9.34
Subagent 76,267 440,448 17,318 — 534,033 $0.48
Cost by token type $1.72 $6.75 $1.34 — — $9.82
Total $9.82
Astra 6 host + Luna 6 implementers + Astra reviewer
Host 105,377 3,028,096 6,461 — — $4.40
Luna implementers 437,280 16,608,512 126,812 — — $0.27
Reviewer 124,609 519,424 6,071 — — $2.07
Cost by token type $2.34 $3.71 $0.69 — — $6.75
Total $6.75
GPT-6 Sol medium coder + Sol reviewer
Coder 79,176 2,611,840 23,160 — — $0.91
Reviewer 50,838 390,528 8,068 — — $0.26
Cost by token type $0.26 $0.60 $0.31 — — $1.17
Total $1.17
Grok 4.7 host + Composer subagents + Grok reviewer
Host 461,084 1,472,896 49,079 — 1,983,059 $1.95
Composer subagents 349,549 5,786,085 89,281 — 6,224,915 $1.56
Cost by token type $1.10 $1.89 $0.52 — — $3.51
Total $3.51
Grok 4.7 medium @ 500k coder + reviewer, 1 run
Session 405,478 13,809,792 183,032 — 14,398,302 $17.63
Cost by token type $1.62 $13.81 $2.20 — — $17.63
Total $17.63
Opus 5.5 host + 18× Haiku 4.5 subagents + Opus reviewer
Host orchestration — — — — — ~$3.20
Haiku subagents (18×) — — — — — ~$16.36
Opus review — — — — — ~$1.18
Cost by token type — — — — — $21.39
Total $21.39