A test of agentic workflows for software engineering
The question: What agent or combination of agents can deliver a spec-consistent PR with real engineering quality at the lowest overall cost and wall-clock time?
Setup
Nine agent workflows were given the same prompt: to implement a detailed written specification in an existing Python codebase.
Specifically, the task was to implement complex structured logging wrapped around a small app with strict accounting requirements. On average, agents wrote ~2,250 new lines of Python code, spanning ~11 new files and 17 files touched in total.
Success was judged by:
- Correctness: linter, formatter, type-checking, and tests all pass, and code was reviewed for hidden bugs.
- Consistency with the spec.
- Hackiness: bypassing linting/type-checking with noqa labels or implementing "fragile" solutions.
- Code readability and maintainability.
- Bonus points for completeness or "future readiness" when the code anticipates future work elegantly.
Evaluations were performed with LLM-as-a-judge using Opus 5.5 and Astra 6 independently, and then averaged.
Harness:
- OpenAI models ran in Codex.
- Anthropic models ran in Claude Code.
- Cursor/SpaceXAI models ran in Cursor.
Costs were calculated at published API rates based on the exact breakdown of token types and amounts recorded in the logs. See the detailed cost breakdown addendum below.
Limitations
- No Fable: special data retention rules are a non-starter.
- No Chinese models: often banned entirely in highly regulated environments.
- Meta Muse Spark 1.3 not tested
- Gemini not tested
- All models ran on "medium" reasoning. See point #6 below.
- Only one trial was run for each workflow.
Results
| Workflow | Cost | Time | Mean quality |
|---|---|---|---|
| GPT-6-Astra + Astra reviewer | $3.99 | 12m | |
| Cursor Composer-2.5 host + Composer subagents + Grok 4.7 medium reviewer | $1.60 | 9m | |
| Opus 5.5 medium @ 1M + Opus reviewer | $6.87 | 65m | |
| Grok 4.7 medium @ 256k coder + Grok reviewer (2 runs) | $9.82 | 60m | |
| Astra 6 host + Luna 6 implementers + Astra reviewer | $6.75 | 40m | |
| GPT-6 Sol medium coder + Sol reviewer | $1.17 | 10m | |
| Grok 4.7 host + Composer subagents + Grok reviewer | $3.51 | 22m | |
| Opus 5.5 host + 18× Haiku 4.5 subagents + Opus reviewer | $21.39 | 110m | — |
My Learnings
-
Opus 5.5 actually wrote the best code.
By every standard of real engineering quality, it was the best. In this experiment, it lost on a technicality: it failed to ensure its linter was passing at the end; namely, it was missing docstrings on test cases. This is a quick fix with prompting, not a limitation of the model. If you are doing foundational engineering work, you want to rely on Opus. Even though it is more expensive and much slower than other "more intelligent" models, it seems to have some secret sauce which makes its work really good.Fable (if you allow it): I'm guessing it would dominate all the others.
-
Arguably the main takeaway here is "just use Astra".
Astra is outstanding. It's ridiculously fast and it's ridiculously token efficient and its quality is excellent. You get frontier intelligence, at speed, and at a very good price. And it doesn't need a lot of special prompting to just get it right.
-
Composer-2.5 is massively underappreciated.
It is a star. It is the only fast, cheap model tested that can actually code with decent speed and quality. It is competing admirably in this benchmark with GPT-6-Astra with just a little code review help from Grok. In this combination, it was significantly faster than Astra, at well under half the price, and at comparable quality.Chinese models, if you allow them, I expect will show a similar value prop.
-
Code reviews are cheap, even with very expensive models.
They don't have long agentic loops racking up millions of cache read tokens. You should always throw the most intelligent models you have at code reviews for a huge ROI.
-
"Cheap" models are not "good enough".
Haiku, Luna, and Sonnet (not pictured) are poor coders. Do not use them. Despite promising benchmarks, in reality they give worse quality, a longer run time, and a higher cost than the "expensive" models.
-
Try bumping up reasoning effort.
Sol was shockingly fast and cheap with medium reasoning. It gave a minimal solution compared to other models, with decent but not great quality. Reasoning is a small portion of the cost. This experiment should be rerun cranking up reasoning to high, xhigh, max. My hypothesis is that cost and time increases will be modest and quality will be much higher, which may turn the tables on these rankings.
-
Try using an expensive model for a quick first draft.
Cost is driven by cache reads during long chains of tool calls, e.g. fixing code in a loop against tests/linters. This experiment should be rerun with a more intelligent model instructed to draft an initial solution without using linters or running tests. Hypothesis: this will give a very strong, intelligently engineered foundation, and then much weaker models could run the long, expensive agentic iterations to pass style and test gates.
Addendum: detailed cost breakdown
Reference pricing
| Model | Input | Cache read | Cache write | Output |
|---|---|---|---|---|
| OpenAI (Codex) | ||||
| GPT-6 Astra | $10.00 | $1.00 | $12.50 | $50.00 |
| GPT-6 Sol | $2.00 | $0.20 | $2.50 | $10.00 |
| GPT-6 Luna | $0.10 | $0.01 | $0.125 | $0.50 |
| Anthropic (Claude Code) | ||||
| Claude Opus 5.5 | $4.00 | $0.20 | $5.00 / $8.00 | $20.00 |
| Claude Haiku 4.5 | $1.00 | $0.10 | $1.25 / $2.00 | $5.00 |
| Cursor | ||||
| Composer 2.5 | $0.50 | $0.20 | — | $2.50 |
| Grok 4.7 (≤256k input) | $2.00 | $0.50 | — | $6.00 |
| Grok 4.7 500k (>256k input) | $4.00 | $1.00 | — | $12.00 |
| Role | Input | Cache read | Output | Cache write | Tokens | Cost |
|---|---|---|---|---|---|---|
| GPT-6-Astra + Astra reviewer | ||||||
| Agent | 76,465 | 1,370,880 | 17,413 | — | — | $3.01 |
| Reviewer | 51,006 | 277,888 | 3,965 | — | — | $0.99 |
| Cost by token type | $1.27 | $1.65 | $1.07 | — | — | $3.99 |
| Total | $3.99 | |||||
| Cursor Composer-2.5 host + Composer subagents + Grok 4.7 medium reviewer | ||||||
| Host | 48,463 | 803,526 | 11,496 | — | 863,485 | $0.21 |
| Composer 1 | 115,467 | 1,958,728 | 18,333 | — | 2,092,528 | $0.50 |
| Composer 2 | 78,433 | 1,068,541 | 17,452 | — | 1,164,426 | $0.30 |
| Composer 3 | 37,532 | 478,105 | 4,847 | — | 520,484 | $0.13 |
| Grok review | 76,535 | 475,776 | 12,658 | — | 564,969 | $0.47 |
| Cost by token type | $0.29 | $1.10 | $0.21 | — | — | $1.60 |
| Total | $1.60 | |||||
| Opus 5.5 medium @ 1M + Opus reviewer | ||||||
| Session | 44,892 | 10,957,605 | 126,000 | 288,836 | 11,418,287 | $6.87 |
| Cost by token type | $0.18 | $2.19 | $2.52 | $2.02* | — | $6.87 |
| Total | $6.87 | |||||
| * Exact breakdown of cache write TTLs were not available, I used $7/million as an approximation. | ||||||
| Grok 4.7 medium @ 256k coder + Grok reviewer (2 runs) | ||||||
| Coder | 786,042 | 13,068,160 | 206,082 | — | 14,060,284 | $9.34 |
| Subagent | 76,267 | 440,448 | 17,318 | — | 534,033 | $0.48 |
| Cost by token type | $1.72 | $6.75 | $1.34 | — | — | $9.82 |
| Total | $9.82 | |||||
| Astra 6 host + Luna 6 implementers + Astra reviewer | ||||||
| Host | 105,377 | 3,028,096 | 6,461 | — | — | $4.40 |
| Luna implementers | 437,280 | 16,608,512 | 126,812 | — | — | $0.27 |
| Reviewer | 124,609 | 519,424 | 6,071 | — | — | $2.07 |
| Cost by token type | $2.34 | $3.71 | $0.69 | — | — | $6.75 |
| Total | $6.75 | |||||
| GPT-6 Sol medium coder + Sol reviewer | ||||||
| Coder | 79,176 | 2,611,840 | 23,160 | — | — | $0.91 |
| Reviewer | 50,838 | 390,528 | 8,068 | — | — | $0.26 |
| Cost by token type | $0.26 | $0.60 | $0.31 | — | — | $1.17 |
| Total | $1.17 | |||||
| Grok 4.7 host + Composer subagents + Grok reviewer | ||||||
| Host | 461,084 | 1,472,896 | 49,079 | — | 1,983,059 | $1.95 |
| Composer subagents | 349,549 | 5,786,085 | 89,281 | — | 6,224,915 | $1.56 |
| Cost by token type | $1.10 | $1.89 | $0.52 | — | — | $3.51 |
| Total | $3.51 | |||||
| Grok 4.7 medium @ 500k coder + reviewer, 1 run | ||||||
| Session | 405,478 | 13,809,792 | 183,032 | — | 14,398,302 | $17.63 |
| Cost by token type | $1.62 | $13.81 | $2.20 | — | — | $17.63 |
| Total | $17.63 | |||||
| Opus 5.5 host + 18× Haiku 4.5 subagents + Opus reviewer | ||||||
| Host orchestration | — | — | — | — | — | ~$3.20 |
| Haiku subagents (18×) | — | — | — | — | — | ~$16.36 |
| Opus review | — | — | — | — | — | ~$1.18 |
| Cost by token type | — | — | — | — | — | $21.39 |
| Total | $21.39 | |||||