Cost per task — Benchmark Sources & Consensus
Total dollar cost to complete a representative agent workflow.
Platforms tracked: Openclaw · Nemoclaw · Ironclaw · Hermes · Claude Cowork · Chatgpt
Consensus across 5 sources
Across 5 sources, cost tracks call shape, harness and model tier far more than platform. The strongest evidence is a controlled comparison holding the model and the task set fixed: three harnesses over 64 SWE-Bench Pro tasks land within 44-53% resolve rate of each other while costing up to 2x apart on the same model, with the expensive one generating 1.35x the input tokens without solving more. Elsewhere: tools run free-local to $100+/month, budget tiers claim 3-50x cheaper tokens, harness choice alone moved token use 4x, and daily tracking of 425 models finds a coding-agent-shaped call costs ~33x a classification-shaped one at the same published rates. No source finds platform identity to be the dominant term; several find the harness is.
All Sources
We aggregate published benchmarks; we never run our own tests and never pick winners. Each row links back to the original publication.
| Source | Date | Finding | Methodology | Quality |
|---|---|---|---|---|
| GitHub | 2026-01-01 | 80+ coding agents surveyed: free local tools to $100+/month for Claude Code; Claude Code ranks highest by adoption; pricing varies widely by workflow and model tier | Survey of 80+ agents with self-reported or public pricing data; SWE-bench scores where available | medium |
| Hacker News | 2026-04-24 | DeepSeek V4-Pro claims open-source SOTA on agentic coding; V4-Flash at $0.14/$0.28/M tokens is 3-50x cheaper than Claude tiers | Self-reported SOTA ranking; official pricing data from announcement | medium |
| capocasa.dev | 2026-07-22 | Same open-weights model, two coding harnesses: 4x difference in token spend on a 10-task SWE-bench subset; leaner one solved one more task | 10 SWE-bench tasks, both agents on GLM 5.2, tokens counted via proxy | medium |
| costpertoken.dev | 2026-09-06 | Tracking 425 models across 58 providers daily: a coding-agent-shaped call (large context in, code out) costs roughly 33x more per call than a bulk-classification-shaped call, and the multiple held across GPT-5.6, Claude Sonnet 5 and Gemini 3.7 Flash | Daily price snapshots across 425 models and 58 providers; two synthetic call shapes priced against each model's published rates | medium |
| aistack | 2026-09-10 | 3 harnesses x 2 models over 64 SWE-Bench Pro tasks: resolve rate is essentially flat at 44-53% across all three harnesses, while cost for the same sweep on the same model varies nearly 2x ($22.80 vs $45 on GLM-5.3-Flash). Harness choice barely moved accuracy and roughly doubled the bill. | 64 real SWE-Bench Pro tasks run through Claude Code, Codex and Pi against Qwen3.8-27B (FP8, 1xH200) and GLM-5.3-Flash (FP8, 4xH200). Each harness worker is a separate developer with its own environment, resources and web access; default harness settings and identical inference-engine configuration. Cost measured as wall time x infrastructure. | high |
How we work
OpenClawDatabase aggregates and links to published benchmarks. We don't run our own tests, and we don't pick winners. Our weekly benchmark-aggregator routine scans 7+ live leaderboards (OpenRouter, Aider, SWE-bench, GAIA, LMSYS, BigCodeBench, MMLU-Pro) plus relevant Reddit and Hacker News threads, then writes structured entries into /assets/benchmarks.json. Every row here links back to the original publication.
← Back to all benchmark tasks · See also: Decision guide · Cost calculator