DeepSWE leaderboard: DeepSWE score against average cost per task for 113 tasks, with the cost-performance frontier highlighted and GPT-6-Astra XHIGH at $6.52 per task

The marketing machine is running at full speed. Every new model is the next breakthrough, and the industry debates which architecture holds the highest raw intelligence. I spent time inside the DeepSWE dataset, and the picture that emerges is economic: GPT-6-Astra XHIGH sits alone on the cost-performance Pareto front, GLM-5.3-Flash is the sensible production default, and the models separate more sharply on cost per solved task than on peak scores.

GPT-6-Astra XHIGH sits alone on the Pareto front

On DeepSWE, GPT-6-Astra XHIGH is the only configuration that truly sits on the Pareto front of cost versus performance. For every other configuration in the mix, you either pay a premium for slightly less output, or accept a performance hit to save a few cents. In autonomous agent work, where every token and every second is billed, that distinction matters more than the theoretical debate about which model is smartest.

The chart above shows the shape of it: 113 tasks, DeepSWE score on the vertical axis, average cost per task on the horizontal axis. GPT-6-Astra XHIGH sits at $6.52 per task near the top of the frontier.

The metric that matters: cost per correctly solved task

The real metric is not peak intelligence. It is how much a task, correctly solved, actually costs. A brilliant model that delivers results at ten times the price of its predecessor is not an industrial revolution; it is a luxury play for deep-pocketed enterprises. Production systems need models that deliver value, not spectacle.

My current pick: GLM-5.3-Flash

My personal top pick remains GLM-5.3-Flash. It is not the most spectacular model on the charts, and it does not dominate every benchmark. It offers the best compromise I can find right now between speed, accuracy, and reasonable inference cost. If you are building production systems, this is the model I would reach for first.

The anomaly: a previous-generation model beats a flagship tier

The data also contains a result that defies simple expectations: GPT-5.6-Luna-Max is more efficient than GPT-6-Astra-Low. An older, heavier configuration delivers better value than the entry tier of the newest flagship. Model scaling is not a linear path, and it may be time to revisit assumptions about which architecture fits which workload.

Follow-up: DeepSeek V4.1 Flash scores 98% of Astra at 1.4% of the cost

Three days after I published this analysis, the team behind the OpenDesign Arena benchmarked DeepSeek V4.1 Flash on everyday design tasks built from user requests. The result reads like this post’s thesis on another dataset: the open model reached 98% of GPT-6 Astra’s score at 1.4% of the cost. Per artifact, that is $0.023 and 5.3 minutes for DeepSeek V4.1 Flash, against $1.61 and 11.1 minutes for Astra, while Claude Fable 5.1 trailed at $3.66 and 12.8 minutes. Every model in the comparison except Astra scored lower and cost more. Full results are in their thread on X.

We benchmarked DeepSeek V4.1 Flash … It reached 98% of GPT-6 Astra’s score at 1.4% of the cost on everyday design tasks based on user requests. Every model except Astra scored lower and cost more. Are open models overtaking closed ones?

OpenDesign (@OpenDesignHQ), on X
OpenDesign Arena chart: DeepSeek V4.1 Flash scores 81.2 of 100 at $0.023 per artifact, just below GPT-6 Astra at 82.7 and $1.61, while every other model scored lower and cost more
Quality score against cost per artifact across 13 models: DeepSeek V4.1 Flash ranks second on score and first on cost. Chart: OpenDesign.

Two cautions before rewriting a shortlist. This is a design benchmark, not DeepSWE: different tasks, different methodology, no direct comparison between the two leaderboards. And the aggregate hides the detail: DeepSeek V4.1 Flash places ahead of Astra on websites and landing pages but tenth on dashboards and admin panels. The discipline stays the same: read the score measured on a workload that resembles yours, then divide it by the cost.

FAQ

What is DeepSWE?

DeepSWE is a benchmark for autonomous coding agents. The leaderboard behind this article scores model configurations across 113 tasks and reports each success rate next to its average cost per task, so every configuration can be read on both axes at once.

Which configuration offers the best cost-performance on DeepSWE?

On the leaderboard I used, GPT-6-Astra XHIGH is the only configuration that truly sits on the Pareto front of cost versus performance. Every other configuration makes you choose: pay a premium for slightly less output, or accept a performance hit to save budget.

Which model should I pick for production work?

My pick is GLM-5.3-Flash. It does not top every chart, but it offers the best compromise between speed, accuracy, and inference cost for production systems. Treat any leaderboard as a starting point, then validate against your own workload before you commit.

Does DeepSeek V4.1 Flash change the DeepSWE conclusion?

No. The DeepSeek V4.1 Flash figures come from OpenDesign’s design benchmark, not from DeepSWE, so the two leaderboards are not directly comparable. It is the same economic pattern on a different workload: an open model at 98% of the leader’s score for 1.4% of its cost per task.

Run the numbers yourself

The dataset behind this analysis is available in an interactive explorer, so you can run your own experiments and form your own conclusions: deepswe.datacurve.ai.

One caution before you generalize from any leaderboard: a benchmark rewards what it measures. I opened the Benchmark Best Practices series with Amdahl’s law and the scaling claims it caps.

So, what is driving your current stack: pure intelligence, or the bottom line? The gap between marketing hype and engineering reality keeps widening, and the only way to bridge it is to look at the hard data.

I posted the headline chart on LinkedIn as well, and the discussion is open there: join the thread.

Related posts

And if you want the same cost-versus-performance discipline applied to your own data platform, that is what my data platform audits are for.


0 Comments

Leave a Reply

Avatar placeholder

Your email address will not be published. Required fields are marked *