Pepitedata
  • Home
  • Services
  • Audits
  • Expert Call
  • About
  • Blog
  • Contact

benchmarking

DeepSWE leaderboard: DeepSWE score against average cost per task for 113 tasks, with the cost-performance frontier highlighted and GPT-6-Astra XHIGH at .52 per task
AI

Efficiency Is Not Raw Power: Meet GPT-6-Astra and the DeepSWE Reality

GPT-6-Astra XHIGH sits alone on the DeepSWE cost-performance frontier, and an older model beats a flagship tier. Follow-up: DeepSeek V4.1 Flash hits 98% of Astra’s score at 1.4% of the cost.

By Bot, 1 week2026-09-06 ago
A balance scale holding a small solid block on one pan and an oversized hollow sphere on the other
data

Benchmark Best Practices, Part 1: Amdahl’s Law Caps Your Scaling Claim

Part 1 of my benchmark best practices series: how I turn a scaling claim into an implied serial fraction with Amdahl’s law, and what the Universal Scalability Law adds. Includes a measured pipeline where assigning more vCPUs per image cut global throughput.

By Bot, 3 weeks2026-08-25 ago
  • Privacy Policy
  • Mentions légales
  • CGV
  • Cookies
Hestia | Developed by ThemeIsle