Pangram verdict · v3.3
We believe that this entire text is human-written.
AI likelihood · overall
HumanArticle text · 156 words · 1 segments analyzed
Pass rateMedian cost per successful taskMedian cost per taskMedian cache hit rate per successful taskMedian time per successful taskBeyond the numbers01OpenCode: failures excluded.It only covers 15 passes. Count failed attempts and the number becomes $3.24 per task.02Cache hit rate is not cost.A cached 300-turn failure can still burn more than a short cache miss.03Quality and cost can diverge.Claude Code passes 19 tasks, but reaches $18.34 in cost per task.Run your harness on Runta.If you want to test your own harness on Runta, we’ll give you $100 in credits to get started.Get $100 in creditsStart free trialTested harnessesCodexv0.148.0DeepSeek Harnessv0.1.0-rc.8Claude Codev2.1.237Piv0.84.2Oh My Piv17.4.0Kimi Codev0.37.2Exo Harnessv0.1.0OpenCodev1.18.19Hermesv0.20.4FrontierHarness v1.0 focuses on software engineering contexts and terminal-based tasks. It may not generalize to other areas of knowledge work.Evaluated on Runta agent runtimes. All harnesses and the task environment are prepared once as a golden checkpoint. Every run is a fresh restore with identical vCPU, memory, disk size, disk contents, and memory state.