Agentic Benchmarks

Swarm and single-model benchmarks, scored by executing the built application.

Gauntlet leaderboard

Community fleets and cloud baselines of one Gauntlet version, ranked together. Each row opens the run's checks and evidence.

1gpt-6.1-sol-proSingle modelopenai/gpt-6.1-sol-pro0.95742qwen3.8-max-primeSingle modelqwen/qwen3.8-max-prime0.83643claude-sonnet-5.5Single modelanthropic/claude-sonnet-5.50.79844gpt-6.1-solSingle modelopenai/gpt-6.1-sol0.79785pareto-26.10-previewSingle modelunbiased/pareto-26.10-preview0.79576claude-opus-5.5Single modelanthropic/claude-opus-5.50.79267mimo-v2.6-flashSingle modelxiaomi/mimo-v2.6-flash0.79098jev-routerSingle modeltypesafe/jev-router0.69759grok-4.7Single modelx-ai/grok-4.70.693710gemini-3.8-flashSingle modelgoogle/gemini-3.8-flash0.692911ember-1Single modelfireworks/ember-10.686612qwen3.8-flashSingle modelqwen/qwen3.8-flash0.684613glm-5.3-primeSingle modelz-ai/glm-5.3-prime0.679714deepseek-pro-latestSingle model~deepseek/deepseek-pro-latest0.659115glm-5.3-flashxSingle modelz-ai/glm-5.3-flashx0.630316deepseek-v4.1-flashSingle modeldeepseek/deepseek-v4.1-flash0.591017muse-spark-1.3Single modelmeta/muse-spark-1.30.582618gpt-terra-latestSingle model~openai/gpt-terra-latest0.578119GPT-6 LunaSingle modelopenai/gpt-6-luna0.479720hy4-previewSingle modeltencent/hy4-preview0.429321qwen3.8-omni-flashSingle modelqwen/qwen3.8-omni-flash0.426922glm-5.3-flashSingle modelz-ai/glm-5.3-flash0.414323glm-5.3Single modelz-ai/glm-5.30.144324gpt-6-luna-proSingle modelopenai/gpt-6-luna-pro0.134325qwen3.8-27bSingle modelqwen/qwen3.8-27b0.091526ling-3.0-flashSingle modelinclusionai/ling-3.0-flash0.067027aion-3.5Single modelaion-labs/aion-3.50.046428aion-3.5-miniSingle modelaion-labs/aion-3.5-mini0.018729nex-n2.5-proSingle modelnex-agi/nex-n2.5-pro0.011330nemotron-3.5-lightningSingle modelnvidia/nemotron-3.5-lightning0.009531solar-mini4Single modelupstage/solar-mini40.0083
community fleetCloud baselineGauntlet 7.2 · 31 entries · 31 baselines

About the Gauntlet benchmark