THE GAUNTLET
A mechanical bake-off ledger for local LLMs on one RTX 3090 — text, vision, code, photo, and hallucination arenas graded by sealed criteria, not vibes.
Last verified:
What is THE GAUNTLET?
The Gauntlet is a competitive benchmarking platform that evaluates and ranks large language models (LLMs) across three practical arenas: newscast writing (TEXT), screen reading and data extraction (VISION), and real bug-fixing from production code (CODE). Models compete directly with transparent, sealed criteria and results are publicly verified with video evidence.
THE GAUNTLET pricing
Pricing model: Freemium
THE GAUNTLET pros
- Multiple real-world evaluation arenas (TEXT, VISION, CODE) with practical task simulation
- Transparent, reproducible methodology using sealed criteria and production pipeline validation
- Public results with video evidence and detailed metrics (speed, accuracy, validity scores)
- Tests against actual production bugs and regression tests—no synthetic tasks
THE GAUNTLET cons
- Limited model coverage—only specific LLM versions tested, doesn't benchmark all available models
- Infrequent result updates shown in ledger (testing appears ongoing but not continuous)
- No information on how to submit or add new models to evaluation
- VISION arena results incomplete—CODE arena results table appears truncated in content
Frequently asked questions about THE GAUNTLET
How are winners determined?
In TEXT: model must sweep both speed and average hedges (weasel wording), not just tie. In VISION: judged on digit recall and avoiding invented numbers. In CODE: test suite is the only judge—fix the regression test and keep full suite green to win.
What are the evaluation criteria?
Criteria are sealed before each run to prevent gaming. They're versioned (v1, v2) and the same production pipeline used for the station's nightly shows grades all responses.
Can I see evidence of the results?
Yes, most recent runs include links to YouTube videos showing the evaluation in progress with real-time results.