Rogue-Bench
LLMs play the game Rogue
Last verified:
What is Rogue-Bench?
Rogue-Bench is a benchmark where AI agents and large language models play the classic roguelike game Rogue to evaluate their reasoning, planning, and sequential decision-making capabilities. The benchmark uses the original Rogue game as a testbed for measuring how well LLMs can navigate procedural dungeons, make strategic choices, and handle permanent death mechanics in a turn-based dungeon crawler. It is designed for AI researchers, machine learning engineers, and anyone working on LLM agent evaluation who needs a challenging environment for testing reasoning abilities.
The benchmark is open-source and Python-based, hosted on GitHub by Ian Whalen. It builds on rogomatic-llm, a proof-of-concept project that demonstrates LLMs playing Rogue using a modified version of the Rogue Collection. The benchmark provides objective evaluation metrics for agent performance in the game environment, allowing researchers to compare different LLM architectures and approaches.
Key features include the use of the classic Rogue game with its procedural generation and permadeath mechanics, integration with LLM agents, and quantitative performance metrics. The benchmark is particularly useful for testing long-horizon planning, resource management, and adaptive strategy development in AI agents.
The tool is free and open-source, available on GitHub for anyone to use and contribute to. It falls under the AI benchmarking category and is part of growing efforts to evaluate LLM capabilities through gameplay rather than traditional benchmark datasets.
Pros include: open-source and free, uses classic Rogue game, evaluates reasoning and planning, Python-based and accessible, objective performance metrics, suitable for LLM agent research, builds on rogomatic-llm proof of concept, allows comparison of different LLMs, tests sequential decision-making, includes procedural dungeon generation, tests permanent death mechanics, turn-based gameplay allows deliberation, available on GitHub for contribution, part of emerging agentic benchmarking field, evaluates long-horizon planning
Cons include: limited to single game (Rogue), requires understanding of Rogue game mechanics, may have limited model support, niche application for roguelike enthusiasts only, evaluation may not generalize to other tasks, requires computational resources for LLM inference, may need technical setup knowledge, limited documentation potentially, focuses on turn-based only not real-time, permadeath may limit extensive testing, procedural generation adds variability complications, primarily for research not production, limited to text-based interface, May require emulated environment setup, narrow scope compared to broader benchmarks
Pricing: Free and open-source, available on GitHub with no paid tiers
FAQs: The benchmark is specifically for evaluating LLMs through gameplay rather than traditional benchmarks, it uses the original Rogue game which adds challenge through procedural generation and permadeath, it's designed for researchers working on AI agent evaluation, it's open-source allowing contributions, it requires Python and technical setup, it tests sequential decision-making capabilities, it allows comparison across different LLM models, it's part of emerging agentic benchmarking approaches, it may not be suitable for production applications, and it focuses on reasoning rather than creative tasks
Rogue-Bench pricing
Pricing model: Freemium
Free and open-source, available on GitHub with no paid tiers or subscription plans
Rogue-Bench pros
- Open-source and completely free
- Uses classic Rogue game
- Evaluates reasoning and planning
- Python-based and accessible
- Objective performance metrics
- Suitable for LLM agent research
- Builds on rogomatic-llm proof of concept
- Allows comparison of different LLMs
- Tests sequential decision-making
- Includes procedural dungeon generation
- Tests permanent death mechanics
- Turn-based gameplay allows deliberation
- Available on GitHub for contribution
- Part of emerging agentic benchmarking field
- Evaluates long-horizon planning
Rogue-Bench cons
- Limited to single game Rogue
- Requires understanding Rogue mechanics
- May have limited model support
- Niche application for roguelike fans
- Evaluation may not generalize elsewhere
- Requires computational resources
- May need technical setup knowledge
- Limited documentation potentially
Frequently asked questions about Rogue-Bench
What is Rogue-Bench?
Rogue-Bench is a benchmark where AI agents and large language models play the classic roguelike game Rogue to evaluate their reasoning, planning, and sequential decision-making capabilities.
Who is Rogue-Bench for?
It is designed for AI researchers, machine learning engineers, and anyone working on LLM agent evaluation who needs a challenging environment for testing reasoning abilities.
Is Rogue-Bench free?
Yes, Rogue-Bench is completely free and open-source, available on GitHub for anyone to use and contribute to.
What programming language is it built with?
Rogue-Bench is Python-based, making it accessible to most machine learning practitioners and researchers.
How does it evaluate LLMs?
It evaluates LLMs by having them play Rogue, measuring their ability to navigate procedural dungeons, make strategic choices, and handle permanent death mechanics in a turn-based environment.
What makes Rogue challenging for LLMs?
Rogue features procedural dungeon generation and permanent death mechanics, requiring long-horizon planning, resource management, and adaptive strategy development.
Can I contribute to Rogue-Bench?
Yes, the benchmark is open-source on GitHub and welcomes contributions from the community.
What models can be tested with Rogue-Bench?
The benchmark can test various LLMs and AI agents, allowing researchers to compare different architectures and approaches.
Does it provide objective metrics?
Yes, Rogue-Bench provides objective performance metrics for agent performance in the game environment.
How does it differ from traditional benchmarks?
Unlike traditional benchmarks, Rogue-Bench evaluates LLM capabilities through gameplay rather than static datasets, testing sequential decision-making in a dynamic environment.