DogLM
Can you pet the dog in an AI-generated game?
Last verified:
What is DogLM?
DogLM is an open-source benchmark framework that measures whether large language models spontaneously generate interactive dog-petting mechanics in video games when prompted to create games with dog characters—testing if models apply real-world game design principles without explicit instruction. It provides a leaderboard comparing 17+ AI models' capabilities in this specific task, with detailed scoring methodology and reproducible results.
DogLM pricing
Pricing model: Freemium
Open-source; costs depend on which model APIs you use to generate games (typically $0.02–$0.30 per game)
DogLM pros
- Comprehensive leaderboard comparing 17+ AI models with detailed scoring breakdowns
- Open-source and reproducible methodology available on GitHub
- Tests spontaneous LLM decision-making beyond explicit instructions
- Detailed interaction type analysis showing different model approaches (petting vs. proximity vs. animation)
DogLM cons
- Extremely narrow benchmark scope (dog-petting in games) with limited generalizability to other LLM capabilities
- Subjective scoring heavily dependent on game output quality and prompt engineering specifics
- Resource-intensive to run across multiple models; costs vary significantly by API ($0.005–$0.32 per game)
- Primarily research-focused; not practical for general users or most real-world applications
Frequently asked questions about DogLM
What does DogLM actually measure?
Whether LLMs spontaneously include dog-petting interactions in AI-generated games when a dog character is present but no explicit instruction covers player-dog interaction.
Where do I access DogLM?
The benchmark framework, games, and detailed methodology are available on GitHub at github.com/mikeushakov/doglm
How is each game scored?
2 points for pettable dogs, 1 point for interactive but non-pettable dogs, 0 for no interaction; failed/unparseable games are excluded. Scores are averaged across 5 runs per model.
Which model scored highest?
Gemini 3.7 Flash with a mean score of 8.2/20, followed by Kimi K3 (5.4) and Claude Opus 5 (5.2)