DogLM

Can you pet the dog in an AI-generated game?

Last verified:

Visit DogLM

What is DogLM?

DogLM is an open-source benchmark framework that measures whether large language models spontaneously generate interactive dog-petting mechanics in video games when prompted to create games with dog characters—testing if models apply real-world game design principles without explicit instruction. It provides a leaderboard comparing 17+ AI models' capabilities in this specific task, with detailed scoring methodology and reproducible results.

DogLM pricing

Pricing model: Freemium

Open-source; costs depend on which model APIs you use to generate games (typically $0.02–$0.30 per game)

DogLM pros

  • Comprehensive leaderboard comparing 17+ AI models with detailed scoring breakdowns
  • Open-source and reproducible methodology available on GitHub
  • Tests spontaneous LLM decision-making beyond explicit instructions
  • Detailed interaction type analysis showing different model approaches (petting vs. proximity vs. animation)

DogLM cons

  • Extremely narrow benchmark scope (dog-petting in games) with limited generalizability to other LLM capabilities
  • Subjective scoring heavily dependent on game output quality and prompt engineering specifics
  • Resource-intensive to run across multiple models; costs vary significantly by API ($0.005–$0.32 per game)
  • Primarily research-focused; not practical for general users or most real-world applications

Frequently asked questions about DogLM

What does DogLM actually measure?

Whether LLMs spontaneously include dog-petting interactions in AI-generated games when a dog character is present but no explicit instruction covers player-dog interaction.

Where do I access DogLM?

The benchmark framework, games, and detailed methodology are available on GitHub at github.com/mikeushakov/doglm

How is each game scored?

2 points for pettable dogs, 1 point for interactive but non-pettable dogs, 0 for no interaction; failed/unparseable games are excluded. Scores are averaged across 5 runs per model.

Which model scored highest?

Gemini 3.7 Flash with a mean score of 8.2/20, followed by Kimi K3 (5.4) and Claude Opus 5 (5.2)

Categories

Use cases

Browse all AI tools on NeedAnAI