Show HN: WorldBuild Bench repo: testing LLM world coherence with 3D games
WorldBuild Bench Same brief, same tools, same harness. Build a 3D game someone would actually want to keep playing. One brief: Sunset Apex, a 3-lap circuit racer.
- ▪WorldBuild Bench Same brief, same tools, same harness.
- ▪Build a 3D game someone would actually want to keep playing.
- ▪One brief: Sunset Apex, a 3-lap circuit racer.
Opening excerpt (first ~120 words) tap to expand
WorldBuild Bench Same brief, same tools, same harness. Build a 3D game someone would actually want to keep playing. One brief: Sunset Apex, a 3-lap circuit racer. Three models. Three completely different games. Real, unedited gameplay. All 27 builds from the round are playable in your browser right now. ▶ Play the builds · Vote in the Arena · Methodology · Round data Why this exists Coding benchmarks usually stop at "does it run." As models move toward "world models," a harder question is spatial, temporal, and causal coherence in a 3D space: does the model understand where things are, stay consistent over time, and when something happens, do the consequences make sense? Those qualities are hard to capture with static benchmark questions.
…
Excerpt limited to ~120 words for fair-use compliance. The full article is at GitHub.