Boundary
What is supplied by the game or workflow, and what is assigned to AI?
THE FOUNDATION MODEL ERA
The next move is
more than playing.
A visual companion to the survey of how foundation models play, model, design, build, adapt, and test games.
Final paper · 120 pages · Concepts, evidence, and an open reading collection
VIDEO INTRODUCTION
A concise visual introduction to the six research roles, their connections, and the evidence needed to assess progress across the game lifecycle.
01:30 · Full presentation · Sound availableGame AI has never been limited to playing. The foundation-model era makes the field’s wider shape impossible to miss.
AI can now act as a player or character, model a world or its players, propose content and rules, operate inside a game engine, change a live experience, and produce testing evidence. The same pretrained model family may touch several of these tasks, but each role still depends on different interfaces, constraints, and forms of evidence.
The survey therefore follows the immediate use of AI output. This separates a proposed mechanic from its implementation, a player forecast from the adaptation it informs, and a tester’s action from the verdict it supports.
See the organizing map02 THE ORGANIZING MAP
The roles classify what an output is used for. The connectors show selective inputs and feedback—not a required architecture.

Six uses of AI output around an executable, player-facing game.
These questions turn the foundation-model era from a date label into a common analytical lens.
What is supplied by the game or workflow, and what is assigned to AI?
Which capabilities transfer, which artifacts can be reused, and what remains setting-specific?
What claims does evaluation support where the output is used?
The timeline traces selected lineages; the two-part knowledge map locates the technical directions and cited systems discussed in the survey.
03 THE ROLE ATLAS
Each chapter follows a different output, operating context, and empirical claim. Select a role to see its research structure and synthesis.

01 / PLAY AND ACT
AI selects actions, plans, or messages inside a game. Foundation models broaden perception and planning, while the path to native controls remains game-specific.
04 CROSS-ROLE DISCUSSION
An implemented exchange shows that information can move between roles. It does not by itself show that competence transfers or that the receiving workflow improves.
Language, multimodal, and code interfaces connect high-level intent to constrained plans, tools, actions, or learned dynamics. Reliability still depends on what the receiving game or component must execute.
A trajectory, specification, profile, or test trace can be consumed downstream without proving competence under a new game, engine, interface, player population, or task. Transfer requires naming what changed and testing whether the capability survives.
Compatibility, actual use, and benefit are separate checks. A connection earns a stronger claim only when the intended downstream outcome improves against a suitable baseline, with cost and failure propagation made visible.
One route separates executable game state from generated appearance: a coding agent builds a graybox, the engine executes rules and updates state, and a state-conditioned model renders what players see. Programmable World Model ↗ demonstrates a related loop with a lightweight engine; AlayaRenderer-Flash ↗ brings neural rendering to a live game engine. The full UE/Unity development-to-rendering workflow is a research direction, not an already validated end-to-end system.
05 EVIDENCE & RESEARCH DIRECTIONS
The six roles do not share a single measure of success. Each card pairs the main evaluation target with the boundary that current research still needs to cross.
Measure completion, progress, efficiency, and coordination.
Next test vary appearance, controls, rules, and partner conventions under fixed interaction and real-time budgets.
Measure visual, mechanics, and persistent-state fidelity; reference-game policy return; and held-out player behavior.
Next test revisit changed worlds and distinguish stable player tendencies from temporary context.
Measure validity, diversity, constraints, and designer control.
Next test follow real revision work, retained alternatives, correction effort, and the resulting player experience.
Measure launchability, requirements, repair, and regression.
Next test use multi-version projects and evaluate whether a new developer can reproduce, revise, and hand off the work.
Measure latency, consistency, controllability, and player outcomes.
Next test run repeated sessions with non-adaptive controls, resumed saves, model updates, and reversible interventions.
Measure validated defects, coverage, verdict accuracy, and reproduction cost.
Next test combine diverse exploration with independent oracles, held-out defect families, and relevant player samples.
Perceptual quality asks whether generated video looks and flows plausibly. Mechanics correctness checks whether actions produce rule-consistent changes. Persistent state asks whether those facts survive revisits and long interaction. FVD measures video-feature distributions; when reference game or engine state is available, health, collisions, inventory, and event traces can be checked separately.
Changing a map, a mode, or a game tests different things. A new scene may probe visual transfer; an altered control scheme probes action transfer; unfamiliar rules require reasoning that neither demonstrates. Cross-game claims should name the held-out unit and report the behavior that survived it.
PLAY NINE AI-CRAFTED WORLDS
Nine browser games span physics puzzles, strategy, racing, platforming, and compact experimental worlds. Start with three new multi-stage adventures, then explore six earlier prototypes.
WORLD 01
Launch three companions with distinct mid-air abilities, expose structural weak points, and topple six increasingly elaborate sky fortresses.
WORLD 02
Grow energy, combine six plant roles, upgrade defenses, and hold five lanes against increasingly coordinated mechanical swarms.
WORLD 03
Race six karts across three circuits, charge nitro through drifts, collect items, and turn every corner into an overtaking opportunity.
WORLD 04
Push planters onto moonlit goals, hold pressure plates, open the bridge, and undo mistakes while solving three compact spatial puzzles.
WORLD 05
Read the next storm footprint, manage scarce energy, activate every beacon, and return safely using dash, shield, and wait.
WORLD 06
Guide a winged dragon across moonlit platforms, collect starlight, and use two model-composed abilities to reach the portal.
WORLD 07
Search a snowy archipelago, bring stranded companions aboard, and return them safely to the lighthouse in limited-capacity trips.
WORLD 08
Dodge persistent pursuers, gather starlight, and combine a repelling pulse with a short shield to protect the garden.
WORLD 09
Traverse a revised route of floating islands, avoid crimson traps, collect starlight, and reach the final portal.
06 PROJECTS, IN VIEW
Six roles, seen through public project imagery. Each example links to its original source and is described at the level its evidence supports.

From specialist policies to language-guided agents, this literature studies how AI perceives a game, chooses actions, and coordinates with other players.
GT Sophy brings learned racing behavior into Gran Turismo 7, connecting research in real-time control with a game people can play.
Illustrations from the cited projects; rights remain with their respective owners.
← → to navigate collections07 THE READING ROOM
Search 436 current manuscript references and eight separately tracked public additions by topic, year, or keyword. Role filters distinguish primary research roles from supporting foundations and context.
444 references in the bibliography
Try another title or topic, or reset your filters.
08 AN OPEN COLLECTION
Know a paper that belongs here? Help make the collection more useful to researchers, designers, and developers.