THE FOUNDATION MODEL ERA

AI for Gamesin the Foundation
Model Era

The next move is
more than playing.

A visual companion to the survey of how foundation models play, model, design, build, adapt, and test games.

Final paper · 120 pages · Concepts, evidence, and an open reading collection

01 / A WORLD IN THE MAKINGConcept artwork
444references
06roles across the lifecycle
03questions across every role
Watch the overview

Authors & affiliations

Meng Luo1Yanlin Li1Hao Li1Hongzhan Lin1Pengfei Zhou1Tianjie Ju1Ran Zhang2Yeying Jin1Mong-Li Lee1Wynne Hsu1

1 National University of Singapore2 Nanyang Technological University

The paper,
in 90 seconds.

A concise visual introduction to the six research roles, their connections, and the evidence needed to assess progress across the game lifecycle.

01:30 · Full presentation · Sound available
AI for Games in the Foundation Model EraPaper overview film

The game is no longer
only given.

Game AI has never been limited to playing. The foundation-model era makes the field’s wider shape impossible to miss.

AI can now act as a player or character, model a world or its players, propose content and rules, operate inside a game engine, change a live experience, and produce testing evidence. The same pretrained model family may touch several of these tasks, but each role still depends on different interfaces, constraints, and forms of evidence.

The survey therefore follows the immediate use of AI output. This separates a proposed mechanic from its implementation, a player forecast from the adaptation it informs, and a tester’s action from the verdict it supports.

See the organizing map

Six roles around
a playable game.

The roles classify what an output is used for. The connectors show selective inputs and feedback—not a required architecture.

Diagram organizing AI for Games into six roles: design, build and maintain, model players and games, test and evaluate, generate and adapt at runtime, and play and act
Figure 2 · Survey overview

Six uses of AI output around an executable, player-facing game.

THREE QUESTIONS, REPEATED ACROSS EVERY ROLE

These questions turn the foundation-model era from a date label into a common analytical lens.

01

Boundary

What is supplied by the game or workflow, and what is assigned to AI?

02

Transfer & reuse

Which capabilities transfer, which artifacts can be reused, and what remains setting-specific?

03

Evidence

What claims does evaluation support where the output is used?

FROM THE MANUSCRIPT

The timeline traces selected lineages; the two-part knowledge map locates the technical directions and cited systems discussed in the survey.

Inside each
research role.

Each chapter follows a different output, operating context, and empirical claim. Select a role to see its research structure and synthesis.

Research directions for AI that plays and acts
Research directions for AI that plays and acts: specialist-to-generalist policies, test-time adaptation, control hierarchies, and NPCs and teammates.

01 / PLAY AND ACT

AI that Plays and Acts

AI selects actions, plans, or messages inside a game. Foundation models broaden perception and planning, while the path to native controls remains game-specific.

Output
Actions · plans · messages
Applications
Players · teammates · NPCs
Main claim
Action quality
CHAPTER STRUCTURE
  • Player and generalist agents
  • Learning and control hierarchies
  • Test-time adaptation and memory
  • Opponents, teammates, NPCs, and companions
SURVEY SYNTHESIS
  • A common interface relocates the learning problem.
  • Adaptation includes the information supplied.
  • Playing well with others is a separate transfer problem.
Representative anchorsVoyager ↗SIMA 2 ↗NitroGen ↗

What connects—and
what actually transfers.

An implemented exchange shows that information can move between roles. It does not by itself show that competence transfers or that the receiving workflow improves.

01

Broader interfaces, game-specific structure

Language, multimodal, and code interfaces connect high-level intent to constrained plans, tools, actions, or learned dynamics. Reliability still depends on what the receiving game or component must execute.

02

Artifact reuse is not capability transfer

A trajectory, specification, profile, or test trace can be consumed downstream without proving competence under a new game, engine, interface, player population, or task. Transfer requires naming what changed and testing whether the capability survives.

03

Downstream benefit needs its own test

Compatibility, actual use, and benefit are separate checks. A connection earns a stronger claim only when the intended downstream outcome improves against a suitable baseline, with cost and failure propagation made visible.

Interaction tracesLearned environmentsSpecifications & rulesExecution & test evidencePlayer models & profiles
AN EMERGING CONNECTION · DESIGN / BUILD / MODEL / RUNTIME

One route separates executable game state from generated appearance: a coding agent builds a graybox, the engine executes rules and updates state, and a state-conditioned model renders what players see. Programmable World Model ↗ demonstrates a related loop with a lightweight engine; AlayaRenderer-Flash ↗ brings neural rendering to a live game engine. The full UE/Unity development-to-rendering workflow is a research direction, not an already validated end-to-end system.

The claim decides
the test.

The six roles do not share a single measure of success. Each card pairs the main evaluation target with the boundary that current research still needs to cross.

01

Play & Act

Measure completion, progress, efficiency, and coordination.

Next test vary appearance, controls, rules, and partner conventions under fixed interaction and real-time budgets.

02

Model Players & Games

Measure visual, mechanics, and persistent-state fidelity; reference-game policy return; and held-out player behavior.

Next test revisit changed worlds and distinguish stable player tendencies from temporary context.

03

Design

Measure validity, diversity, constraints, and designer control.

Next test follow real revision work, retained alternatives, correction effort, and the resulting player experience.

04

Build & Maintain

Measure launchability, requirements, repair, and regression.

Next test use multi-version projects and evaluate whether a new developer can reproduce, revise, and hand off the work.

05

Generate & Adapt at Runtime

Measure latency, consistency, controllability, and player outcomes.

Next test run repeated sessions with non-adaptive controls, resumed saves, model updates, and reversible interventions.

06

Test & Evaluate

Measure validated defects, coverage, verdict accuracy, and reproduction cost.

Next test combine diverse exploration with independent oracles, held-out defect families, and relevant player samples.

WORLD-MODEL FIDELITY

Three tests, not one score.

Perceptual quality asks whether generated video looks and flows plausibly. Mechanics correctness checks whether actions produce rule-consistent changes. Persistent state asks whether those facts survive revisits and long interaction. FVD measures video-feature distributions; when reference game or engine state is available, health, collisions, inventory, and event traces can be checked separately.

WHAT TRANSFERS?

A new map is not a new game.

Changing a map, a mode, or a game tests different things. A new scene may probe visual transfer; an altered control scheme probes action transfer; unfamiliar rules require reasoning that neither demonstrates. Cross-game claims should name the held-out unit and report the behavior that survived it.

Don’t just read it.
Play it.

Nine browser games span physics puzzles, strategy, racing, platforming, and compact experimental worlds. Start with three new multi-stage adventures, then explore six earlier prototypes.

Cloud Sling physics puzzle with three companions facing a wooden and glass robot fortress WORLD 01
Physics puzzle6 fortresses

Cloud Sling 云岛弹射

Launch three companions with distinct mid-air abilities, expose structural weak points, and topple six increasingly elaborate sky fortresses.

Drag · aimRelease · launchSpace · ability
DesignBuildPlay
Play now
Sprout Guard five-lane garden defense with flowers, cacti, stones, and approaching robot bugs WORLD 02
Lane defense18 waves

Sprout Guard 芽芽保卫战

Grow energy, combine six plant roles, upgrade defenses, and hold five lanes against increasingly coordinated mechanical swarms.

Click · plant1–6 · selectSpace · collect
DesignBuildPlay
Play now
Neon Kart 3D arcade race with six karts, a coastal circuit, item boxes, and a minimap WORLD 03
3D arcade racing3 tracks

Neon Kart 霓虹卡丁

Race six karts across three circuits, charge nitro through drifts, collect items, and turn every corner into an overtaking opportunity.

WASD · driveShift · driftSpace / E · boost & item
DesignBuildPlay
Play now
Moonlight Garden spatial puzzle with a robot, planters, moonlit goals, and a bridge WORLD 04
Spatial puzzle3 levels

Moonlight Garden 月光花园

Push planters onto moonlit goals, hold pressure plates, open the bridge, and undo mistakes while solving three compact spatial puzzles.

WASD / arrowsZ · undoR · restart
DesignBuildTest
Play now
Aurora Outpost turn-based strategy board with a robot, beacons, energy, and storm cells WORLD 05
Turn-based strategy3 levels

Aurora Outpost 极光哨站

Read the next storm footprint, manage scarce energy, activate every beacon, and return safely using dash, shield, and wait.

WASD / arrowsShift · dashE · shieldSpace · wait
DesignBuildTest
Play now
A small purple dragon crossing moonlit snow platforms beneath a starry sky WORLD 06
Platform adventureAI-created

Lulu's Snowy Letter 露露的雪夜来信

Guide a winged dragon across moonlit platforms, collect starlight, and use two model-composed abilities to reach the portal.

A / D · moveW · jumpQ / E · abilities
DesignBuildPlay
Play now
A small rescue boat moving through a snowy archipelago toward a lighthouse WORLD 07
Ocean rescueAI-revised

The Lighthouse Is Still Shining 灯塔还亮着

Search a snowy archipelago, bring stranded companions aboard, and return them safely to the lighthouse in limited-capacity trips.

WASD · steerShift · dashQ / E · abilities
DesignBuildPlay
Play now
A snowy garden survival arena with a small fox, pursuers, and collectible starlight WORLD 08
Survival arenaAI-revised

Garden Watch 花园守望

Dodge persistent pursuers, gather starlight, and combine a repelling pulse with a short shield to protect the garden.

WASD · moveShift · dashQ / E · abilities
DesignBuildPlay
Play now
A tiny adventurer crossing bright floating islands toward a portal in the sky WORLD 09
Platform journeyAI-revised

A Letter from the Clouds 云端来信

Traverse a revised route of floating islands, avoid crimson traps, collect starlight, and reach the final portal.

A / D · moveW · jumpQ / E · abilities
DesignBuildPlay
Play now
Playable in the browserNo installation

Ideas made
concrete.

Six roles, seen through public project imagery. Each example links to its original source and is described at the level its evidence supports.

Gran Turismo 7 racing scene from an official GT Sophy announcement
Gran Turismo 7 / GT SophyOriginal project
01 / PLAY AND ACT116 works

Learning to act.
Learning to cooperate.

From specialist policies to language-guided agents, this literature studies how AI perceives a game, chooses actions, and coordinates with other players.

Generalist policiesPlanning & memoryHuman–AI teamwork
IN THE IMAGE

GT Sophy brings learned racing behavior into Gran Turismo 7, connecting research in real-time control with a game people can play.

Explore this role

Illustrations from the cited projects; rights remain with their respective owners.

← → to navigate collections

Follow your curiosity.

Search 436 current manuscript references and eight separately tracked public additions by topic, year, or keyword. Role filters distinguish primary research roles from supporting foundations and context.

444 references in the bibliography

Download BibTeX ↓Newest first

The next discovery
might be yours.

Know a paper that belongs here? Help make the collection more useful to researchers, designers, and developers.

Suggest a paper