From standing still to moving right.
v001 copies the input state. v002 moves the object marked YOU, matching the real next state for this action.
From interaction to executable world models
Online Self-Supervised Dynamics Discovery
for Executable World Models
1Korea University2Kyung Hee University
Same world. Unfamiliar words. Rules discovered through experience.
Baba Is You is a puzzle game whose rules are movable word blocks. BABA IS YOU means you control Baba; FLAG IS WIN makes touching a flag a win. Rearrange the words into ROCK IS YOU, and you control the rocks instead.
Explore the original Baba Is You on Steam ↗
Baba in Wonderland keeps the same mechanics but changes familiar property words: YOU becomes STRANGE. Without a dictionary, rule descriptions, or a learning reward, the agent has to discover what these words do by acting and observing the result.
Alice learns a world model written in Python: give it a state and an action, and it predicts the next state. The research measures how accurately this code predicts changes in the world. In the demo below, you or a saved solution choose the actions; the learned program predicts their outcomes.
BABAISYOU
BABAISSTRANGE
The label changes. The behavior stays the same.
Executable world models can be read, edited, executed, and reused for planning, but only if the program captures the environment's transition law rather than semantic shortcuts in its surface vocabulary. We study online executable world-model learning under prior misalignment, where an agent must induce state-dependent dynamics from interaction evidence alone, without rule descriptions, reward signals, or trustworthy lexical priors.
We introduce Alice, a closed-loop system that treats failed candidate updates as structural signal: when a candidate explains a new transition but loses previously explained ones, the preservation conflict reveals dynamics that the current program had conflated. Alice refines these conflicts into hypothesis classes that both provide compact, class-stratified preservation counterexamples for update and guide frontier exploration toward transitions that are novel and underrepresented with respect to the current program.
We evaluate Alice on Baba in Wonderland, a prior-misaligned variant of Baba Is You that preserves simulator dynamics while replacing semantically meaningful rule-property labels with unrelated words. Experiments show that Alice substantially improves executable world-model learning under prior misalignment, and ablations show that both class refinement and class-aware exploration contribute.
Two moments on the same map, replayed through saved programs. First, a model learns to move Baba. Then, it learns that pushing one object can move a whole chain.
DEFAULT WORLD YOU controls Baba. PUSH makes an object pushable.
Preparing the comparison images…
v001 copies the input state. v002 moves the object marked YOU, matching the real next state for this action.
v002 moves Baba into the rock's cell without pushing either object. By v005, the program moves both rocks together and matches the real next state.
Fixed images from the supplied Default World programs on the first map, cropped to the action. Each row uses the same input and RIGHT action; the resolved prediction exactly matches the simulator. Red dashed outlines mark errors. These are illustrative replays, not a claim that training encountered these exact scenes.
Choose a map below. Solve it yourself or use Auto to replay a solution.
What if the words stopped helping?
YOU identifies what you control, PUSH what you can push, and WIN the goal. Start here, then switch to Wonderland to keep the scene and change the vocabulary.
Original saved Python, executed for the prediction on the left. Changes compares this revision with the previous saved revision.
Each world opens with its latest supplied program and remembers your selection. Changing worlds or versions keeps your game in place. Saved programs predict each move from the actual pre-action state; they are not trained live here. Revision numbers are specific to each run, not LLM-call counts. These examples do not measure full-benchmark accuracy.
A fix can correct one prediction while breaking another. Alice uses these conflicts to separate past experiences into finer groups, called hypothesis classes. The same groups guide which old behaviors to preserve in the next code update and which unfamiliar dynamics to explore next.
Interact with the environment and find a transition the current program cannot explain.
A rejected update separates preserved and lost transitions into finer hypothesis classes.
Use the same classes to choose preservation examples and seek underrepresented dynamics.
Representative counterexamples guide the next code revision. Acceptance is checked against all previously explained transitions.
Frontier candidates are scored by embedding novelty and expected coverage of rare hypothesis classes.
Exact next-state prediction, including the uncommon transitions that are easy to overlook.
| Method | All accuracy ↑ | Balanced accuracy ↑ | LLM calls |
|---|
All accuracy is exact one-step state matching. Balanced accuracy uses a class-reduced set grouped by action and state-change signature. Results use GPT-5.4, a 100-call cap, and one run per configuration; they are point estimates. Evaluation details ↗
@article{seo2026baba,
title={Baba in Wonderland: Online Self-Supervised Dynamics Discovery for Executable World Models},
author={SeungWon Seo and DongHeun Han and SeongRae Noh and HyeongYeop Kang},
journal={arXiv preprint arXiv:2605.16725},
year={2026},
url={https://arxiv.org/abs/2605.16725}
}Questions about the paper or code? Get in touch ↗