A curated collection of papers, datasets, benchmarks, and code on spatial and 3D world models for reasoning, generation, prediction, planning, and embodied intelligence.
World Models
│
├── Representation
│ ├── Spatial World Models
│ ├── 3D World Models
│ ├── Video World Models
│ └── Physical World Models
│
├── Capabilities
│ ├── Spatial Memory
│ ├── Cognitive Maps
│ ├── Prediction
│ ├── Reasoning
│ └── Planning
│
└── Applications
├── Embodied AI
├── Robotics
├── Autonomous Driving
└── Vision-Language(-Action) Models
- Spatial & 3D World Models
- Video World Models
- Spatial Memory & Cognitive Maps
- Embodied World Models
- Autonomous Driving World Models
- Planning with World Models
- Spatial Reasoning Models
- 3D Scene Representations
- Benchmarks & Datasets
-
Beyond Pixel Histories, "Beyond Pixel Histories: World Models with Persistent 3D State".
-
MindJourney, "MindJourney: Test-Time Scaling with World Models for Spatial Reasoning".
-
Learning 3D Persistent Embodied World Models (2025) [Paper]
-
3D-Belief: Embodied Belief Inference via Generative 3D World Modeling (2026) [Paper]
-
Embody4D: A Generalist 4D World Model for Embodied AI (2026) [Paper]
- Genie: Generative Interactive Environments (2024) [Paper]
- Genie 2: A Large-Scale Foundation World Model (2024) [Project]
- VideoPoet: A Large Language Model for Zero-Shot Video Generation (2024) [Paper]
- Cognitive Mapping and Planning for Visual Navigation (2018) [Paper]
- Active Neural SLAM (2020) [Paper]
- Neural Topological SLAM for Visual Navigation (2020) [Paper]
- Embodied World Models Emerge from Navigational Tasks in Open-Ended Environments (2025) [Paper]
- GAIA-1: A Generative World Model for Autonomous Driving (2023) [Paper]
- GAIA-1: A Generative World Model for Autonomous Driving (2023) [Paper]
- DriveWM: Driving into the Future with World Models [Paper]
- World4Drive [Paper]
- OccWorld: Learning Occupancy World Models for Autonomous Driving [Paper]
- GenCAD: Generative World Models for Autonomous Driving [Paper]
- MuZero: Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model (2020) [Paper]
- Dreamer: Dream to Control (2020) [Paper]
- TD-MPC: Learning to Plan in Latent Space (2022) [Paper]
- TD-MPC2: Scalable, Generalist World Models for Control (2024) [Paper]
- SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities (2024) [Paper]
- SAT: Spatial Aptitude Training for Multimodal Language Models (2025) [Paper]
- BLINK: Multimodal Large Language Models Can See but Not Perceive (2024) [Paper]
- NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis (2020) [Paper]
- Instant-NGP: Instant Neural Graphics Primitives (2022) [Paper]
- 3D Gaussian Splatting for Real-Time Radiance Field Rendering (2023) [Paper]
- VSR
- BLINK
- SAT
- SPARE3D
- Habitat
- BEHAVIOR-1K
- Open X-Embodiment
- nuScenes
- Waymo Open Dataset
- Argoverse 2
Contributions are welcome. Please submit a pull request to add papers, datasets, benchmarks, codebases, or surveys related to spatial and 3D world models.