Success Rate (SR)
Fraction of tasks completed by visiting all required stops in order and submitting an accepted claim_done.
Agentic spatial intelligence
Benchmarking Spatial Agentic Intelligence in the Wild
The project film
Recorded gameplay across multiple models and settings. Music included.
| Model / effort |
|---|
Fraction of tasks completed by visiting all required stops in order and submitting an accepted claim_done.
Average fraction of ordered intermediate stops reached, measuring partial progress.
Successful completion weighted by route efficiency; shorter successful routes receive higher scores.
Average full-episode 3D distance across all 180 tasks, including failures, sampled at approximately 1 Hz.
Average number of agent turns per task, excluding automatic history-summary responses.
180 tasks per model: 125 outdoor and 55 indoor. SR, CC, and SPL are percentages; PL is the mean full-episode 3D distance.
Each task asks an agent to follow a natural-language instruction and visit destinations in order. The agent chooses its own route and submits a completion claim after the final stop.

S marks the start; numbered markers and dashed lines show visit order.
Watch complete standard-setting runs and their recorded trajectories.
Recording and trajectory, synchronized.
Complete game recordings and positions from start to termination, uniformly accelerated on one clock. All runs use standard settings. Each pair shares the same task and map. Pauses are retained. The map shows horizontal movement, while the video shows the game view and changes in height. Selected examples, not aggregate results.
Hover to play · 6-second clips
Original in-game recordings · Hover to preview, click to enlarge
These environments present diverse terrain and spatial constraints, including uneven ground, water boundaries, narrow corridors, and multilevel spaces.
Ascending stairs, climbing ladders, opening fence gates, or activating buttons to pass through doors.
The agent receives screenshots and execution feedback.
Agent loopControl the player through APIs, keyboard and mouse input.
Use the in-game map to inspect explored areas by panning, zooming or searching.
Run tools through Bash, one operation at a time or combined in shell commands or Python scripts.
An independent evaluator checks position against waypoint coordinates once per second, including during model inference.
White House · Plaza Hotel · Ueno Park
Recorded routes, original observations and agent programs.
In both cases, Astra redirects its search and follows an alternative route to the target, while the unsuccessful models continues exploring within the same local region.
Each route unfolds in its recorded order, normalized to the same animation length. Pauses are omitted; playback does not compare real-world running times or speeds. S / E mark the recorded endpoints; gold stars mark targets.
Together, these cases show how agents use programming to turn visual map information into explicit coordinates and individual movement commands into reusable action sequences.
@misc{cao2026mineodyssey,
title = {Mine Odyssey: Benchmarking Spatial Agentic Intelligence in the Wild},
author = {Cao, Yuxuan and Li, Junlong and Li, Hao and He, Junxian},
year = {2026},
url = {https://mine-odyssey.github.io/}
}