From Amap (AutoNavi/Alibaba Maps) CV Lab — streaming 3D reconstruction from video input only
Key insight: uses a fixed 12-frame local context window instead of complex long-range memory mechanisms
At each timestep: caches KV features from preceding 11 frames, predicts point map in current camera coords, estimates adjacent relative pose, recovers global trajectory via sequential pose composition
Per-frame computation is independent of sequence length — scales to very long videos (up to 22,000 frames)
Includes a motion-visual rotation refiner and composition-aware pose loss to limit drift
Results: 24.45 FPS on H100, 6.71 GiB memory; ATE 4.35m on Oxford Spires (no loop closure)
Optional loop closure module using DINOv2-SALAD for trajectory refinement on sequences with revisited regions