GitHub
ai-toolscreative-ai
Original source

ABot-Recon: Streaming 3D Reconstruction from Video

Summary

  • From Amap (AutoNavi/Alibaba Maps) CV Lab — streaming 3D reconstruction from video input only
  • Key insight: uses a fixed 12-frame local context window instead of complex long-range memory mechanisms
  • At each timestep: caches KV features from preceding 11 frames, predicts point map in current camera coords, estimates adjacent relative pose, recovers global trajectory via sequential pose composition
  • Per-frame computation is independent of sequence length — scales to very long videos (up to 22,000 frames)
  • Includes a motion-visual rotation refiner and composition-aware pose loss to limit drift
  • Results: 24.45 FPS on H100, 6.71 GiB memory; ATE 4.35m on Oxford Spires (no loop closure)
  • Optional loop closure module using DINOv2-SALAD for trajectory refinement on sequences with revisited regions
  • Python API: simple ABotRecon.from_pretrained() → model.infer(images) interface
  • Uses FlashInfer for paged KV-cache acceleration, falls back to PyTorch SDPA
  • Requires: Linux, Python 3.10+, PyTorch 2.5.1, CUDA 12.1