R3: streaming feed-forward 3D reconstruction via relative-pose regression — 372M params, 20+ FPS, thousands of frames, open-sourced for NeurIPS 2026

Congrong Xu announced R3 (3D Reconstruction via Relative Regression) is accepted to NeurIPS 2026 (their first first-author conference paper), fully open-sourced (github.com/KevinXu02/R3, Apache-2.0, arXiv:2605.26519). Core idea: instead of regressing every camera in one global frame, a lightweight pairwise pose MLP on a partially fine-tuned Depth Anything 3 backbone predicts confidence-weighted pairwise relative poses, assembled into a consistent global trajectory downstream; a single learned confidence per edge (decoupled into rotation/translation) drives loss weighting, pose aggregation and keyframe-bank management. No recurrent state, no TTT modules, no extra transformer. With 372M parameters (about 1/3 of recent 1B-class models), R3 matches or surpasses SOTA streaming methods on pose estimation and dense reconstruction at 20+ FPS, scaling to thousands of frames under bounded memory; two checkpoints: r3 (indoor/small-coverage) and r3_long (outdoor/long trajectories).

![Fish Audio upgrades its ASR model: speaker identification and inline emotion cues such as [laughter]](/static/img/twitter/FishAudio_2104617131920748857.card.webp)



