ICLR 2026

Mango-GS: Enhancing Spatio-Temporal Consistency in Dynamic Scenes Reconstruction using Multi-Frame Node-Guided 4D Gaussian Splatting

Tingxuan Huang1 Haowei Zhu1 Jun-hai Yong1 Hao Pan1 Bin Wang1,2

1School of Software, Tsinghua University   2BNRist

Cook SpinachN3V
Sear SteakN3V
Peel BananaHyperNeRF
ChickenHyperNeRF

Overview

Abstract

Reconstructing dynamic 3D scenes with photorealistic detail and strong temporal coherence remains a significant challenge. Existing Gaussian splatting approaches for dynamic scene modeling often rely on per-frame optimization, which can overfit to instantaneous states instead of capturing underlying motion dynamics. To address this, we present Mango-GS, a multi-frame, node-guided framework for high-fidelity 4D reconstruction.

Mango-GS leverages a temporal Transformer to model motion dependencies within a short window of frames, producing temporally consistent deformations. Temporal modeling is confined to a sparse set of control nodes for efficiency. Each node uses a decoupled canonical position and latent code, providing stable semantic anchors for motion propagation and preventing correspondence drift under large motion.

Approach

Multi-frame node-guided dynamics

Sparse control nodes model motion over short temporal windows, then propagate temporally coherent deformation to the dense Gaussian representation.

Mango-GS pipeline showing multi-frame sampling, sparse control nodes, temporal attention, and Gaussian deformation.
Mango-GS combines decoupled control nodes, temporal attention, input masking, and multi-frame supervision in an end-to-end reconstruction pipeline.
01

Decoupled control nodes

Canonical positions provide stable spatial anchors while latent codes carry motion features without coupling identity to instantaneous deformation.

02

Temporal attention

A compact Transformer learns dependencies across neighboring frames on sparse nodes instead of operating over millions of Gaussians.

03

Robust multi-frame training

Grouped frame sampling, input masking, and multi-frame objectives improve temporal consistency under fast motion and partial observation.

Multi-view video

N3V results

Novel-view renderings for all six scenes. Videos are encoded at 2x dataset playback speed.

Coffee Martini

193.10 FPS

Flame Salmon

154.62 FPS

Cook Spinach

64.18 FPS

Cut Roasted Beef

70.09 FPS

Flame Steak

72.79 FPS

Sear Steak

61.15 FPS

Monocular video

HyperNeRF results

Novel-view renderings across topology-changing and articulated motion sequences, shown at the fixed HyperNeRF review speed.

Peel Banana

174.35 FPS

Chicken

160.84 FPS

3D Printer

138.89 FPS

Broom

360.39 FPS

Real-time rendering

Fast multi-frame inference

Mean deform-and-render speed recorded for the ten displayed scene/view outputs on NVIDIA RTX 3090 GPUs.

N3V mean 102.66 HyperNeRF mean 208.62 Median 146.76

145.0FPS

Reference

Citation

Please cite Mango-GS when using the code, models, or results.

@inproceedings{huang2026mangogs,
  title     = {Mango-GS: Enhancing Spatio-Temporal Consistency in Dynamic Scenes Reconstruction using Multi-Frame Node-Guided 4D Gaussian Splatting},
  author    = {Huang, Tingxuan and Zhu, Haowei and Yong, Jun-hai and Pan, Hao and Wang, Bin},
  booktitle = {International Conference on Learning Representations},
  year      = {2026}
}