Dynamic-Robust Photometric-Semantic Reconstruction for Open-Vocabulary 3D Scene Understanding

Boyu Cai1,2 Li Yang2* Yan Xu3 Wei Liu2 Nian Liu2 Sikui Zhang2 Yan Wang4 Chunfeng Yuan2 Weiming Hu1,2
1ShanghaiTech University 2Institute of Automation, Chinese Academy of Sciences 3The Chinese University of Hong Kong 4Deepeleph Intelligent Technology

* Corresponding author

ECCV 2026

Overview of dynamic-robust photometric-semantic reconstruction

TL;DR: We jointly reconstruct clean novel views and open-vocabulary semantics from sparse, unposed observations of dynamic scenes.

Abstract

The integration of novel view synthesis (NVS) and open-vocabulary segmentation (OVS) has recently yielded powerful feed-forward 3D foundation models. However, their inherent reliance on static-scene assumptions leads to severe misalignment of spatial features in unconstrained dynamic environments. To bridge this critical gap, we propose SPAR, a novel joint semantic-geometric encoding architecture that explicitly isolates transient dynamic noise prior to latent space aggregation. Furthermore, we introduce a dynamic-region-aware end-to-end training paradigm that structurally couples motion estimation with multi-view visual and semantic learning. This unified approach enables the network to inherently resolve motion conflicts and distill multi-view consistent, temporally stable scene representations from dynamic inputs.

Extensive experiments on the challenging D-RE10K benchmark demonstrate that SPAR achieves state-of-the-art performance. Our end-to-end approach yields a PSNR of 22.15 dB and 23.33 dB given only 3 and 4 input views, respectively. Despite being trained in a self-supervised manner, our model achieves an mIoU of 88.5% for motion mask prediction. Semantic synthesis learning also consistently enhances photometric fidelity in novel-view rendering.

Method

SPAR framework for photometric-semantic reconstruction
SPAR couples a Cross-View Dynamic Region Predictor (CV-DRP) with joint photometric-semantic scene encoding. Dynamic regions are suppressed before static multi-view evidence is aggregated and rendered into target-view RGB images and semantic features.

Results

Dynamic Novel-View Synthesis on D-RE10K

Input Views PSNR ↑ SSIM ↑ LPIPS ↓
219.970.6270.339
322.150.7020.283
423.330.7390.263
Novel-view synthesis and semantic segmentation comparisons
Qualitative comparison of novel-view synthesis and semantic segmentation on dynamic scenes.
Dynamic region prediction comparisons
Dynamic region prediction on diverse multi-view sequences. The public evaluator uses the raw CV-DRP output shown as Ours w/o Refine; SAM2 refinement is disabled.

Takeaways

Semantics Helps Reconstruction

Semantic learning supplies structural cues that improve photometric scene reconstruction and novel-view rendering.

Reconstruction Reveals Dynamics

Multi-view reconstruction inconsistencies expose transient content and provide self-supervision for learning dynamic masks without ground-truth motion labels.

Video Results

Each scene shows synchronized novel-view image synthesis and open-vocabulary semantic segmentation along the same camera trajectory.

Scene 01

Novel View Synthesis
Semantic Segmentation

Scene 02

Novel View Synthesis
Semantic Segmentation

Scene 03

Novel View Synthesis
Semantic Segmentation

Scene 04

Novel View Synthesis
Semantic Segmentation

Scene 05

Novel View Synthesis
Semantic Segmentation

Scene 06

Novel View Synthesis
Semantic Segmentation

Scene 07

Novel View Synthesis
Semantic Segmentation

Scene 08

Novel View Synthesis
Semantic Segmentation

Scene 09

Novel View Synthesis
Semantic Segmentation

Scene 10

Novel View Synthesis
Semantic Segmentation

Scene 11

Novel View Synthesis
Semantic Segmentation

Scene 12

Novel View Synthesis
Semantic Segmentation

Scene 13

Novel View Synthesis
Semantic Segmentation

Scene 14

Novel View Synthesis
Semantic Segmentation

Scene 15

Novel View Synthesis
Semantic Segmentation

Scene 16

Novel View Synthesis
Semantic Segmentation

Scene 17

Novel View Synthesis
Semantic Segmentation

Scene 18

Novel View Synthesis
Semantic Segmentation

Scene 19

Novel View Synthesis
Semantic Segmentation

Scene 20

Novel View Synthesis
Semantic Segmentation

BibTeX

@inproceedings{cai2026spar,
  title     = {Dynamic-Robust Photometric-Semantic Reconstruction for
               Open-Vocabulary 3D Scene Understanding},
  author    = {Cai, Boyu and Yang, Li and Xu, Yan and Liu, Wei and
               Liu, Nian and Zhang, Sikui and Wang, Yan and
               Yuan, Chunfeng and Hu, Weiming},
  booktitle = {European Conference on Computer Vision (ECCV)},
  year      = {2026}
}

Acknowledgements

This project builds upon LVSM, RayZer, and WildRayZer. We thank their authors for releasing the code.