RGB-Pointmap Pretraining for Unified 3D Scene Understanding
Paper • 2604.02546 • Published
UniScene3D is a transformer-based encoder that learns unified scene representations from multi-view colored pointmaps, jointly modeling image appearance and geometry. It extends pretrained CLIP models to learn representations that effectively combine complementary information from images and pointmaps, generalizing across diverse 3D scene understanding tasks.
If you find this work useful, please cite:
@article{mao2026rgb,
title={RGB-Pointmap Pretraining for Unified 3D Scene Understanding},
author={Mao, Ye and Luo, Weixun and Huang, Ranran and Jing, Junpeng and Mikolajczyk, Krystian},
journal={The 19th European Conference on Computer Vision},
year={2026}
}