Applies post hoc to recent VLM families using spatial structure already present in their intermediate features.
Abstract
Vision-Language Models (VLMs) enable robots to follow open-language instructions. However, dense VLM embeddings have shown to be noisy and lack spatial consistency. This is problematic for robotic applications, which require simultaneous reasoning over semantics and 3D space. We examine spatial structure across recent VLMs and propose ReSiReg, a feature reconstruction method that uses spatially consistent VLM intermediates to improve dense language-grounded retrieval. ReSiReg clusters intermediates into visual prototypes, derives their language descriptors, and reconstructs each patch as a soft mixture of prototype-level language embeddings. We evaluate quantitatively on open-vocabulary semantic segmentation and 3D mapping across backbones, and qualitatively in real-world manipulation scenes. Quantitative results show improved dense retrieval; manipulation scenes show more spatially consistent target activations. We further provide a compact 25M dense VLM for robotic applications, substantially smaller than and competitive with ViT-B baselines.
Reconstructs dense output features from visual prototypes without increasing the underlying model size.
Supports online 2D and 3D robotics tasks, with our 25M VLM exceeding 30 FPS in batched Jetson inference.
Method
From visual prototypes to language embeddings
ReSiReg transfers spatial structure from model intermediates into the model’s language-grounded feature space through soft clustering and reconstruction.
- Probe VLMExtract spatially consistent intermediate tokens and language-aligned output tokens from the same frozen model.
- Discover PrototypesReduce and cluster intermediate tokens into hard assignments and soft patch-to-prototype affinities.
- Ground in LanguagePool output tokens over each visual prototype, optionally incorporating a segmentation prior.
- Reconstruct SoftlyRepresent each patch as a mixture of prototype-level language embeddings for spatial consistency.
Results
Evaluation across backbones and tasks
Quantitatively evaluated on 2D open-vocabulary segmentation and 3D semantic mapping, with qualitative evaluations on real-world manipulation scenes.
ReSiReg Full average gain across evaluated backbones for outdoor open-vocabulary semantic segmentation.
The largest reported ORAD-3D gain, improving a RADIOv4 + modified SigLIP 2 dense backbone from 19.11 to 28.89 mIoU.
A compact dense VLM built for robotic deployment, substantially smaller than and competitive with ViT-B alternatives.
ReSiReg Lite and the compact ReSiReg Mini backbone surpass real-time throughput on NVIDIA Jetson Orin.
Qualitative Manipulation Results
Interactive playground
Try ReSiReg in your browser
Based on our findings, we train ReSiReg Mini, a dense ViT-S VLM whose vision path has 25M parameters. Explore ReSiReg Mini on common robotics tasks in the interactive demo hosted on Hugging Face Spaces.
Citation
Cite ReSiReg
If ReSiReg supports your research, please cite our work.
Two closely related works complement ReSiReg: From Language Priors to Field Adaptation: Preference Learning for Traversability Estimation learns traversability preferences directly in the ReSiReg Mini feature space, enabling sample-efficient domain adaptation. OTAS is a methodological precursor to ReSiReg's alignment mechanism for open-vocabulary semantic segmentation and reconstruction. It can be viewed as a hard-clustering variant of ReSiReg, with DINOv2 providing fv and CLIP providing fvl.
@article{schwaiger2026resireg,
title = {ReSiReg: Towards Spatially Consistent Semantics in
Language-Conditioned Robotic Tasks},
author = {Schwaiger, Simon and Seyser, David and Scherl, Alessandro
and W{\"o}ber, Wilfried and Steinbauer-Wagner, Gerald},
journal = {arXiv preprint arXiv:2606.19088},
year = {2026},
eprint = {2606.19088},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
url = {https://arxiv.org/abs/2606.19088},
doi = {10.48550/arXiv.2606.19088}
}
@article{schwaiger2026traversability,
title = {From Language Priors to Field Adaptation:
Preference Learning for Traversability Estimation},
author = {Schwaiger, Simon and Seyser, David and Scherl, Alessandro
and Ajanovi{\'c}, Zlatan and W{\"o}ber, Wilfried
and Steinbauer-Wagner, Gerald},
journal = {arXiv preprint arXiv:2610.02974},
year = {2026},
eprint = {2610.02974},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
url = {https://arxiv.org/abs/2610.02974},
doi = {10.48550/arXiv.2610.02974}
}
@inproceedings{schwaiger2026otas,
author = {Schwaiger, Simon and Thalhammer, Stefan and W{\"o}ber, Wilfried
and Steinbauer-Wagner, Gerald},
title = {{OTAS}: Open-vocabulary Token Alignment for Outdoor Segmentation},
booktitle = {2026 IEEE International Conference on Robotics and Automation (ICRA)},
year = {2026},
pages = {8707--8714},
doi = {10.1109/ICRA57385.2026.11696067},
url = {https://doi.org/10.1109/ICRA57385.2026.11696067}
}