ReSiReg: Towards Spatially Consistent Semantics in Language-Conditioned Robotic Tasks

Simon Schwaiger1,2 David Seyser2 Alessandro Scherl2,3 Wilfried Wöber4 Gerald Steinbauer-Wagner1

1 Graz University of Technology 2 University of Applied Sciences Technikum Wien
3 University of Alicante 4 University of Natural Resources and Life Sciences Vienna

CoRL 2026

Comparison of backbone, self-calibration, OTAS, ReSiReg Lite, and ReSiReg Full features for an outdoor bridge scene and the query gravel path.

ReSiReg is a feature reconstruction method that turns noisy dense vision-language embeddings into spatially consistent, language-grounded representations. Based on our findings, we also provide a densely grounded Vit-S sized VLM with only 25M parameters in the vision path for robotic applications.

Abstract

Vision-Language Models (VLMs) enable robots to follow open-language instructions. However, dense VLM embeddings have shown to be noisy and lack spatial consistency. This is problematic for robotic applications, which require simultaneous reasoning over semantics and 3D space. We examine spatial structure across recent VLMs and propose ReSiReg, a feature reconstruction method that uses spatially consistent VLM intermediates to improve dense language-grounded retrieval. ReSiReg clusters intermediates into visual prototypes, derives their language descriptors, and reconstructs each patch as a soft mixture of prototype-level language embeddings. We evaluate quantitatively on open-vocabulary semantic segmentation and 3D mapping across backbones, and qualitatively in real-world manipulation scenes. Quantitative results show improved dense retrieval; manipulation scenes show more spatially consistent target activations. We further provide a compact 25M dense VLM for robotic applications, substantially smaller than and competitive with ViT-B baselines.

Backbone-Agnostic

Applies post hoc to recent VLM families using spatial structure already present in their intermediate features.

No Added Parameters

Reconstructs dense output features from visual prototypes without increasing the underlying model size.

Robot-Ready

Supports online 2D and 3D robotics tasks, with our 25M VLM exceeding 30 FPS in batched Jetson inference.

Method

From visual prototypes to language embeddings

ReSiReg transfers spatial structure from model intermediates into the model’s language-grounded feature space through soft clustering and reconstruction.

ReSiReg pipeline: foundation model probing, visual prototype clustering, masked pooling, and prototype-reconstructed language embedding.
  1. Probe VLMExtract spatially consistent intermediate tokens and language-aligned output tokens from the same frozen model.
  2. Discover PrototypesReduce and cluster intermediate tokens into hard assignments and soft patch-to-prototype affinities.
  3. Ground in LanguagePool output tokens over each visual prototype, optionally incorporating a segmentation prior.
  4. Reconstruct SoftlyRepresent each patch as a mixture of prototype-level language embeddings for spatial consistency.

Results

Evaluation across backbones and tasks

Quantitatively evaluated on 2D open-vocabulary segmentation and 3D semantic mapping, with qualitative evaluations on real-world manipulation scenes.

+4.12 average mIoU points on ORAD-3D

ReSiReg Full average gain across evaluated backbones for outdoor open-vocabulary semantic segmentation.

+9.78 mIoU points on RADIOv4

The largest reported ORAD-3D gain, improving a RADIOv4 + modified SigLIP 2 dense backbone from 19.11 to 28.89 mIoU.

25M parameters in ReSiReg Mini vision path

A compact dense VLM built for robotic deployment, substantially smaller than and competitive with ViT-B alternatives.

>30 FPS batched embedded throughput

ReSiReg Lite and the compact ReSiReg Mini backbone surpass real-time throughput on NVIDIA Jetson Orin.

Qualitative Manipulation Results

Qualitative comparison across EUPE, CLIP, RADIOv3, and dino.txt showing improved spatial consistency and transparent bottle similarity after ReSiReg reconstruction.

Interactive playground

Try ReSiReg in your browser

Based on our findings, we train ReSiReg Mini, a dense ViT-S VLM whose vision path has 25M parameters. Explore ReSiReg Mini on common robotics tasks in the interactive demo hosted on Hugging Face Spaces.

ReSiReg Playground Open full screen ↗

Citation

Cite ReSiReg

If ReSiReg supports your research, please cite our work.

Two closely related works complement ReSiReg: From Language Priors to Field Adaptation: Preference Learning for Traversability Estimation learns traversability preferences directly in the ReSiReg Mini feature space, enabling sample-efficient domain adaptation. OTAS is a methodological precursor to ReSiReg's alignment mechanism for open-vocabulary semantic segmentation and reconstruction. It can be viewed as a hard-clustering variant of ReSiReg, with DINOv2 providing fv and CLIP providing fvl.

@article{schwaiger2026resireg,
  title     = {ReSiReg: Towards Spatially Consistent Semantics in
               Language-Conditioned Robotic Tasks},
  author    = {Schwaiger, Simon and Seyser, David and Scherl, Alessandro
               and W{\"o}ber, Wilfried and Steinbauer-Wagner, Gerald},
  journal   = {arXiv preprint arXiv:2606.19088},
  year      = {2026},
  eprint    = {2606.19088},
  archivePrefix = {arXiv},
  primaryClass  = {cs.RO},
  url       = {https://arxiv.org/abs/2606.19088},
  doi       = {10.48550/arXiv.2606.19088}
}
@article{schwaiger2026traversability,
  title     = {From Language Priors to Field Adaptation:
               Preference Learning for Traversability Estimation},
  author    = {Schwaiger, Simon and Seyser, David and Scherl, Alessandro
               and Ajanovi{\'c}, Zlatan and W{\"o}ber, Wilfried
               and Steinbauer-Wagner, Gerald},
  journal   = {arXiv preprint arXiv:2610.02974},
  year      = {2026},
  eprint    = {2610.02974},
  archivePrefix = {arXiv},
  primaryClass  = {cs.RO},
  url       = {https://arxiv.org/abs/2610.02974},
  doi       = {10.48550/arXiv.2610.02974}
}
@inproceedings{schwaiger2026otas,
  author    = {Schwaiger, Simon and Thalhammer, Stefan and W{\"o}ber, Wilfried
               and Steinbauer-Wagner, Gerald},
  title     = {{OTAS}: Open-vocabulary Token Alignment for Outdoor Segmentation},
  booktitle = {2026 IEEE International Conference on Robotics and Automation (ICRA)},
  year      = {2026},
  pages     = {8707--8714},
  doi       = {10.1109/ICRA57385.2026.11696067},
  url       = {https://doi.org/10.1109/ICRA57385.2026.11696067}
}