X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
ICLR 2026, Poster
X-VLA is a scalable vision-language-action model designed to learn from heterogeneous data collected across different robotic embodiments. It introduces embodiment-specific learnable soft prompts into a unified Transformer architecture, enabling the model to exploit shared knowledge while preserving embodiment-specific characteristics.
Contribution: Contributed to simulation experiments and real-robot data collection.