Retrieval-Augmented Vision-Language-Action Policies for Cross-Embodiment Generalization in Robotic Manipulation
Jiacheng Geng
Simen Owen Academic Proceedings Series, 2026, vol. 7, 226-236
Abstract:
Robotic manipulation is gradually evolving from performing tasks in fixed scenarios to the flexible deployment of robots, tools, and complex environments. The Vision-Language-Action (VLA) paradigm enables robots to generate executable actions by integrating visual perception with natural language instructions. However, when robots differ in kinematic structures, end-effector configurations, observation perspectives, or action scales, the transferability of learned policies remains severely limited. Existing VLA methods and retrieval-based strategies typically reuse historical task experiences without considering whether those experiences are physically compatible with the current robot embodiment. To address this challenge, this study proposes an enhanced retrieval-augmented VLA framework that explicitly incorporates robot body information into the experience retrieval process. The proposed method improves the applicability of historical trajectories through compatibility screening and target robot action-space mapping. Specifically, the framework retrieves similar operation demonstrations by jointly integrating visual, linguistic, and embodiment features, then selects the most effective experiences based on the physical characteristics of the target robot to generate corresponding actions. Five experiments were conducted on a subset of the publicly available Open X-Embodiment dataset. Results demonstrate that, compared with the strongest baseline, the proposed method achieves notable improvements in action prediction accuracy, cross-embodiment transfer capability, and vision-language retrieval compatibility. These findings indicate that retrieval methods that account for a robot's own physical characteristics can effectively enhance action prediction accuracy, task transferability, and result interpretability in cross-embodiment robotic operations.
Keywords: robotic manipulation; vision-language-action policy; retrieval-augmented learning; cross-embodiment generalization; embodiment-aware retrieval (search for similar items in EconPapers)
Date: 2026
References: Add references at CitEc
Citations:
Downloads: (external link)
https://soapubs.com/index.php/SOAPS/article/view/2550/2320 (application/pdf)
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:axf:soapsa:v:7:y:2026:i::p:226-236
Access Statistics for this article
More articles in Simen Owen Academic Proceedings Series from Scientific Open Access Publishing
Bibliographic data for series maintained by Yuchi Liu ().