Item

Toward Task Generalization in Vision-Language-Action Models through Structured Visual Goal Conditioning

Alghaithi, Maitha Hamdan Saeed
Citations
Altmetric:
Supervisor
Department
Robotics
Embargo End Date
Type
Thesis
Date
2026
License
Language
English
Collections
Research Projects
Organizational Units
Journal Issue
Abstract
Large vision-language models' semantic reasoning capabilities are combined with robotic control in Vision-Language-Action (VLA) models to enable policies that obey natural language instructions in a variety of manipulation tasks. There is, however, a persistent gap between what these models comprehend and what they can consistently perform: policies often correctly interpret task intent but fail during physical manipulation, especially under distribution shifts in object appearance, workspace layout, and scene composition. The goal representation itself is one contributing factor. While standard natural language instructions offer semantic flexibility, they also introduce spatial ambiguity that makes precise manipulation more difficult. To improve the generalization capabilities of VLA policies in pick-and-place tasks, this research investigates structured visual goal conditioning. The proposed approach derives per-timestep pixel-space UV coordinates of the target object through instance segmentation. cThese coordinates are then encoded as succinct language instructions. These instructions define the spatial goal location within the camera frame and the desired gripper action, rather than relying on unstructured language descriptions. This method distinguishes goal specification from object identity and appearance by directly linking task instructions to the policy's visual observation space. The complete pipeline was constructed and evaluated. NVIDIA Isaac Sim's simulation-driven data generation framework produced 9,500 expert demonstration episodes, incorporating detailed goal metadata and systematic scene randomization. These demonstrations served as the training data for SmolVLA, a VLA model comprising 450 million parameters, which was trained using UV-conditioned instructions. For closed-loop policy evaluation, a dual-process deployment architecture overcomes the incompatible dependencies between the simulation runtime and the training framework. Three conditions of increasing distributional difficulty are used to assess the trained policy: visual generalization with new object shapes, distractors, and camera perturbations; clean task transfer with zero-shot stacking and sorting; and training-distribution replay with matched clutter statistics. Distance-based behavioral analysis suggests that UV conditioning may elicit goal-directed approach behavior across all experimental conditions. In particular, the end-effector appears to reduce its distance to the target by approximately 26–32 cm during the initial approach phase, even though the policy does not successfully complete the task under any condition. A consistent failure mode is observed in the form of a vertical bias, where the end-effector remains several centimeters above the target object, preventing successful grasp contact. One possible explanation for this behavior is the absence of explicit depth information in the 2D pixel-space goal representation, which may limit accurate vertical alignment. Cross-condition trends indicate that performance tends to be higher in training-distribution scenes compared to geometrically simpler but distributionally mismatched clean scenes, and may degrade as distributional shift increases. Notably, some episodes under generalization conditions approach the median performance observed in training-distribution settings, suggesting that UV conditioning may retain some robustness under mild perturbations. Overall, these results point toward the potential benefit of incorporating explicit depth information to enable more reliable task completion. They also suggest that pixel-space goal conditioning can provide useful lateral spatial grounding for VLA policies, though further validation is needed.
Citation
Alghaithi, Maitha Hamdan Saeed, "Toward Task Generalization in Vision-Language-Action Models through Structured Visual Goal Conditioning," M.S. Thesis, Robotics, MBZUAI, 2026.
Source
Conference
Keywords
Robotic manipulation, Vision-Language-Action models, UV Goal conditioning, Pick-and-place
Subjects
Source
Publisher
DOI
Full-text link