Item

From Planning to Contact: A Three-Level Vision-Tactile-Language-Action Architecture for Humanoid Manipulation with Advantage-Driven Post-Training

Smirnov, Konstantin
Citations
Altmetric:
Department
Robotics
Embargo End Date
Type
Thesis
Date
2026
License
Language
English
Collections
Research Projects
Organizational Units
Journal Issue
Abstract
Humanoid robots promise general-purpose manipulation in unstructured environments, but bridging high-level semantic reasoning with contact-rich dexterous control remains an open challenge. This thesis presents a three-level Vision–Tactile–Language–Action (VTLA) architecture for the Unitree G1 humanoid with Dex3-1 dexterous hands. The hierarchy comprises: System 2 (Qwen3-VL-8B-Thinking, ∼1 Hz), which performs chain-of-thought task decomposition, object grounding, and progress evaluation; System 1 (GR00T N1.5 DiT, 10–30 Hz), which generates 16-step joint trajectory chunks via flow matching conditioned on System 2 features, depth, and robot state; and System 0 (∼100 Hz), a Mixture-of-Experts reactive policy trained via PPO in Isaac Sim that consumes Dex31 tactile pressure (33 sensing elements per hand) and per-joint torques, producing joint corrections and a compact feedback vector for System 1. Unlike dual-system VLAs such as GR00T and π0 , which run both systems synchronously and leave VLM reasoning unused, the proposed architecture decouples reasoning from control frequency: System 2 runs asynchronously on a remote server at ∼1 Hz while Sytem 1 acts locally at 10–30 Hz using cached semantic features. A wrist-mounted Intel D405 camera provides metric depth to ground the policy spatially beyond what RGB alone can supply. RGB-only datasets are augmented offline with Grounding DINO bounding boxes and Depth Anything V2 monocular depth to extend training coverage. Experiments use three datasets: the Unitree Dex3 real-robot BlockStacking dataset (301 episodes), the Unitree Dex1 simulated dataset (204 episodes), and our own Dex3 Sim dataset (158 episodes collected via VR teleoperation in Isaac Sim) combining wrist RGB, metric depth, and Dex3-1 tactile pressure—a modality combination absent from all prior public humanoid datasets. The principal empirical finding is a System 1 bottleneck: all tested flow-matching VLAs collapse to the dataset mean on 102 –103 episodes due to idle-frame gradient domi- nance and the absence of an object-grounding term in the behavioural-cloning loss. Sys- tem 0 transfers successfully to the physical Dex3-1 hand zero-shot, achieving 82% grasp and 99% release in simulation and demonstrating stable grasping on real hardware. The thesis reports this negative result with a precise mechanistic analysis and five falsifiable contributions that constrain the design space for future dexterous VLA work.
Citation
Smirnov, Konstantin, "From Planning to Contact: A Three-Level Vision-Tactile-Language-Action Architecture for Humanoid Manipulation with Advantage-Driven Post-Training," M.S. Thesis, Robotics, MBZUAI, 2026.
Source
Conference
Keywords
tactile, vision-language-action model, humanoid robot, Reasoning, Learning from experience, Simulation
Subjects
Source
Publisher
DOI
Full-text link