RL-to-RF Framework: Lightweight Demonstration-Free Generative Control for Whole-Body Soft Robotic Pick-and-Place
Alsarraj, Ibrahim Jamal Aldin
Alsarraj, Ibrahim Jamal Aldin
Citations
Altmetric:
Supervisor
Department
Robotics
Embargo End Date
Type
Thesis
Date
2026
License
Language
English
Collections
Research Projects
Organizational Units
Journal Issue
Abstract
Tendon-driven soft robots offer significant potential for safe, compliant, and contact-rich manipulation due to their intrinsic compliance, hyper-redundancy, and underactuated structure. These properties enable whole-body interaction, adaptive grasping, and force distribution that are difficult to achieve with rigid robotic systems. However, achieving reliable whole-body pick-and-place manipulation under contact-induced robot–object state uncertainty remains a major challenge. The nonlinear, high-dimensional, and history dependent dynamics of tendon-driven soft robots, combined with uncertain contact interactions, complicate accurate modeling and limit the effectiveness of traditional model-based control approaches. Furthermore, learning-based methods such as RL, while promising, often suffer from high computational cost, reward sensitivity, and limited generalization across object sizes and task variations. This thesis investigates the hypothesis that de-
coupling behavior discovery from policy generalization can enable scalable, sensorless, and generalizable manipulation in tendon-driven soft robots. To this end, we propose a Reinforcement Learning to Rectified Flow (RL-to-RF) framework for pick-and-place manipulation with sensorless size-conditioned placing, using SpiRob, a spiral-inspired tendon-driven soft robot as the experimental platform. In the first phase, a hierarchical Soft Actor–Critic (SAC) framework autonomously learns grasping and transport behaviors from scratch in simulation through carefully designed geometry-based reward functions. A grasp surrogate based on enclosure geometry enables stable whole-body grasp formation under contact uncertainty. In the second phase, only two successful transport-and-place rollouts generated by RL, together with limited prior knowledge in the form of a lightweight goal map, are transformed into a conditional generative controller using RF via flow matching. This transformation yields a lightweight model capable of generating open-loop actuation trajectories conditioned on inferred object size. Object size is inferred directly from extreme curling kinematics using motor-encoded tendon displacements, enabling sensorless size-conditioned placing without vision or embedded tactile sensors. The proposed frame-
work is validated in simulation by analyzing the scalability limitations of direct RL, where increasing the number of size-conditioned targets substantially increases training time (approximately 100% increase when scaling from 2 to 5 targets). The learned RL rollouts are then transformed into an RF model, which demonstrates significantly improved control trajectory generalization compared to deterministic and Gaussian MLP baselines, reducing placement error by up to 90% and improving grasp stability from 13–20% to 100%. Finally,
the framework is verified experimentally on a physical SpiRob platform, achieving a 100% success rate in sensorless size-conditioned pick-and-place manipulation. The system further demonstrates robustness to initial object pose perturbations within a circular region of radius 300% of the cylinder radius and dynamic scalability under execution-time scaling ranging from approximately 8.3% to 1200% of nominal duration and beyond.
Citation
Alsarraj, Ibrahim Jamal Aldin, "RL-to-RF Framework: Lightweight Demonstration-Free Generative Control for Whole-Body Soft Robotic Pick-and-Place," M.S. Thesis, Robotics, MBZUAI, 2026.
Source
Conference
Keywords
Demonstration free learning, RL-to-RF, Soft Robotic Grasping
