Embodiment-Aware Skill Learning for Diverse Robots

Mingyo Seo
Ph.D. Dissertation, The University of Texas at Austin

Supervisors: Yuke Zhu, Luis Sentis

Dissertation | Defense Slides


Defense Recording


Abstract

Robotic systems are rapidly diversifying in both form and application, spanning humanoids, mobile manipulators, and industrial platforms. Achieving generalist robot autonomy therefore requires skills that can be shared across robots while accounting for embodiment-specific constraints. Recent efforts have pursued this goal by scaling data across multiple platforms, learning structural differences and invariances from large datasets. However, robots, as engineered systems, exhibit behavior shaped by design choices in hardware, control, and representation. Rather than relying on learning algorithms to infer this structure from data alone, we develop embodiment-aware skill learning, which explicitly incorporates these system-level factors as structural priors to guide learning toward physically feasible and transferable behaviors.

This dissertation presents a holistic approach that combines hardware interfaces, control abstractions, and embodiment-aware representations. Our findings show that incorporating structural priors complements and strengthens skill learning, enabling robots to accumulate and transfer skills across embodiments while maintaining physical feasibility. This accelerates progress toward generalist robot autonomy, supporting scalable skill acquisition and reliable performance across a broad spectrum of rapidly evolving robot platforms.


Part I: Designing Robot Hardware for Consistent Physical Interaction


A core challenge in robot autonomy is that each platform operates within its own domain, making skills hardware-specific and difficult to transfer. While most hardware components are fixed, interaction with the environment occurs primarily through end-effectors, making them a natural locus for co-design with sensorimotor abstractions. In Part I, we develop a series of end-effector grippers [LEGATO, FORTE] that unify how robots grasp, sense, and manipulate objects. These grippers standardize visual and tactile perception as well as contact interactions across morphologies, allowing abstraction layers to interpret tactile and force feedback consistently. Serving as a shared interface for skill learning, they establish a unified action-observation space that supports cross-embodiment transfer, skill reuse, and scalable data collection. These works illustrate that hardware is not merely a constraint but an active component of robot learning.


Part II: Bridging Control Architectures and Skill Learning


High-degree-of-freedom robots such as humanoids and legged systems offer rich capabilities but are difficult to control due to high-dimensional state-action spaces and complex dynamics. These challenges make skill learning inefficient, as even simple tasks can require large datasets. Introducing abstractions through intermediate controllers can reduce this complexity while preserving expressive robot motion and control. In Part II, we develop hybrid learning frameworks [TRILL, PRELUDE] that manage complex whole-body dynamics within lower-dimensional action spaces by optimizing control outputs under dynamic constraints. These control-driven abstractions simplify demonstrations, improve data efficiency, and enable scalable skill learning on robot systems whose dynamics would otherwise be difficult to learn directly. These works demonstrate how combining model-based and learning-based control architectures enables scalable skill learning by allowing each component to be developed in the domain where it is most effective.


Part III: Learning Embodiment-Aware Representations for Skill Transfer


Achieving scalable robot autonomy requires skills that transfer across diverse robot embodiments while adapting to each platform’s physical constraints. While shared hardware interfaces and control abstractions enable transferable interactions, successful deployment still requires adapting learned behaviors to differences in morphology and kinematics. In Part III, we develop embodiment-aware representations that capture both task intent and embodiment-specific feasibility, enabling skills learned from demonstrations to transfer across robots. Object-centric representations can extract task intent from demonstrations by capturing what should be accomplished in the scene, independent of the robot embodiment [OKAMI]. However, executing this intent on a new robot requires accounting for embodiment-specific feasibility, including morphology, kinematics, self-collision constraints, and joint limits. To address this, we develop C-space-based embodiment-aware representations that support feasible trajectory generation on the target robot [PRESTO]. These works demonstrate how embodiment-aware representations enable flexible skill transfer across robots by separating task intent from embodiment-specific motion generation while fully leveraging each robot’s capabilities.


Future Work


Achieving generalist robot autonomy requires robots to accumulate knowledge across embodiments, transfer that knowledge to new systems, and compose it into behaviors that generalize across tasks and platforms. A central question is how skills can be reused and composed when both task structure and robot embodiment vary. Skill reuse and composition are not purely algorithmic, but are fundamentally shaped by robot morphology, sensing and actuation capabilities, and action abstractions. Building on the structural priors explored in this dissertation, future research will study how control abstractions, hardware design, and representations jointly determine which skill compositions are feasible, reusable, and transferable across robots. This research direction aims to advance the next generation of generalist robot autonomy.

Contact

For questions, please contact Mingyo Seo.