<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Robot Perception and Learning Lab</title>
    <description>RPL Homepage</description>
    <link>https://ut-austin-rpl.github.io/rpl.github.io/</link>
    <atom:link href="https://ut-austin-rpl.github.io/rpl.github.io/feed.xml" rel="self" type="application/rss+xml" />
    
      <item>
        <title>EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data</title>
        <description>&lt;p&gt;Human behavior is among the most scalable sources of data for learning physical intelligence, yet how to effectively leverage it for dexterous manipulation remains unclear. While prior work demonstrates human to robot transfer in constrained settings, it is unclear whether large scale human data can support fine grained, high degree of freedom dexterous manipulation. We present EgoScale, a human to dexterous manipulation transfer framework built on large scale egocentric human data. We train a Vision Language Action (VLA) model on over 20,854 hours of action labeled egocentric human video, more than 20 times larger than prior efforts, and uncover a log linear scaling law between human data scale and validation loss. This validation loss strongly correlates with downstream real robot performance, establishing large scale human data as a predictable supervision source. Beyond scale, we introduce a simple two stage transfer recipe: large scale human pretraining followed by lightweight aligned human robot mid training. This enables strong long horizon dexterous manipulation and one shot task adaptation with minimal robot supervision. Our final policy improves average success rate by 54% over a no pretraining baseline using a 22 DoF dexterous robotic hand, and transfers effectively to robots with lower DoF hands, indicating that large scale human motion provides a reusable, embodiment agnostic motor prior.&lt;/p&gt;
</description>
        <pubDate>Tue, 10 Nov 2026 00:00:00 +0000</pubDate>
        <link>https://ut-austin-rpl.github.io/rpl.github.io/publications/2026/02/18/zheng-arxiv26-egoscale/</link>
        <guid isPermaLink="true">https://ut-austin-rpl.github.io/rpl.github.io/publications/2026/02/18/zheng-arxiv26-egoscale/</guid>
      </item>
    
      <item>
        <title>Learning Dexterous Manipulation Using Contact Wrench Guidance From Human Demonstration</title>
        <description>&lt;p&gt;Dexterous robot manipulation can benefit from the abundance of human demonstrations, but transferring such demonstrations to robot policies remains challenging. We present Contact Wrench Guidance from Human Demonstration in Robotic Dexterous Manipulation (CHORD), a framework for long-horizon manipulation of rigid and articulated objects with reinforcement learning. The key idea is object-centric contact wrench space guidance: we represent human and robot motions by the forces and torques they can induce on the object, enabling similarity to be measured by the induced instantaneous motions. This guidance makes reinforcement learning more scalable for contact-rich dexterous manipulation. We further introduce a large-scale simulation benchmark with 4,739 bimanual dexterous manipulation tasks, constructed from motion-capture datasets and reconstructed in-house videos. Evaluated on 1,831 benchmark tasks, CHORD achieves an average success rate of 82.12%, demonstrating strong scalability. CHORD also generalizes to whole-body manipulation from hand-only and third-person demonstrations, achieving a 90.77% success rate, and the learned policies transfer to the real world in both open-loop and closed-loop settings.&lt;/p&gt;
</description>
        <pubDate>Mon, 09 Nov 2026 00:00:00 +0000</pubDate>
        <link>https://ut-austin-rpl.github.io/rpl.github.io/publications/2026/11/09/zhu-corl26-chord/</link>
        <guid isPermaLink="true">https://ut-austin-rpl.github.io/rpl.github.io/publications/2026/11/09/zhu-corl26-chord/</guid>
      </item>
    
      <item>
        <title>T-Rex: Tactile-Reactive Dexterous Manipulation</title>
        <description>&lt;p&gt;The ability to react dynamically to tactile signals has long been considered crucial to agile human-level dexterity. Yet contemporary learning-based Vision-Language-Action (VLA) models for robotic manipulation generally either overlook the tactile modality or are limited to encoders with static cues, due in part to the scarcity of diverse training data and standardized evaluation, architectural constraints in current VLA models, and limitations of static tactile encoders. In this paper, we push the frontier of tactile-reactive manipulation by addressing all of these limitations. We propose a large-scale, 100-hour tactile-rich dataset collected via a novel, data-efficient recipe that prioritizes elementary motor primitives. To effectively exploit naturally high-frequency touch signals without sacrificing the existing capabilities of existing VLAs, we introduce a variable-rate Mixture-of-Transformers (MoT) architecture equipped with a novel temporal tactile VQ-VAE encoder. We demonstrate the effectiveness of tactile-reactive policies on 12 manipulation tasks requiring delicate force control and deformable object manipulation, achieving over 30% higher average success rate than the strongest baseline.&lt;/p&gt;
</description>
        <pubDate>Mon, 09 Nov 2026 00:00:00 +0000</pubDate>
        <link>https://ut-austin-rpl.github.io/rpl.github.io/publications/2026/11/09/niu-corl26-trex/</link>
        <guid isPermaLink="true">https://ut-austin-rpl.github.io/rpl.github.io/publications/2026/11/09/niu-corl26-trex/</guid>
      </item>
    
      <item>
        <title>GaP: A Graph-as-Policy Multi-Agent Self-Learning Harness For Variational Automation Tasks</title>
        <description>&lt;p&gt;For robots to work reliably in commercial and industrial applications, can recent advances in agentic coding systems combine interpretable robot programming with the open-world adaptability of model-free policies? We focus on “Variational Automation” (VA), a class of tasks that have larger variations in object geometry and pose than fixed automation. Model-free policies often struggle to close the reliability gap for VA tasks, which must be executed persistently and reliably in commercial and industrial applications. Motivated by prior work on Task and Motion Planning (TAMP) and the Robot Operating System (ROS), we introduce Graph-as-Policy (GaP), a multi-agent coding harness that generates directed computation graphs with perception, planning, and control nodes from a Modular Open Robot Skill Library (MORSL). GaP then generates an internal simulation environment to rehearse task instances with different graphs in parallel to iteratively refine the graph structure and parameters to improve success rates and throughput. Evaluation with 8 new open VA task benchmarks, 4 in-simulation and 4 in real-world, suggests that GaP can achieve success rates that significantly outperform baselines. Details, code, and data can be found online: https://graph-robots.github.io/gap/&lt;/p&gt;
</description>
        <pubDate>Mon, 09 Nov 2026 00:00:00 +0000</pubDate>
        <link>https://ut-austin-rpl.github.io/rpl.github.io/publications/2026/11/09/chen-corl26-gap/</link>
        <guid isPermaLink="true">https://ut-austin-rpl.github.io/rpl.github.io/publications/2026/11/09/chen-corl26-gap/</guid>
      </item>
    
      <item>
        <title>RoboTTT: Context Scaling for Robot Policies</title>
        <description>&lt;p&gt;Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-Training Robot Policies (RoboTTT), a robot model and training recipe that scale visuomotor context to 8K timesteps, three orders of magnitude beyond state-of-the-art policies, without growing inference latency. At this context length, we unlock new robot capabilities: one-shot in-context imitation from human video demonstrations, on-the-fly policy improvement, robustness to perturbations, and stronger performance on multi-stage, long-horizon tasks. We also observe, for the first time, steady gains in closed-loop performance as pretraining context length scales. At its core, RoboTTT integrates Test-Time Training into robot foundation models such as Vision-Language-Action policies, yielding a sequence model whose recurrent state consists of fast weights, parameters updated by gradient descent during both training and inference, compressing histories into weight space and retrieving contextual information for long-context conditioning. To scale training context length, the recipe combines sequence action forcing with truncated backpropagation through time. On challenging real-robot manipulation tasks, RoboTTT improves overall performance by 87% over the single-step context baseline and fully completes a five-minute, ten-stage assembly task, which no baseline ever does. RoboTTT trained with 8K-timestep context outperforms the same model pretrained with 1K timesteps by 62%, suggesting context length as a new scaling axis for robot foundation models.&lt;/p&gt;
</description>
        <pubDate>Mon, 09 Nov 2026 00:00:00 +0000</pubDate>
        <link>https://ut-austin-rpl.github.io/rpl.github.io/publications/2026/07/16/jiang-arxiv26-robottt/</link>
        <guid isPermaLink="true">https://ut-austin-rpl.github.io/rpl.github.io/publications/2026/07/16/jiang-arxiv26-robottt/</guid>
      </item>
    
      <item>
        <title>SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation</title>
        <description>&lt;p&gt;Training and evaluating robot policies in the real world is costly and difficult to scale. We introduce SimFoundry, a modular and automated system for zero-shot real-to-sim scene construction from a video. SimFoundry generates sim-ready digital twins and supports object, scene, and task editing, enabling the automated generation of diverse digital cousins: affordance-preserving variations of reconstructed real-world scenes. Policies trained on SimFoundry data transfer zero-shot to challenging real tasks involving multi-step manipulation, articulated object interaction, and bimanual interaction, and its digital cousins facilitate generalization to new real-world conditions. Across 7 manipulation tasks and 5 policy architectures, SimFoundry simulation evaluations strongly predict real-world performance, with mean Pearson correlation 0.911 and mean maximum ranking violation 0.018. When evaluating sim-trained policies zero-shot in the real world, policies trained with object, scene, and task cousins in simulation show average task success rate improvements of 17%, 21%, and 40%, respectively.&lt;/p&gt;
</description>
        <pubDate>Mon, 09 Nov 2026 00:00:00 +0000</pubDate>
        <link>https://ut-austin-rpl.github.io/rpl.github.io/publications/2026/06/26/ranawaka-arxiv26-simfoundry/</link>
        <guid isPermaLink="true">https://ut-austin-rpl.github.io/rpl.github.io/publications/2026/06/26/ranawaka-arxiv26-simfoundry/</guid>
      </item>
    
      <item>
        <title>ENPIRE: Agentic Robot Policy Self-Improvement in the Real World</title>
        <description>&lt;p&gt;Achieving dexterous robotic manipulation in the real world heavily relies on human supervision and algorithm engineering, which becomes a central bottleneck in the pursuit of general physical intelligence. Although emerging coding agents can generate code to automate algorithm search, their successes remain largely confined in digital environments. We conjecture that the missing abstraction to automate robotics research is a repeatable feedback loop for real-world policy improvement: reset the scene, execute a policy, verify the outcome, and refine the next iteration. To bridge this gap, we introduce ENPIRE, a harness framework for coding agents that instantiates this physical feedback routine with four core modules: an Environment module (EN) for automatic reset and verification, a Policy Improvement module (PI) that launches policy refinement, a Rollout module (R) to evaluate policies with one or multiple physical robots operating in parallel, and an Evolution module (E) in which coding agents analyze logs, consult literature, improve training infrastructure and algorithm code to address failure modes. This closed-loop system transforms real-world manipulation learning into a controllable optimization procedure, minimizing human effort while allowing fair ablations across training recipe and agent variants. Powered by ENPIRE, frontier coding agents can autonomously train a policy to achieve a 99% success rate on challenging, dexterous manipulation tasks, such as organizing a pin box, fastening a zip tie, and tool use, a process that further accelerates when we dispatch an agent team on a robot fleet. Our results suggest a practical and scalable path toward deploying coding agents to autonomously advancing robotics in the physical world.&lt;/p&gt;
</description>
        <pubDate>Mon, 09 Nov 2026 00:00:00 +0000</pubDate>
        <link>https://ut-austin-rpl.github.io/rpl.github.io/publications/2026/06/18/xiao-arxiv26-enpire/</link>
        <guid isPermaLink="true">https://ut-austin-rpl.github.io/rpl.github.io/publications/2026/06/18/xiao-arxiv26-enpire/</guid>
      </item>
    
      <item>
        <title>GRAIL: Generating Humanoid Loco-Manipulation from 3D Assets and Video Priors</title>
        <description>&lt;p&gt;Scaling humanoid loco-manipulation requires robot-compatible demonstrations across diverse objects, whole-body motions, and scene geometries, but teleoperation and motion capture are difficult to scale because each collection depends on physical setups, instrumented actors, and robot operation. We present GRAIL, a digital generation pipeline that remains fully virtual until deployment: it composes 3D assets, simulator-ready scenes, and priors from video foundation models (VFMs) to synthesize interactions without rebuilding physical environments or teleoperating the robot. Rather than reconstructing unconstrained in-the-wild videos, GRAIL starts from fully specified 3D configurations in which object geometry, camera parameters, metric scale, environment depth, and a robot-proportioned character are known before video generation and reused during reconstruction. This privileged setup better conditions 4D recovery, allowing model-based object tracking, human motion estimation, and interaction-aware optimization to reconstruct metric 4D human-object interaction (HOI) trajectories with reduced depth ambiguity and morphology mismatch. We retarget the recovered motions to a humanoid robot and train complementary task-general trackers: an object-aware latent adaptor for manipulation and a scene-aware tracker for terrain traversal. GRAIL produces over 20,000 sequences spanning pick-up, object manipulation, sitting, and terrain traversal. Using only GRAIL-generated data, we train egocentric visual policies through a sim-to-real pipeline and deploy them on a Unitree G1 humanoid, achieving 84% real-world success on diverse object pick-up and 90% success on stair-climbing.&lt;/p&gt;
</description>
        <pubDate>Mon, 09 Nov 2026 00:00:00 +0000</pubDate>
        <link>https://ut-austin-rpl.github.io/rpl.github.io/publications/2026/06/03/xie-arxiv26-grail/</link>
        <guid isPermaLink="true">https://ut-austin-rpl.github.io/rpl.github.io/publications/2026/06/03/xie-arxiv26-grail/</guid>
      </item>
    
      <item>
        <title>HumanoidMimicGen: Data Generation for Loco-Manipulation via Whole-Body Planning</title>
        <description>&lt;p&gt;Imitation learning is a promising approach for training humanoid robots to both walk and manipulate, but it requires a large number of demonstrations, which are time-intensive and difficult to collect via teleoperation. Existing data-generation algorithms can automatically synthesize demonstrations for manipulators, but they are ineffective on humanoids because their high-dimensional composite action spaces involve arms, legs, and torsos. We present HumanoidMimicGen, a method for generating humanoid legged loco-manipulation data. Our method adapts contact-rich whole-body skills from a handful of source demonstrations to new states, generalizing across changes in object pose. By interleaving these single- and dual-arm skills with whole-body locomotion and manipulation planning, the method generates stable, collision-free data across diverse scenes and layouts. To evaluate our approach, we introduce a new simulated loco-manipulation benchmark containing nine diverse tasks that test humanoid loco-manipulation capabilities. There, we demonstrate that HumanoidMimicGen automatically generates large datasets for imitation learning and enables a systematic study of how data generation and policy learning decisions impact model performance. We show that whole-body visuomotor policies co-trained with data generated by HumanoidMimicGen outperform those trained only on real-world data by 20%.&lt;/p&gt;
</description>
        <pubDate>Mon, 09 Nov 2026 00:00:00 +0000</pubDate>
        <link>https://ut-austin-rpl.github.io/rpl.github.io/publications/2026/05/26/lin-arxiv26-humanoidmimicgen/</link>
        <guid isPermaLink="true">https://ut-austin-rpl.github.io/rpl.github.io/publications/2026/05/26/lin-arxiv26-humanoidmimicgen/</guid>
      </item>
    
      <item>
        <title>DreamZero: World Action Models are Zero-shot Policies</title>
        <description>&lt;p&gt;State-of-the-art Vision-Language-Action (VLA) models excel at semantic generalization but struggle to generalize to unseen physical motions in novel environments. We introduce DreamZero, a World Action Model (WAM) built upon a pretrained video diffusion backbone. Unlike VLAs, WAMs learn physical dynamics by predicting future world states and actions, using video as a dense representation of how the world evolves. By jointly modeling video and action, DreamZero learns diverse skills effectively from heterogeneous robot data without relying on repetitive demonstrations. This results in over 2x improvement in generalization to new tasks and environments compared to state-of-the-art VLAs in real robot experiments. Crucially, through model and system optimizations, we enable a 14B autoregressive video diffusion model to perform real-time closed-loop control at 7Hz. Finally, we demonstrate two forms of cross-embodiment transfer: video-only demonstrations from other robots or humans yield a relative improvement of over 42% on unseen task performance with just 10-20 minutes of data. More surprisingly, DreamZero enables few-shot embodiment adaptation, transferring to a new embodiment with only 30 minutes of play data while retaining zero-shot generalization.&lt;/p&gt;
</description>
        <pubDate>Mon, 09 Nov 2026 00:00:00 +0000</pubDate>
        <link>https://ut-austin-rpl.github.io/rpl.github.io/publications/2026/02/17/ye-arxiv26-dreamzero/</link>
        <guid isPermaLink="true">https://ut-austin-rpl.github.io/rpl.github.io/publications/2026/02/17/ye-arxiv26-dreamzero/</guid>
      </item>
    
  </channel>
</rss>
