Detect, Evaluate, and Compose
Making Every Interaction Count: Efficient Real-World RL for Robot Policies
anonymous
Abstract
Real-world reinforcement learning (RL) is crucial for adapting robotic control policies to physical environments. Yet, it faces several challenges: costly environment interactions and human supervision, sparse rewards in long-horizon tasks, and underutilized suboptimal experience from autonomous policy rollouts. To address this, we introduce Detect, Evaluate, and Compose (DEC), an RL framework that minimizes human intervention while enabling efficient autonomous policy improvement. DEC first trains a dynamics-aware discriminator to detect failures and determine when to trigger human takeover and resume autonomous execution. It then evaluates a robust value function to tackle the sparse reward challenge by adopting trajectory-held-out bootstrapping and discounted partial returns. For policy optimization, DEC learns complementary positive and negative sub-policies to exploit both successful and suboptimal experience, and compose them at inference time to enable controllable action generation. Across simulated manipulation environment and long-horizon, dexterous bimanual tasks in the real world, DEC consistently outperforms existing RL and policy fine-tuning methods with limited human intervention, demonstrating that effective real-world RL can make every interaction count.
Succeeding in complex, long horizon real world bimanual tasks
DEC performs efficient RL fine-tuning of robot policies for complex, long-horizon real-world tasks requiring dexterous bimanual manipulation. We demonstrate this capability on a dual-arm ARX A5 robot across four tasks spanning spatial generalization, insertion under limited visual observability, in-hand adjustment, and long-horizon assembly.
Loading videos…
These tasks are considerably challenging. Take PenAssembly as an example: the robot must assemble a pen through a long-horizon sequence of coordinated bimanual manipulation, including grasping, reorientation, alignment, and insertion. The task demands high positional and rotational precision, especially during close-contact assembly. Small errors in grasp pose or intermediate placement can propagate across stages, making subsequent operations increasingly difficult and requiring precise coordination throughout the entire sequence. For the USBInsertion task, the robot must additionally infer the connector orientation and execute the corresponding in-hand precise reorientation before precise insertion.
Loading videos…
Methodology
Our proposed DEC framework contains three stages of algorithmic designs—Detect, Evaluate, and Compose.
- Detect. A discriminator trained under a non-negative PU-learning objective with dynamics-aware representations anticipates the eventual success likelihood of each (s, a) pair and triggers human intervention when success is unlikely.
- Evaluate. We augment sparse outcome reward with discriminator feedback, fit horizon-conditioned partial returns, and learn values with trajectory-held-out bootstrapping.
- Compose. We utilize positive-negative policy optimization and composition to update the policy. Score functions are learned separately from desirable and suboptimal behaviors and flexibly combined at inference time.
The updated policy then collects new trajectories, which are used to refine all three stages in the next iteration.
How it works in reality?
To maximize the utility of limited human supervision, interventions should occur only when they are likely to prevent rollout failure. If the policy is likely to succeed after an (s, a) pair, autonomous execution is sufficient; otherwise, a timely takeover can substantially increase the chance of task success. We thus train a discriminator to anticipate the eventual outcome following each (s, a) and trigger human intervention in advance when the predicted success likelihood is low, as well as return control to the policy once the likelihood recovers.
Loading discriminator demonstration…
Why it matters? Here are some real failures cases.The discriminator is critical for real-world tasks, where failures occur frequently, exhibit diverse modes, and are often difficult for humans to anticipate early enough for timely intervention.
Loading videos…
Results
1. Efficient policy improvement with human intervention.
DEC consistently improves task success with limited environment interactions and human guidance, while requiring the least intervention time among human-in-the-loop methods. Although directly finetuning a pretrained generative policy is challenging with limited experience, DEC effectively uses human guidance to improve the generative policy itself, which in turn enhances autonomous rollout quality and makes subsequent interaction effective.
At inference time, controls the combination of positive and negative branch predictions.
Here, is the noisy action chunk, is the current state, and is generative-model time. and are the positive and negative branch predictions.
PenAssembly and USBInsertion, we retain only the base policy and RLT as baselines, since the other methods achieve limited performance under the same interaction budget.2. Accurate discriminator and separable representations.
We construct a test set from held-out trajectories to evaluate the discriminator's ability to anticipate whether a given (s, a) pair will eventually lead to failure. Our discriminator achieves strong detection performance across simulated and real-world experiments, with particularly favorable recall relative to existing methods.The t-SNE visualization and representation comparisons further support the idea of dynamics-aware pretraining.
Dynamics-aware pretraining predicts the next chunk state from the current state-action representation.
contains successful trajectories, is the action-chunk length, is the encoder, and is the dynamics predictor.
3. Robust value learning
We visualize state-value predictions along representative trajectories from the real-world PutSausageInPot task. Specifically, we separately examine trajectory-held-out bootstrapping and learned discounted partial returns. The following figure plots model predicted values against normalized trajectory progress, with corresponding observations shown above each panel.
Trajectory-held-out bootstrapping. The left panel compares value estimates with and without trajectory-held-out bootstrapping on a rollout that ultimately fails after the sausage is dropped. With trajectory-level exclusion, bootstrap targets are supplied by value networks trained on other trajectory splits. The resulting values increase as the robot makes progress and decrease sharply at the dropping event, whereas the variant without holdout assigns nearly uniform low values throughout the displayed segment. This suggests that held-out bootstrapping helps distinguish promising intermediate states from the eventual failed outcome, mitigating trajectory-specific overfitting.
Discounted partial returns. The right panel compares value learning using learned discounted partial returns against 3-step TD on a successful rollout. Learned partial returns yield a smoother increase in value toward task completion, while 3-step TD exhibits noticeable fluctuations, particularly during the final stage. This smoother progression supports the use of learned partial returns to stabilize long-range value propagation and provide more consistent evaluations of intermediate task progress.
The partial return accumulates discounted rewards over n action chunks of length H.
Value targets average learned partial returns and trajectory-held-out bootstrap estimates over N backup horizons.
is the observed discounted partial return, while is its learned prediction. is the discount factor, is the reward, and is the number of backup horizons. is a held-out target network trained on other trajectory splits.
PutSausageInPot task. Left: Effect of trajectory-held-out bootstrapping on a failed rollout; the vertical dotted line marks the dropping event. Right: Comparison of learned discounted partial returns and 3-step TD on a successful rollout. Observations above each panel illustrate the corresponding task progression.