Detect, Evaluate, and Compose

Making Every Interaction Count: Efficient Real-World RL for Robot Policies

anonymous

Abstract

Real-world reinforcement learning (RL) is crucial for adapting robotic control policies to physical environments. Yet, it faces several challenges: costly environment interactions and human supervision, sparse rewards in long-horizon tasks, and underutilized suboptimal experience from autonomous policy rollouts. To address this, we introduce Detect, Evaluate, and Compose (DEC), an RL framework that minimizes human intervention while enabling efficient autonomous policy improvement. DEC first trains a dynamics-aware discriminator to detect failures and determine when to trigger human takeover and resume autonomous execution. It then evaluates a robust value function to tackle the sparse reward challenge by adopting trajectory-held-out bootstrapping and discounted partial returns. For policy optimization, DEC learns complementary positive and negative sub-policies to exploit both successful and suboptimal experience, and compose them at inference time to enable controllable action generation. Across simulated manipulation environment and long-horizon, dexterous bimanual tasks in the real world, DEC consistently outperforms existing RL and policy fine-tuning methods with limited human intervention, demonstrating that effective real-world RL can make every interaction count.

Succeeding in complex, long horizon real world bimanual tasks

DEC performs efficient RL fine-tuning of robot policies for complex, long-horizon real-world tasks requiring dexterous bimanual manipulation. We demonstrate this capability on a dual-arm ARX A5 robot across four tasks spanning spatial generalization, insertion under limited visual observability, in-hand adjustment, and long-horizon assembly.

These tasks are considerably challenging. Take PenAssembly as an example: the robot must assemble a pen through a long-horizon sequence of coordinated bimanual manipulation, including grasping, reorientation, alignment, and insertion. The task demands high positional and rotational precision, especially during close-contact assembly. Small errors in grasp pose or intermediate placement can propagate across stages, making subsequent operations increasingly difficult and requiring precise coordination throughout the entire sequence. For the USBInsertion task, the robot must additionally infer the connector orientation and execute the corresponding in-hand precise reorientation before precise insertion.

Methodology

Our proposed DEC framework contains three stages of algorithmic designs—Detect, Evaluate, and Compose.

  1. Detect. A discriminator trained under a non-negative PU-learning objective with dynamics-aware representations anticipates the eventual success likelihood of each (s, a) pair and triggers human intervention when success is unlikely.
  2. Evaluate. We augment sparse outcome reward with discriminator feedback, fit horizon-conditioned partial returns, and learn values with trajectory-held-out bootstrapping.
  3. Compose. We utilize positive-negative policy optimization and composition to update the policy. Score functions are learned separately from desirable and suboptimal behaviors and flexibly combined at inference time.

The updated policy then collects new trajectories, which are used to refine all three stages in the next iteration.

How it works in reality?

To maximize the utility of limited human supervision, interventions should occur only when they are likely to prevent rollout failure. If the policy is likely to succeed after an (s, a) pair, autonomous execution is sufficient; otherwise, a timely takeover can substantially increase the chance of task success. We thus train a discriminator to anticipate the eventual outcome following each (s, a) and trigger human intervention in advance when the predicted success likelihood is low, as well as return control to the policy once the likelihood recovers.

Loading discriminator demonstration…

Why it matters? Here are some real failures cases.The discriminator is critical for real-world tasks, where failures occur frequently, exhibit diverse modes, and are often difficult for humans to anticipate early enough for timely intervention.

Results

1. Efficient policy improvement with human intervention.

DEC consistently improves task success with limited environment interactions and human guidance, while requiring the least intervention time among human-in-the-loop methods. Although directly finetuning a pretrained generative policy is challenging with limited experience, DEC effectively uses human guidance to improve the generative policy itself, which in turn enhances autonomous rollout quality and makes subsequent interaction effective.

ϵ~(aτ,st,τ)=(1+w)ϵθ+(aτ,st,τ)−wϵθ−(aτ,st,τ)\tilde{\epsilon}(\mathbf{a}_{\tau},\mathbf{s}_t,\tau)=(1+w)\epsilon_{\theta^+}(\mathbf{a}_{\tau},\mathbf{s}_t,\tau)-w\epsilon_{\theta^-}(\mathbf{a}_{\tau},\mathbf{s}_t,\tau)

At inference time, w≥0w\geq 0 controls the combination of positive and negative branch predictions.

Here, aτ\mathbf{a}_{\tau} is the noisy action chunk, st\mathbf{s}_t is the current state, and τ\tau is generative-model time. ϵθ+\epsilon_{\theta^+} and ϵθ−\epsilon_{\theta^-} are the positive and negative branch predictions.

Results for simulation tasks. DEC demonstrates both highest performance and interaction efficiency across the two simulation tasks.
Results for real-world tasks. DEC shows the strongest performance on complex real-world manipulation tasks. For PenAssembly and USBInsertion, we retain only the base policy and RLT as baselines, since the other methods achieve limited performance under the same interaction budget.

2. Accurate discriminator and separable representations.

We construct a test set from held-out trajectories to evaluate the discriminator's ability to anticipate whether a given (s, a) pair will eventually lead to failure. Our discriminator achieves strong detection performance across simulated and real-world experiments, with particularly favorable recall relative to existing methods.The t-SNE visualization and representation comparisons further support the idea of dynamics-aware pretraining.

Ldyn(φ,ψ)=E(st,at,st+H)∼D+ ⁣[∥fψ(φ(st,at))−st+H∥22]\mathcal{L}_{\mathrm{dyn}}(\varphi,\psi)=\mathbb{E}_{(\mathbf{s}_t,\mathbf{a}_t,\mathbf{s}_{t+H})\sim\mathcal{D}^{+}}\!\left[\left\|f_{\psi}(\varphi(\mathbf{s}_t,\mathbf{a}_t))-\mathbf{s}_{t+H}\right\|_2^2\right]

Dynamics-aware pretraining predicts the next chunk state from the current state-action representation.

D+\mathcal{D}^{+} contains successful trajectories, HH is the action-chunk length, φ\varphi is the encoder, and fψf_{\psi} is the dynamics predictor.

Failure detection on simulated and real-world tasks. We report AUROC, accuracy (Acc.), and recall (Rec.). DEC achieves the highest recall and best or tied-best accuracy across all four tasks.
t-SNE visualization of state-action representations. Dynamics-aware pretraining yields clearer separation between success-like and failure-related (s, a) feature pairs than TACO (Zheng et al., 2023) and RPT (Dong et al., 2025). Light green denotes the positive set, while dark green and red denote the success- and failure-related subsets of the unlabeled set, respectively.
Comparison of latent representations for failure detection. Here, we provide the results of using different representations for downstream PU learning; Our dynamic-aware pretraining outperforms other methods in these experiments.

3. Robust value learning

We visualize state-value predictions along representative trajectories from the real-world PutSausageInPot task. Specifically, we separately examine trajectory-held-out bootstrapping and learned discounted partial returns. The following figure plots model predicted values against normalized trajectory progress, with corresponding observations shown above each panel.

Trajectory-held-out bootstrapping. The left panel compares value estimates with and without trajectory-held-out bootstrapping on a rollout that ultimately fails after the sausage is dropped. With trajectory-level exclusion, bootstrap targets are supplied by value networks trained on other trajectory splits. The resulting values increase as the robot makes progress and decrease sharply at the dropping event, whereas the variant without holdout assigns nearly uniform low values throughout the displayed segment. This suggests that held-out bootstrapping helps distinguish promising intermediate states from the eventual failed outcome, mitigating trajectory-specific overfitting.

Discounted partial returns. The right panel compares value learning using learned discounted partial returns against 3-step TD on a successful rollout. Learned partial returns yield a smoother increase in value toward task completion, while 3-step TD exhibits noticeable fluctuations, particularly during the final stage. This smoother progression supports the use of learned partial returns to stabilize long-range value propagation and provide more consistent evaluations of intermediate task progress.

Gt:t+nHH=∑i=0n−1(γH)i∑k=0H−1γkrt+iH+kG^H_{t:t+nH}=\sum_{i=0}^{n-1}(\gamma^H)^i\sum_{k=0}^{H-1}\gamma^k r_{t+iH+k}

The partial return accumulates discounted rewards over n action chunks of length H.

Tt(Vφˉj)=1N∑n=1N[Gξ(st,st+nH,n)+γnHVφˉj(st+nH)]\mathcal{T}_t(V_{\bar{\varphi}_j})=\frac{1}{N}\sum_{n=1}^{N}\left[G_{\xi}(\mathbf{s}_t,\mathbf{s}_{t+nH},n)+\gamma^{nH}V_{\bar{\varphi}_j}(\mathbf{s}_{t+nH})\right]

Value targets average learned partial returns and trajectory-held-out bootstrap estimates over N backup horizons.

GHG^H is the observed discounted partial return, while GξG_{\xi} is its learned prediction. γ\gamma is the discount factor, rr is the reward, and NN is the number of backup horizons. VφˉjV_{\bar{\varphi}_j} is a held-out target network trained on other trajectory splits.

visualization on the real-world PutSausageInPot task. Left: Effect of trajectory-held-out bootstrapping on a failed rollout; the vertical dotted line marks the dropping event. Right: Comparison of learned discounted partial returns and 3-step TD on a successful rollout. Observations above each panel illustrate the corresponding task progression.

For more experiments and methodological details, please refer to the paper. Thank you.