Back to library

Robotics / NVIDIA

Self-Improving Vision-Language-Action Models with Data Generation via Residual RL

Three key questions about this paper

What problem does Self-Improving Vision-Language-Action Models with Data Generation via Residual RL address?

Wenli Xiao et al. · ICLR 2026 conference paper · Primary paper

What evidence supports the main claim in Self-Improving Vision-Language-Action Models with Data Generation via Residual RL?

In 30 randomized cube-pickup trials, the PLD-trained policy succeeded 30/30 times, compared with 16/30 for RLPD-generated data and 10/30 for additional human data. All three reached 30/30 on peg insertion. The observed difference appeared when the cube entered a corner state: PLD had collected recoveries from that region. [Paper §4.4; Fig. 8]

What limitation should readers know about Self-Improving Vision-Language-Action Models with Data Generation via Residual RL?

The useful dataset is neither raw failure footage nor specialist-only perfection. It contains successful hybrid trajectories: the generalist’s approach, the difficult state it produced, and the specialist-assisted recovery. In the paper’s trajectory visualizations, specialist-only data is concentrated and far from base behavior, while PLD data stays nearer the base distribution and covers varied recoveries. [Paper §§3.2, 4.5; Fig. 10]

2 new free reports left todaySubscribe to Pro for unlimited reading and 10 new paper explanations each month.Upgrade to Pro