Skill-Space Shooting for Autonomous Robot Policy Improvement

Anonymous ICLR 2027 Submission

Video supplement Policy-only evaluations

Methods stay side by side. After the shared initial comparison where shown, only the blue-outlined evaluation plays; the others remain frozen. Context runs at 4× or 8×; captioned highlights play at 1×.

Interpreting These Videos

These are policy-only evaluations, without online repair assistance: changes between iterations reflect what the policy has learned. Where shown, iteration 0 plays together. Later, each method takes a turn while the others remain frozen in the same row, before the comparison advances to the next selected iteration. The blue outline and captions direct attention to the active evaluation. HUMAN denotes a policy trained with additional demonstrations; DSRL is the reinforcement-learning baseline. On-screen labels identify each method, training iteration, and playback speed.