Haoran Hao, Shahram Najam Syed, Jeffrey Ichnowski, Jeff Schneider
Proposes FAR, a framework that lets robots learn from test-time failures to adapt behavior, recover autonomously, and continually improve their policy.
Deployed robot policies inevitably encounter failures. Naive retries repeat mistakes, and many existing recovery methods rely on human intervention.
FAR combines Failure-Contrastive Preference Adaptation to steer policies away from unsuccessful behaviors using preference learning data from failures, with lightweight action perturbations during retries to encourage local exploration. Successful recovery trajectories are incorporated into a training loop for continual policy improvement.
Experiments show FAR substantially improves success rates and robustness, with average gains of 17.6% in simulation and 11.7% in the real world over standard diffusion policies. It also significantly improves data efficiency during continual policy improvement by exploiting informative failure cases.