Reward Learning

UBP2 animation
UBP2: Uncertainty-Balanced Preference Planning for Efficient Preference-based Reinforcement Learning.

arXiv:2606.19328, (arXiv), 2026.
Preference-based RL provides an approach to learning reward models from pairwise comparisons of behaviors, bypassing the need for explicit reward design. However, existing methods typically rely on passive data collection and suffer from poor sample efficiency, especially during the early stages of learning. We introduce a model-based approach that actively directs exploration by jointly reasoning over uncertainties in the reward, dynamics, and value functions. Our method, Uncertainty-Balanced Preference Planning (UBP2), uses ensembles of reward, dynamics, and value function models to evaluate candidate trajectories according to a unified score that combines expected reward, terminal value, and epistemic uncertainty. Planning under this objective yields an explicit tradeoff between exploitation and information acquisition without requiring ad hoc exploration heuristics. Under standard regularity assumptions, we establish sublinear regret guarantees for both finite-horizon and infinite-horizon settings. Empirically, experiments on the Meta-World benchmark show UBP2 achieves substantially higher sample efficiency than model-free preference-based methods and non-optimistic model-based baselines.
@misc{nabail2026ubp2,
 archiveprefix = {arXiv},
 author = {Mohamed Nabail and Leo Cheng and Jingmin Wang and Nicholas Rhinehart},
 eprint = {2606.19328},
 primaryclass = {cs.LG},
 title = {UBP2: Uncertainty-Balanced Preference Planning for Efficient Preference-based Reinforcement Learning},
 url = {https://arxiv.org/abs/2606.19328},
 year = {2026}
}
Residual Reward Models for Preference-based Reinforcement Learning
Residual Reward Models for Preference-based Reinforcement Learning.

arXiv:2507.00611, (arXiv), 2025.
Preference-based Reinforcement Learning (PbRL) provides a way to learn high-performance policies in environments where the reward signal is hard to specify, avoiding heuristic and time-consuming reward design. However, PbRL can suffer from slow convergence speed since it requires training in a reward model. Prior work has proposed learning a reward model from demonstrations and fine-tuning it using preferences. However, when the model is a neural network, using different loss functions for pre-training and fine-tuning can pose challenges to reliable optimization. In this paper, we propose a method to effectively leverage prior knowledge with a Residual Reward Model (RRM). An RRM assumes that the true reward of the environment can be split into a sum of two parts: a prior reward and a learned reward. The prior reward is a term available before training, for example, a user’s ``best guess’’ reward function, or a reward function learned from inverse reinforcement learning (IRL), and the learned reward is trained with preferences. We introduce state-based and image-based versions of RRM and evaluate them on several tasks in the Meta-World environment suite. Experimental results show that our method substantially improves the performance of a common PbRL method. Our method achieves performance improvements for a variety of different types of prior rewards, including proxy rewards, a reward obtained from IRL, and even a negated version of the proxy reward. We also conduct experiments with a Franka Panda to show that our method leads to superior performance on a real robot. It significantly accelerates policy learning for different tasks, achieving success in fewer steps than the baseline. The videos are presented at https://sunlighted.github.io/RRM-web/.
@article{cao2025residualrewardmodelspreferencebased,
      title={Residual Reward Models for Preference-based Reinforcement Learning}, 
      author={Chenyang Cao and Miguel Rogel-García and Mohamed Nabail and Xueqian Wang and Nicholas Rhinehart},
      year={2025},
      eprint={2507.00611},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2507.00611}, 
}
Deep Imitative Models for Flexible Inference, Planning, and Control.

International Conference on Learning Representations, (ICLR), 2020.
Imitation Learning (IL) is an appealing approach to learn desirable autonomous behavior. However, directing IL to achieve arbitrary goals is difficult. In contrast, planning-based algorithms use dynamics models and reward functions to achieve goals. Yet, reward functions that evoke desirable behavior are often difficult to specify. In this paper, we propose Imitative Models to combine the benefits of IL and goal-directed planning. Imitative Models are probabilistic predictive models of desirable behavior able to plan interpretable expert-like trajectories to achieve specified goals. We derive families of flexible goal objectives, including constrained goal regions, unconstrained goal sets, and energy-based goals. We show that our method can use these objectives to successfully direct behavior. Our method substantially outperforms six IL approaches and a planning-based approach in a dynamic simulated autonomous driving task, and is efficiently learned from expert demonstrations without online data collection. We also show our approach is robust to poorly specified goals, such as goals on the wrong side of the road.
@inproceedings{rhinehart2020deep,
 author = {Rhinehart, Nicholas and McAllister, Rowan and Levine, Sergey},
 booktitle = {International Conference on Learning Representations (ICLR)},
 title = {Deep Imitative Models for Flexible Inference, Planning, and Control},
 year = {2020}
}

Jointly Forecasting and Controlling Behavior by Learning from High-Dimensional Data
Jointly Forecasting and Controlling Behavior by Learning from High-Dimensional Data.

2019.
Achieving a precise predictive understanding of the future is difficult, yet widely studied in the natural sciences. Significant research activity has been dedicated to building testable models of cause and effect. From a certain view, the ability to forecast the universe is the “holy grail”; the ultimate goal of science. If we had it, we could anticipate, and therefore (at least implicitly) understand all observable phenomena. The human capability to forecast offers complementary motivation. Critical to our intelligence is our ability to plan behaviors by considering how our actions are likely to result in future payoff, especially in the presence of other collaborative and competitive agents. In this work, we seek to computationally model the future in the presence of agent behavior given rich observations of the environment. The brunt of our focus is to reason about what agents could do, instead of other sources of stochasticity. This focus on future agent behavior allows us to tightly couple and jointly perform forecasting and control. The field of Computer Vision (CV) is focused on designing algorithms to automatically understand images, videos, and other perceptual data. However, the field’s effort to-date focuses on non-interactive, present-focused tasks [79, 81, 158, 184]. Most CV contributions are algorithms to answer questions like “what is that”, and “what happened”, rather than “what could happen”, or “how could I achieve X”. Computer Vision has under-explored reasoning about the interactive and decision-based nature of the world. In contrast, Reinforcement Learning (RL) prioritizes modeling interactions and decisions by focusing on how to design algorithms to evoke behavior that maximizes a scalar reward signal. The resulting learning agents, in order to perform well, must have an understanding of how their current behaviors will affect their prospects of future reward. However, in the dominant paradigm of model-free RL [218], agents reason implicitly about the future. In contrast, model-based RL learns one-step dynamics to estimate “what could happen in the near future”. Yet model-based RL primarily focuses on control, rather than explicitly forecasting a single agent (let alone multiple agents). In this thesis, we consider the problem of designing algorithms to enable computational systems to (1) forecast future behavior of intelligent agents given rich observations of their environments, as well as to (2) use this reasoning for control. We believe these two problems should be tightly integrated and jointly considered, and use them to structure this thesis. We define forecasting to be the problem of estimating the set of possible outcomes of a system, whereas control is the problem of producing actions that generate a single outcome of a system. We often use Imitation Learning and Reinforcement Learning to formulate and situate our work. We contribute forecasting and control approaches to excel in diverse, realistic, single-agent, and multi-agent domains. The first part of the thesis focuses on progressively designing more capable forecasting models. We proceed through approaches to (1) forecast single actions of daily behavior by developing matrix factorization models [169], (2) forecast goal-driven action trajectories of daily behavior by developing Online Inverse Reinforcement Learning models [168, 170], (3) forecast motion trajectories of vehicles by developing a deep reversible generative models [171, 174]. The second part of the thesis focuses on progressively designing more capable models that tightly couple forecasting and control. We discuss (4) forecasting as auxiliary supervision for implicitly-planned control [228], (5) forecasting and explicitly planning with the same model [176], and (6) forecasting and planning future interactions of multiple agents [175].
@phdthesis{rhinehart2019jointly,
 author = {Rhinehart, Nicholas},
 school = {Carnegie Mellon University},
 title = {Jointly Forecasting and Controlling Behavior by Learning from High-Dimensional Data},
 year = {2019}
}

R2P2: A Reparameterized Pushforward Policy for Diverse, Precise Generative Path Forecasting.

Proceedings of the European Conference on Computer Vision, (ECCV), 2018.
We propose a method to forecast a vehicle’s ego-motion as a distribution over spatiotemporal paths, conditioned on features (e.g., from LIDAR and images) embedded in an overhead map. The method learns a policy inducing a distribution over simulated trajectories that is both “diverse” (produces most of the likely paths) and “precise” (mostly produces likely paths). This balance is achieved through minimization of a symmetrized cross-entropy between the distribution and demonstration data. By viewing the simulated-outcome distribution as the pushforward of a simple distribution under a simulation operator, we obtain expressions for the cross-entropy metrics that can be efficiently evaluated and differentiated, enabling stochastic-gradient optimization. We propose concrete policy architectures for this model, discuss our evaluation metrics relative to previously-used degenerate metrics, and demonstrate the superiority of our method relative to state-of-the-art methods in both the Kitti dataset and a similar but novel and larger real-world dataset explicitly designed for the vehicle forecasting domain.
@inproceedings{rhinehart2018r2p2,
 author = {Rhinehart, Nicholas and Kitani, Kris M. and Vernaza, Paul},
 booktitle = {Proceedings of the European Conference on Computer Vision (ECCV)},
 pages = {772--788},
 title = {R2P2: A Reparameterized Pushforward Policy for Diverse, Precise Generative Path Forecasting},
 year = {2018}
}

Human-Interactive Subgoal Supervision for Efficient Inverse Reinforcement Learning
Human-Interactive Subgoal Supervision for Efficient Inverse Reinforcement Learning.

Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems, (AAMAS), 2018.
Humans are able to understand and perform complex tasks by strategically structuring tasks into incremental steps or sub-goals. For a robot attempting to learn to perform a sequential task with critical subgoal states, these subgoal states can provide a natural opportunity for interaction with a human expert. This paper analyzes the benefit of incorporating a notion of subgoals into Inverse Reinforcement Learning (IRL) with a Human-In-The-Loop (HITL) framework. The learning process is interactive, with a human expert first providing input in the form of full demonstrations along with some subgoal states. These subgoal states defines a set of sub-tasks for the learning agent to complete in order to achieve the final goal. The learning agent queries for partial demonstrations corresponding to each sub-task as needed when the learning agent struggles with individual sub-task. The proposed Human Interactive IRL (HI-IRL) framework is evaluated on several discrete path-planning tasks. We demonstrate that subgoal-based interactive structuring of the learning task results in significantly more efficient learning, requiring only a fraction of the demonstration data needed for learning the underlying reward function with a baseline IRL model.
@inproceedings{pan2018human,
 author = {Pan, Xinlei and Ohn-Bar, Eshed and Rhinehart, Nicholas and Xu, Yan and Shen, Yilin and Kitani, Kris M.},
 booktitle = {Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems},
 organization = {International Foundation for Autonomous Agents and Multiagent Systems},
 pages = {1380--1387},
 title = {Human-Interactive Subgoal Supervision for Efficient Inverse Reinforcement Learning},
 year = {2018}
}

First-Person Activity Forecasting from Video with Online Inverse Reinforcement Learning.

IEEE Transactions on Pattern Analysis and Machine Intelligence, (PAMI), 2018.
We address the problem of incrementally modeling and forecasting long-term goals of a first-person camera wearer: what the user will do, where they will go, and what goal they seek. In contrast to prior work in trajectory forecasting, our algorithm, Darko, goes further to reason about semantic states (will I pick up an object?), and future goal states that are far in terms of both space and time. Darko learns and forecasts from first-person visual observations of the user’s daily behaviors via an Online Inverse Reinforcement Learning (IRL) approach. Classical IRL discovers only the rewards in a batch setting, whereas Darko discovers the transitions, rewards, and goals of a user from streaming data. Among other results, we show Darko forecasts goals better than competing methods in both noisy and ideal settings, and our approach is theoretically and empirically no-regret.
@article{rhinehart2018first,
 author = {Rhinehart, Nicholas and Kitani, Kris},
 journal = {IEEE Transactions on Pattern Analysis and Machine Intelligence},
 publisher = {IEEE},
 title = {First-Person Activity Forecasting from Video with Online Inverse Reinforcement Learning},
 year = {2018}
}

First-Person Activity Forecasting with Online Inverse Reinforcement Learning.

The IEEE International Conference on Computer Vision, (ICCV), 2017.
Best Paper Honorable Mention
We address the problem of incrementally modeling and forecasting long-term goals of a first-person camera wearer: what the user will do, where they will go, and what goal they seek. In contrast to prior work in trajectory forecasting, our algorithm, DARKO, goes further to reason about semantic states (will I pick up an object?), and future goal states that are far in terms of both space and time. DARKO learns and forecasts from first-person visual observations of the user’s daily behaviors via an Online Inverse Reinforcement Learning (IRL) approach. Classical IRL discovers only the rewards in a batch setting, whereas DARKO discovers the states, transitions, rewards, and goals of a user from streaming data. Among other results, we show DARKO forecasts goals better than competing methods in both noisy and ideal settings, and our approach is theoretically and empirically no-regret.
@inproceedings{rhinehart2017first,
 author = {Rhinehart, Nicholas and Kitani, Kris M.},
 booktitle = {The IEEE International Conference on Computer Vision (ICCV)},
 pages = {3716--3725},
 title = {First-Person Activity Forecasting with Online Inverse Reinforcement Learning},
 year = {2017}
}

Learning Action Maps of Large Environments Via First-Person Vision.

The IEEE Conference on Computer Vision and Pattern Recognition, (CVPR), 2016.
When people observe and interact with physical spaces, they are able to associate functionality to regions in the environment. Our goal is to automate dense functional understanding of large spaces by leveraging sparse activity demonstrations recorded from an ego-centric viewpoint. The method we describe enables functionality estimation in large scenes where people have behaved, as well as novel scenes where no behaviors are observed. Our method learns and predicts “Action Maps”, which encode the ability for a user to perform activities at various locations. With the usage of an egocentric camera to observe human activities, our method scales with the size of the scene without the need for mounting multiple static surveillance cameras and is well-suited to the task of observing activities up-close. We demonstrate that by capturing appearance-based attributes of the environment and associating these attributes with activity demonstrations, our proposed mathematical framework allows for the prediction of Action Maps in new environments. Additionally, we offer a preliminary glance of the applicability of Action Maps by demonstrating a proof-of concept application in which they are used in concert with activity detections to perform localization.
@inproceedings{rhinehart2016learning,
 author = {Rhinehart, Nicholas and Kitani, Kris M.},
 booktitle = {The IEEE Conference on Computer Vision and Pattern Recognition (CVPR)},
 title = {Learning Action Maps of Large Environments Via First-Person Vision},
 year = {2016}
}