The paper proposes a method to infer multi-objective rewards from preferences.
problem Modeling preferences based on multiple, often competing objectives.
method Modeling priorities lexicographically and inferring multi-objective rewards from observed preferences.
result Lexicographically-ordered rewards provide a better understanding of preferences and improve policies.
Bayesian optimization agent learns user preferences from pairwise comparisons.
problem Learning user preferences from unknown and infinite choices.
method Sequential Bayesian optimization with pairwise comparisons.
result Optimal agent strategy minimizes remaining system uncertainty.
Bayesian method predicts individual and crowd preferences from small data.
problem Difficult to predict preferences from limited personal data and noisy labels.
method Combines matrix factorization with Gaussian processes for scalable inference.
result Method predicts preferences for new users and items not in training set.
SLHF uses sequential game theory to optimize preferences from human feedback.
problem Optimizing preferences from human feedback in sequential settings.
method SLHF frames the problem as a sequential-move game between Leader and Follower, decomposing the optimization into refinement and adversarial optimization.
result SLHF achieves strong alignment across diverse preference datasets and scales to large models.
RL agents optimize only specified features; this project infers unmentioned preferences from the state of the environment.
problem RL agents are indifferent to features not specified in a reward function, leading to unconsidered preferences.
method Developed an algorithm based on Maximum Causal Entropy IRL to infer preferences and side effects from the state of the environment.
result Information from the initial state can infer both side effects to avoid and preferences for environment organization.
We tackle the problem of constructive preference elicitation, that is the problem of learning user preferences over very large decision problems, involving a combinatorial space of possible outcomes. In this setting, the suggested configuration is synthesized on-the-fly by solving a constrained optimization problem, wh…
Study recovers investor preferences from portfolio data using synthetic data and robust optimization.
problem Recovering latent investor preferences from observed portfolio allocations under uncertainty.
method Inverse portfolio optimization framework integrating robust optimization and regret-based inference.
result Accurate recovery of transaction cost parameters and partial identifiability of ESG penalties under preference misspecification and market shocks.
Improved model for analyzing topics, sentiments, and user preferences in online reviews.
problem Inefficient processing of large-scale online review datasets.
method Developed variational inference models (vTSPRA, svTSPRA, ovTSPRA) for faster and more efficient processing of large datasets.
result The new models (svTSPRA, ovTSPRA) achieve better performance and faster convergence compared to the original TSPRA model.
Paper develops methods to optimize policies directly from human feedback without reward inference.
problem Challenges in RLHF, including reward model overfitting and distribution shift.
method Develops two algorithms for RLHF without reward inference, using zeroth-order gradient approximators.
result Establishes polynomial convergence rates and outperforms existing methods in numerical experiments.
Automates debiasing for large language model evaluations through Fisher random walk.
problem Rigorous and scalable evaluation of large language models.
method Semiparametric efficient estimator using Fisher random walk for weighted residual balancing.
result Efficient estimation of contextual preference scores for large language models.
This work frames reward modelling from preferences as a causal problem.
problem Reward modelling from preference data for AI alignment.
method Causal inference approach to identify challenges and assumptions.
result Causally-inspired approaches improve model robustness.
Enhances robo-advisors with client investment preference inference.
problem Accurately inferring clients' investment preferences from past activities.
method Stochastic control framework with continuous-time model and discounting scheme.
result Proves sufficient conditions for client investment preference identifiability.
The study infers risk preferences from portfolio choices and measures portfolio efficiency.
problem Measuring the efficiency of household investment portfolios based on risk preferences.
method Statistical analysis of portfolio choices and demographic information over six years.
result Implied risk aversion increases with wealth and financial literacy, impacting portfolio efficiency.
Meta-Router optimizes LLM selection using gold-standard and preference-based data.
problem Training a high-quality LLM router with combined data sources is challenging due to bias and scarcity.
method Developed an integrative causal router training framework to correct bias and improve routing accuracy.
result Our approach delivers more accurate routing and improves the trade-off between cost and quality.
Robot learns user preferences from brain signals.
problem Decoding user preferences for robot motions from brain signals.
method Proposes a novel approach using electroencephalography to decode user preferences from brain signals.
result Brain signals can reliably infer user preferences for robot trajectories.
Novel representer theorem for metric and preference learning in RKHSs.
problem Metric and preference learning problems in Hilbert spaces.
method Regularization with respect to task structure norm, RKHS representation, and novel algorithm.
result Significant performance improvement over baseline methods in real-world rank inference benchmarks.
Enhances preference learning by incorporating response times into binary choices.
problem Limited information from binary choices about preference strength.
method Combines choices and response times using the EZ diffusion model.
result Response times improve utility estimation for strong preferences.
Kernel ridge regression inference for nonstandard data.
problem Inferential theory for kernel ridge regression with nonstandard data.
method Constructs valid and sharp confidence sets using anti-symmetric multipliers.
result Develops a test for match effects in school matching mechanisms.
PGRec improves recommendation by modeling user-item preferences as a graph and embedding it for better predictions.
problem Sparse user-item data in recommender systems.
method PGRec models user-item preferences as a PrefGraph, then uses deep learning and factorization to embed and predict user preferences.
result PGRec outperforms state-of-the-art methods by up to 3.2% in NDCG@10.
Paper proposes a new RLHF framework for human preference learning.
problem Handling dependent online human preference outcomes with dynamic contexts.
method Two-stage algorithm with ε-greedy followed by exploitation; anti-concentration inequalities and matrix martingale concentration techniques. result Our method achieves optimal regret bound and asymptotic normality of estimators.
Fashion preference is a fuzzy concept that depends on customer taste, prevailing norms in fashion product/style, henceforth used interchangeably, and a customer's perception of utility or fashionability, yet fashion e-retail relies on algorithmically generated search and recommendation systems that process structured d…
Optimizes AI learning with limited human feedback budgets.
problem Optimizing allocation of a fixed annotation budget for AI learning.
method Preference-Calibrated Active Learning (PCAL) using semi-parametric inference.
result Proves asymptotic optimality and robustness of the PCAL estimator.
Agents learn state ambiguity from non-linear sensor data using Gaussian approximations.
problem Learning state representation from non-linear sensor data.
method Second-order Taylor approximation of Gaussian distribution for non-linear measurement functions.
result Induces a preference for states based on inferability from observations.
OSIL learns safe policies from unsafe demonstrations.
problem Offline safe imitation learning with implicit safety.
method Formulates CMDP, infers safety from non-preferred trajectories, learns cost model.
result OSIL learns safer policies without degrading reward performance.
Unsupervised model predicts facial attractiveness with high accuracy.
problem Capturing the complexity of facial attractiveness through machine learning.
method Infer probabilistic models of facial preferences using Maximum Entropy and neural networks.
result High prediction accuracy in gender classification of sculpting subjects.
Given a set of pairwise comparisons, the classical ranking problem computes a single ranking that best represents the preferences of all users. In this paper, we study the problem of inferring individual preferences, arising in the context of making personalized recommendations. In particular, we assume that there are …
Proposes a method to infer ranking properties and top-K rankings with uncertainty quantification.
problem General uncertainty quantification in ranking problems.
method Combinatorial inference framework for the Bradley-Terry-Luce model, generalized to multiple testing.
result Minimax optimal method for inferring top-K rankings with FDR control.
Mitigates biases in reward models using variational inference.
problem Spurious correlations in reward models that align large language models with human preferences.
method Formulates data-generating process, identifies non-spurious latent variables, and uses variational inference to recover them.
result Effective mitigation of spurious correlation issues, yielding more robust reward models.
Hierarchical Partial-Order Models for Ranking
problem Rank aggregation combining ordered lists
method Hierarchical partial-order models
result Bayesian inference for latent poset hierarchy
A tutorial on variational inference for high-dimensional models.
problem Approximating marginal likelihood and posterior in Bayesian models.
method Parametric approach to variational inference.
result Variational inference is now preferred for high-dimensional models and large datasets.
This paper protects rankings from differential privacy breaches.
problem Leakage of personal information in rankings.
method Develops ε-ranking differential privacy and a multistage ranking algorithm.
result Establishes the connection between Mallows model and ε-ranking differential privacy.
The paper proposes a method to learn and leverage contextual preference distributions for better decision-making.
problem Heterogeneous and context-dependent human preferences in decision-making problems.
method A sequential learning-and-optimization pipeline using a bounded-variance score function gradient estimator to train a predictive model mapping contextual features to preference distributions.
result The approach reduces average post-decision surprise by up to 25 times compared to risk-averse baselines in a ridesharing environment.
Generative model reveals hidden interaction preferences in networks.
problem Separate analysis of community and hierarchy overlooks real-world network complexities.
method Generative model based on node preferences and hierarchical structures exploiting network sparsity.
result Model accurately identifies overall node preferences and discerns subsets with different behaviors.
A new framework enables real-time task trade-off control.
problem Conflict between multiple related tasks in a fixed model capacity.
method Formulates MTL as a preference-conditioned multiobjective optimization problem; uses a hypernetwork-based neural network.
result A single model can handle different trade-off preferences among multiple tasks.
MAXMINLCB optimizes unknown target functions with preference feedback using a Stackelberg game approach.
problem Optimizing unknown target functions with pairwise comparisons and human feedback.
method MAXMINLCB, a zero-sum Stackelberg game, balances exploration and exploitation.
result MAXMINLCB consistently outperforms existing algorithms with a rate-optimal regret guarantee.
FSPO optimizes synthetic preferences for LLM personalization.
problem Personalizing large language models for diverse users.
method FSPO reframes reward modeling as a meta-learning problem, using few labeled preferences and synthetic data.
result FSPO achieves high winrates in personalized responses, both synthetic and real.
New method optimizes policies without assuming known link functions between preferences and rewards.
problem Policy alignment with unknown and unrestricted link functions.
method Formulates an f-divergence-constrained reward maximization problem, learning policies directly. result Induces a semiparametric single-index binary choice model for policy alignment.
Model-based collaborative filtering analyzes user-item interactions to infer latent factors that represent user preferences and item characteristics in order to predict future interactions. Most collaborative filtering algorithms assume that these latent factors are static, although it has been shown that user preferen…
Demon aligns diffusion models without retraining or backpropagation.
problem Aligning diffusion models with user preferences.
method Stochastic optimization to control noise distribution.
result Significantly improves aesthetics scores for text-to-image generation.
The paper learns personalized thermal preferences using Bayesian active learning.
problem Learning personalized thermal preferences from occupant feedback.
method Bayesian active learning with unimodality constraints on Gaussian process.
result The method requires fewer observations to learn optimal temperature preferences.
Estimates users' preference for a site over others using engagement data.
problem Lack of data on users' interactions with other sites makes it hard to estimate preferences for a focal site.
method Uses Hierarchical Bayes Method with two estimation techniques: Markov Chain Monte Carlo and Stochastic Gradient with Langevin Dynamics.
result Good support found for the approach to computing personalized share of engagement.
Generates low-dimensional node vectors for graphs with privacy while preserving structural preferences.
problem Publishing graph node vectors can leak sensitive individual information.
method SE-PrivGEmb, a skip-gram based technique with a unified noise tolerance mechanism and negative sampling probabilities.
result Our method outperforms existing methods in structural equivalence and link prediction tasks.
Method distills reward and strategies from diverse demonstrators.
problem Reward ambiguity and heterogeneity in human demonstrations.
method Reward network distillation to infer task goal and strategies.
result Better recovery of task and strategy rewards.
New method learns decisions from collective preferences without individual covariates.
problem Making decisions online without individual covariates.
method Collaborative filtering, matrix completion bandit, ε-greedy policy, online gradient descent, inverse propensity weighting.
result Method outperforms benchmarks and reveals new discoveries.
TripDecoder recovers metro routes and travel times from smart card data.
problem Recover unknown routes and travel times in metro systems.
method Decouples two inference tasks: travel time and route preference.
result TripDecoder improves accuracy and efficiency compared to competitors.
New results for modeling voter probabilities in elections.
problem Modeling voter preferences with aggregate data and individual covariates.
method Maximum likelihood estimation for Poisson binomial distribution, approximated with heteroscedastic Gaussian.
result Existence and curvature results for the MLE of the Poisson binomial likelihood.
Models for recommender systems use latent factors to explain the preferences and behaviors of users with respect to a set of items (e.g., movies, books, academic papers). Typically, the latent factors are assumed to be static and, given these factors, the observed preferences and behaviors of users are assumed to be ge…
Extends Mallows model to handle item indifference in rankings.
problem Real data often contains item indifference, challenging strict preference assumptions.
method Proposes Clustered Mallows Model (CMM) to accommodate item indifference.
result CMM provides a flexible representation of rank collections with ordered clusters.