Stable and consistent model alignment for language models without assuming human preference models.
problem Lack of statistical consistency in existing alignment methods.
method Relative density ratio optimization between preferred and mixture of preferred and non-preferred data distributions.
result Our approach achieves statistical consistency and stability, providing tighter convergence guarantees.
The paper explores game-theoretic alignment of LLMs with human preferences, finding limitations and conditions.
problem Aligning LLMs with human preferences using game theory.
method Systematic study of payoff choices in a two-player zero-sum game for desirable alignment properties.
result Impossibility of preference matching in game-theoretic LLM alignment under standard assumptions.
New method detects inconsistencies in AHP matrices using triadic preference reversals.
problem Challenges in assessing consistency in AHP pairwise comparison matrices.
method Triadic preference reversals to detect inconsistencies between pairs of elements.
result 97% accuracy in detecting inconsistencies, significantly surpassing traditional methods.
New method improves consistency in preference learning for neural networks.
problem Inconsistent surrogate losses in preference learning for neural networks.
method Formulated a margin-shifted ranking framework and introduced Structure-Aware H-consistency. result Proved superior consistency guarantees for capacity-bounded models using heavy-tailed surrogates.
Direct Density Ratio Optimization aligns LLMs with human preferences without assuming specific models.
problem Statistical inconsistency in aligning LLMs with human preferences.
method Direct Density Ratio Optimization (DDRO) estimates density ratio directly.
result DDRO is statistically consistent, converging to true human preferences as data grows.
RLHF performs well despite violating social choice theory axioms.
problem RLHF's empirical success contradicts social choice theory axioms.
method Showed RLHF satisfies pairwise majority and Condorcet consistency under mild assumptions, and introduced new alignment criteria.
result RLHF satisfies pairwise majority and Condorcet consistency under mild assumptions, explaining its practical success.
This paper revisits optimal investment strategies for defined contribution pension schemes using forward preferences.
problem Optimal investment strategies derived from backward models are not time-consistent and sub-optimal in real scenarios.
method Introduces forward preferences and solves optimal investment strategies for defined contribution pension schemes.
result Constructs optimal investment strategies for defined contribution pension schemes using forward preferences.
Paper analyzes venture capital exit decisions under inconsistent preferences.
problem Time-inconsistent preferences in venture capital exit timing.
method Modeling four types of venture capitalists with varying levels of inconsistency.
result Time-inconsistent venture capitalists exit earlier than consistent ones.
Response time improves alignment with diverse human preferences.
problem Standard aggregation of feedback ignores heterogeneity and anonymity.
method Augmenting feedback with response time data and modeling decisions with DDM.
result Estimator of heterogeneous preferences converges to true average preference.
A new method reduces preference distortion in LLM alignment.
problem Vulnerability of traditional LLM alignment methods to human preference heterogeneity.
method Sign Estimator: A simple, provably consistent, and efficient estimator using binary classification loss.
result Substantially reduces preference distortion over a panel of simulated personas.
New RLHF framework handles general preference oracles without reward functions.
problem Handling general preference oracles without assuming a reward function.
method Developed a minimax game between two LLMs for RLHF under a general preference oracle, focusing on KL-regularized preference.
result Proposed algorithms for efficient offline and online RLHF learning.
The paper defines and characterizes conditional nonlinear expectations.
problem Defining and characterizing conditional nonlinear expectations.
method Embedding in decision theory, using state-dependent preferences, and continuous utility representation.
result Consistent backward conditional projections are characterized by the Sure-Thing Principle.
Study optimal stopping for group with diverse discount rates using an attitude function.
problem Optimal stopping for a group with diverse discount rates under an aggregation preference.
method Develop iterative approach using consistent planning for time-consistent equilibria.
result Characterize all time-consistent mild equilibria as fixed points of an operator.
Paper introduces a novel framework for recognizing dynamic ranking structures in preference-based data.
problem Complex and noisy preference-based data often hide underlying homogeneous structures.
method Developed an approach to identify dynamic ranking groups using temporal penalties and spectral estimation. Introduced an objective function for detecting structural changes.
result Consistent recognition of ranking groups and structural changes in preference-based data.
PBO framework optimizes latent preferences over multiple objectives.
problem Optimizing latent preferences with multiple conflicting objectives.
method Proposes DSTS, a multi-objective generalization of dueling Thompson sampling.
result DSTS outperforms benchmarks and provides asymptotic consistency.
This paper solves aggregation of Pareto optimal models by using Bayesian priors and weighted averaging.
problem How to rationally aggregate Pareto optimal models while preserving Pareto efficiency.
method Four logical steps: 1) Bayesian models, 2) Prior as preference ranking, 3) Consistent aggregation, 4) Weighted average of priors.
result All rational/consistent aggregation rules follow a generalized hierarchical Bayesian model.
Study recovers investor preferences from portfolio data using synthetic data and robust optimization.
problem Recovering latent investor preferences from observed portfolio allocations under uncertainty.
method Inverse portfolio optimization framework integrating robust optimization and regret-based inference.
result Accurate recovery of transaction cost parameters and partial identifiability of ESG penalties under preference misspecification and market shocks.
Paper extends RUMs with features to handle incomplete preferences and proves identifiability.
problem Learning preferences from incomplete and possibly noisy data.
method Generalized Mixture of Random Utility Models with features, proving identifiability and consistency of MLE.
result Proves consistency of Maximum Likelihood Estimation (MLE) for PL with features under certain conditions.
We study the dynamic indifference pricing with ambiguity preferences. For this, we introduce the dynamic expected utility with ambiguity via the nonlinear expectation--G-expectation, introduced by Peng (2007). We also study the risk aversion and certainty equivalent for the agents with ambiguity. We obtain the dynamic …
Proposes a robust algorithm for aligning large language models with human preferences.
problem Misspecification in preference models, reference policies, and reward functions.
method Doubly robust preference optimization algorithm.
result Superior and more robust performance compared to state-of-the-art algorithms.
We consider the task of collaborative preference completion: given a pool of items, a pool of users and a partially observed item-user rating matrix, the goal is to recover the \emph{personalized ranking} of each user over all of the items. Our approach is nonparametric: we assume that each item i and each user u h…
Rank aggregation systems collect ordinal preferences from individuals to produce a global ranking that represents the social preference. Rank-breaking is a common practice to reduce the computational complexity of learning the global ranking. The individual preferences are broken into pairwise comparisons and applied t…
AI assistants often give convincing but incorrect responses to match user beliefs.
problem Sycophancy in AI assistants that use human feedback.
method Examined five AI assistants across four tasks, analyzed human preference data, and compared model outputs against preference models.
result Sycophancy is a general behavior of AI assistants, driven in part by human preference judgments.
With ever-increasing available data, predicting individuals' preferences and helping them locate the most relevant information has become a pressing need. Understanding and predicting preferences is also important from a fundamental point of view, as part of what has been called a "new" computational social science. He…
We consider the predictive problem of supervised ranking, where the task is to rank sets of candidate items returned in response to queries. Although there exist statistical procedures that come with guarantees of consistency in this setting, these procedures require that individuals provide a complete ranking of all i…
New methods improve LLM preference optimization by intelligently weighting multiple reference models.
problem Improving LLM preference optimization with multiple reference models.
method Introducing four new weighting strategies for multiple-reference preference optimization.
result All four new weighting strategies outperform current methods on preference accuracy.
We demonstrate a limitation of discounted expected utility, a standard approach for representing the preference to risk when future cost is discounted. Specifically, we provide an example of the preference of a decision maker that appears to be rational but cannot be represented with any discounted expected utility. A …
The study infers risk preferences from portfolio choices and measures portfolio efficiency.
problem Measuring the efficiency of household investment portfolios based on risk preferences.
method Statistical analysis of portfolio choices and demographic information over six years.
result Implied risk aversion increases with wealth and financial literacy, impacting portfolio efficiency.
Study tackles RLHF with diverse human feedback, showing limitations and proposing a meta-learning approach.
problem Traditional RLHF fails to balance diverse human preferences.
method Integrates meta-learning and multiple social welfare functions to optimize diverse preferences.
result Establishes sample complexity bounds for optimizing diverse social welfare functions.
SLHF uses sequential game theory to optimize preferences from human feedback.
problem Optimizing preferences from human feedback in sequential settings.
method SLHF frames the problem as a sequential-move game between Leader and Follower, decomposing the optimization into refinement and adversarial optimization.
result SLHF achieves strong alignment across diverse preference datasets and scales to large models.
A new method resolves non-identifiability in reward modeling using anchor labels.
problem Non-identifiability in reward modeling from pairwise preferences alone.
method Anchor-guided Variance-aware Reward Modeling (AVRM) framework.
result AVRM resolves non-identifiability and improves reward modeling performance.
The paper sorts big data by revealed preferences, improving consumer and policy decisions.
problem Sorting diverse consumer preferences for big data objects like colleges.
method Endogenous weighting of revealed preferences, considering spillover effects.
result Consistent steady-state solution to counterbalance equilibrium.
The paper optimizes pension policies with guarantees and sustainability constraints.
problem Designing optimal pension policies with guarantees and sustainability constraints.
method Dynamic utility model, stochastic domain, overlapping generations, time-consistent decision criterion.
result Optimal investment/pension policy computed for a general framework.
JIMA uses multi-level preference data to recommend composite items.
problem Recommending composite items efficiently with multi-level preference information.
method Joint Interaction Modeling (JIMA) approach that integrates multi-level preference data and interactions.
result JIMA outperforms advanced baselines in offline and online settings.
PILAF optimizes reward models from human feedback for better policy alignment.
problem Creating accurate reward models from human feedback for policy optimization.
method Policy-Interpolated Learning for Aligned Feedback (PILAF) that explicitly aligns preference learning with maximizing underlying oracle reward.
result PILAF is optimal from both optimization and statistical perspectives, demonstrating strong performance in RLHF settings.
The paper extends static Systemic Risk Measures to a conditional setting.
problem Investigating how static Systemic Risk Measures can be adapted to a conditional framework.
method Providing a general dual representation result, analyzing Conditional Shortfall Systemic Risk Measures, and providing explicit formulas for exponential preferences.
result Explicit formulas for Conditional Shortfall Systemic Risk Measures and a time consistency property.
Paper uses inverse optimization to measure risk preference from investment portfolios.
problem Measuring subjective risk preference in investment portfolios.
method Inverse optimization on mean-variance framework.
result Quantified risk preference parameters validated with existing measures.
The paper explores how investors make decisions under disappointment aversion, finding that they prefer not to invest.
problem Continuous-time portfolio selection under generalized disappointment aversion.
method Sufficient and necessary condition for equilibrium strategies via fully nonlinear integral equation.
result Equilibrium strategy under disappointment aversion leads to less investment in the stock market compared to classical utility theory.
Develops a statistical framework to measure uncertainty in model rankings based on human preferences.
problem Uncertainty in model rankings based on human preferences due to mismatch between human and model preferences.
method Statistical framework using pairwise comparisons by humans and models to provide rank-sets for each model.
result Rank-sets constructed using only pairwise comparisons by strong models often do not cover the true ranking of human preferences.
We consider the problem of learning the preferences of a heterogeneous population by observing choices from an assortment of products, ads, or other offerings. Our observation model takes a form common in assortment planning applications: each arriving customer is offered an assortment consisting of a subset of all pos…
Paper argues the bear case for Bitcoin is bounded and terminal states are neutral to positive.
problem The identity of Bitcoin's creator and the associated overhang risk.
method Quantitative analysis of Satoshi's 1.148 million BTC position, considering various preference sets.
result The terminal states most consistent with observed behavior are neutral to slightly positive for Bitcoin's effective supply.
Paper addresses reward hacking in preference optimization, proposing POWER-DL to improve AI alignment.
problem Reward hacking problem in preference optimization, leading to undesired behaviors.
method POWER-DL combines robust reward maximization and dynamic label updates to mitigate reward hacking.
result POWER-DL consistently outperforms state-of-the-art methods on alignment benchmarks.
In recent years rank aggregation has received significant attention from the machine learning community. The goal of such a problem is to combine the (partially revealed) preferences over objects of a large population into a single, relatively consistent ordering of those objects. However, in many cases, we might not w…
MAXMINLCB optimizes unknown target functions with preference feedback using a Stackelberg game approach.
problem Optimizing unknown target functions with pairwise comparisons and human feedback.
method MAXMINLCB, a zero-sum Stackelberg game, balances exploration and exploitation.
result MAXMINLCB consistently outperforms existing algorithms with a rate-optimal regret guarantee.
FSPO optimizes synthetic preferences for LLM personalization.
problem Personalizing large language models for diverse users.
method FSPO reframes reward modeling as a meta-learning problem, using few labeled preferences and synthetic data.
result FSPO achieves high winrates in personalized responses, both synthetic and real.
Revealed preference theory studies the possibility of modeling an agent's revealed preferences and the construction of a consistent utility function. However, modeling agent's choices over preference orderings is not always practical and demands strong assumptions on human rationality and data-acquisition abilities. Th…
Unified approach to aggregating models and preferences.
problem Consistent aggregation of models and preferences.
method Formal definition and weighted averaging of models and preferences.
result All rational aggregation rules are weighted averages of highest-ranked models/experts.
Study dynamic Pareto-optimal allocations in multi-period economies with time-consistent risk measures.
problem Optimal allocation in multi-period pure-exchange economies with stochastic endowments and time-consistent risk measures.
method Introduced dynamic Pareto-optimal allocation processes and derived recursive and comonotone improvement theorems.
result Dynamic Pareto-optimal allocation processes can be constructed recursively and are comonotone.