Response time improves alignment with diverse human preferences.
problem Standard aggregation of feedback ignores heterogeneity and anonymity.
method Augmenting feedback with response time data and modeling decisions with DDM.
result Estimator of heterogeneous preferences converges to true average preference.
Study shows how diverse investors' learning and preferences shape financial markets.
problem Understanding how diverse investor behaviors and preferences affect market dynamics.
method Developed a multi-agent reinforcement learning framework with heterogeneous preferences and learning mechanisms.
result Diverse investors develop differentiated strategies through interaction, leading to realistic market dynamics.
Study optimal investment decisions for diverse risk-tolerant agents.
problem Optimizing investment choices for agents with varying risk preferences.
method Characterizes optimal behavior using certainty equivalents and lognormal risks.
result Derives optimal decision menus under known and uncertain preference distributions.
A new method reduces preference distortion in LLM alignment.
problem Vulnerability of traditional LLM alignment methods to human preference heterogeneity.
method Sign Estimator: A simple, provably consistent, and efficient estimator using binary classification loss.
result Substantially reduces preference distortion over a panel of simulated personas.
This work frames reward modelling from preferences as a causal problem.
problem Reward modelling from preference data for AI alignment.
method Causal inference approach to identify challenges and assumptions.
result Causally-inspired approaches improve model robustness.
Paper introduces Functional Effects Models to account for individual heterogeneity in panel data.
problem Accounting for preference heterogeneity in panel data with machine learning.
method Functional Effects Models using gradient boosting decision trees and deep neural networks to learn individual-specific preference parameters.
result Functional Effects Models outperform traditional models in learning inter-individual heterogeneity and predictive performance.
There is an increasing interest in estimating heterogeneity in causal effects in randomized and observational studies. However, little research has been conducted to understand heterogeneity in an instrumental variables study. In this work, we present a method to estimate heterogeneous causal effects using an instrumen…
The paper sorts big data by revealed preferences, improving consumer and policy decisions.
problem Sorting diverse consumer preferences for big data objects like colleges.
method Endogenous weighting of revealed preferences, considering spillover effects.
result Consistent steady-state solution to counterbalance equilibrium.
Study risk sharing among agents with varying risk preferences.
problem Risk sharing among agents with heterogeneous risk measures.
method Derive explicit solutions for inf-convolution and counter-monotonic inf-convolution under varying risk seeking.
result Explicit solutions for inf-convolution and counter-monotonic inf-convolution can be represented by a generalization of distortion risk measures.
Generalizes risk sharing models to a continuum of agents.
problem Risk sharing among a large number of heterogeneous agents.
method Modeling agents as points in a measure space, using risk measures on a probability space, and deriving dual representations.
result Explicit formulas for specific risk measures (entropic and expected shortfall) and applications to Pareto efficiency.
LoCo-RLHF models diverse human feedback with contextual information.
problem Heterogeneous human feedback from diverse contexts and preferences.
method Low-rank contextual preference model, PRS policy.
result LoCo-RLHF achieves tighter sub-optimality gap than existing methods.
We propose an extended public goods interaction model to study the evolution of cooperation in heterogeneous population. The investors are arranged on the well known scale-free type network, the Barabási-Albert model. Each investor is supposed to preferentially distribute capital to pools in its portfolio based on the …
Method distills reward and strategies from diverse demonstrators.
problem Reward ambiguity and heterogeneity in human demonstrations.
method Reward network distillation to infer task goal and strategies.
result Better recovery of task and strategy rewards.
Method tackles uncertainty in reward models for LLMs from heterogeneous human feedback.
problem Uncertainty in reward models for LLMs from heterogeneous human feedback.
method Heterogeneous preference framework and alternating gradient descent algorithm.
result Established theoretical guarantees for estimator convergence and asymptotic distribution.
In market modeling, one often treats buyers as a homogeneous group. In this paper we consider buyers with heterogeneous preferences and products available in many variants. Such a framework allows us to successfully model various market phenomena. In particular, we investigate how is the vendor's behavior influenced by…
New framework estimates treatment effects based on preferences.
problem Estimating treatment effects with flexible outcomes.
method Preference-based Conditional Treatment Effect (CPTE) framework.
result CPTE provides interpretable targets and new identifiability conditions.
Recent years have witnessed an increased focus on interpretability and the use of machine learning to inform policy analysis and decision making. This paper applies machine learning to examine travel behavior and, in particular, on modeling changes in travel modes when individuals are presented with a novel (on-demand)…
Bayesian model identifies three types of travelers adapting to feedback.
problem Capturing adaptive, feedback-driven travel behavior in heterogeneous individuals.
method Latent Class Reinforcement Learning (LCRL) model with Variational Bayes estimation.
result Three distinct traveler classes identified: context-dependent, persistent exploitative, and exploratory.
GBS uses machine learning to design products based on consumer preferences.
problem Designing products to meet consumer preferences.
method GBS is a discrete choice experiment that uses machine learning to adaptively construct paired comparison questions.
result GBS outperforms existing methods in accuracy and sample efficiency.
This paper develops, in a Brownian information setting, an approach for analyzing the preference for information, a question that motivates the stochastic differential utility (SDU) due to Duffie and Epstein [Econometrica 60 (1992) 353-394]. For a class of backward stochastic differential equations (BSDEs) including th…
The dynamics of many socioeconomic systems is determined by the decision making process of agents. The decision process depends on agent's characteristics, such as preferences, risk aversion, behavioral biases, etc.. In addition, in some systems the size of agents can be highly heterogeneous leading to very different i…
Bayesian framework learns latent preference archetypes for many-objective optimization.
problem Expanding space of trade-offs and context-dependent human values.
method Dirichlet-process mixture model for latent preference archetypes, hybrid queries for efficient information.
result Mixture-aware Bayesian optimization outperforms standard methods on synthetic and real-world benchmarks.
New study shows personalized content recommendations can lead to polarization of user preferences.
problem Personalized content recommendations can alter user preferences, leading to polarization.
method Used a model of preference dynamics to explore how personalized content affects user preferences.
result Standard reward maximization algorithms achieve only constant regret in personalized recommendation environments.
This paper analyzes consumer choices over lunchtime restaurants using data from a sample of several thousand anonymous mobile phone users in the San Francisco Bay Area. The data is used to identify users' approximate typical morning location, as well as their choices of lunchtime restaurants. We build a model where res…
The paper proposes a method to learn and leverage contextual preference distributions for better decision-making.
problem Heterogeneous and context-dependent human preferences in decision-making problems.
method A sequential learning-and-optimization pipeline using a bounded-variance score function gradient estimator to train a predictive model mapping contextual features to preference distributions.
result The approach reduces average post-decision surprise by up to 25 times compared to risk-averse baselines in a ridesharing environment.
Upper bounds on utility for managing heterogeneous collectivised funds.
problem Managing pension funds with diverse investor preferences and mortality.
method Axiomatic approach to define optimal management strategies.
result Asymptotically optimal strategies for maximizing investor utility.
We consider the problem of learning the preferences of a heterogeneous population by observing choices from an assortment of products, ads, or other offerings. Our observation model takes a form common in assortment planning applications: each arriving customer is offered an assortment consisting of a subset of all pos…
Generative model reveals hidden interaction preferences in networks.
problem Separate analysis of community and hierarchy overlooks real-world network complexities.
method Generative model based on node preferences and hierarchical structures exploiting network sparsity.
result Model accurately identifies overall node preferences and discerns subsets with different behaviors.
Optimizes pension mix of PAYGO, EET, and individual savings.
problem Balancing PAYGO, EET, and individual savings in funded pension schemes.
method Solves a Nash equilibrium between pension participants and government, considering age-dependent preferences and optimal asset allocation.
result Identifies critical ages and optimal contribution rates for maximizing overall utility.
In this paper we model the problem of learning preferences of a population as an active learning problem. We propose an algorithm can adaptively choose pairs of items to show to users coming from a heterogeneous population, and use the obtained reward to decide which pair of items to show next. We provide computational…
We propose the Heterogeneous Thurstone Model (HTM) for aggregating ranked data, which can take the accuracy levels of different users into account. By allowing different noise distributions, the proposed HTM model maintains the generality of Thurstone's original framework, and as such, also extends the Bradley-Terry-Lu…
We study the market selection hypothesis in complete financial markets, populated by heterogeneous agents. We allow for a rich structure of heterogeneity: individuals may differ in their beliefs concerning the economy, information and learning mechanism, risk aversion, impatience and 'catching up with Joneses' preferen…
We develop a finite horizon continuous time market model, where risk averse investors maximize utility from terminal wealth by dynamically investing in a risk-free money market account, a stock written on a default-free dividend process, and a defaultable bond, whose prices are determined via equilibrium. We analyze fi…
Autonomous systems can substantially enhance a human's efficiency and effectiveness in complex environments. Machines, however, are often unable to observe the preferences of the humans that they serve. Despite the fact that the human's and machine's objectives are aligned, asymmetric information, along with heterogene…
Develops methods to correct bias in AI feedback for more accurate alignment.
problem Systematic bias in AI feedback compared to human labels.
method Two debiased alignment methods: DDPO and DIPO.
result Methods improve alignment efficiency and performance close to human-labeled data.
FedConPE improves conversational recommender systems efficiency and privacy.
problem Efficiently eliciting user preferences in interactive systems with heterogeneous clients.
method Phase elimination-based federated conversational bandit algorithm with adaptive key term construction.
result Minimizes uncertainty across all dimensions in feature space and offers improved efficiency and privacy.
Study on self-consuming generative models with diverse human curation, focusing on convergence and stability.
problem Analyzing self-consuming generative models with heterogeneous human curation.
method Investigates the asymptotic behavior of retraining dynamics using nonlinear Perron--Frobenius theory and Banach contraction mapping.
result Improves convergence results and provides stability and non-stability analyses for the model.
Optimizes investment strategies for retirees with longevity risk.
problem Maximizing retirement savings under longevity risk for a group of investors.
method Analytic and numerical solutions for investment strategies in both discrete and continuous time models.
result Analytic formulae for optimal investment strategies in both discrete and continuous time models.
This paper characterizes the equilibrium in a continuous time financial market populated by heterogeneous agents who differ in their rate of relative risk aversion and face convex portfolio constraints. The model is studied in an application to margin constraints and found to match real world observations about financi…
We study consumption behaviour in systems with heterogeneous interacting agents. Two different models are introduced, respectively with long and short range interactions among agents. At any time step an agent decides whether or not to consume a good, doing so if this provides positive utility. Utility is affected by i…
The paper optimizes reinsurance under uncertain dependence among insurers.
problem Designing Pareto-optimal reinsurance contracts in a market with uncertain dependence.
method Robust optimization approach assuming known marginal distributions and unspecified dependence structure.
result Characterization of optimal indemnity schedules under worst-case scenario and derivation of optimal two-parameter layer contracts for independent risks.
Adaptive reward models capture individual preferences from human feedback.
problem Learning a reward model that can be specialised to a user.
method Empirical risk minimisation and PAC bound analysis.
result Adaptive reward models benefit from the heterogeneity of user preferences.
The paper optimizes pension policies with guarantees and sustainability constraints.
problem Designing optimal pension policies with guarantees and sustainability constraints.
method Dynamic utility model, stochastic domain, overlapping generations, time-consistent decision criterion.
result Optimal investment/pension policy computed for a general framework.
Study risk sharing with Lambda VaR under diverse beliefs.
problem Risk sharing among agents with different beliefs.
method Use Lambda Value-at-Risk as preference, analyze under heterogeneous beliefs.
result Explicit formulas for risk sharing under various belief scenarios.
This study develops a dynamic inverse optimization framework to recover hidden, time-varying preferences from observed allocation trajectories.
problem The gap between classical optimization theory and real-world practice, especially in the presence of drift and shocks.
method Dynamic inverse optimization framework using a drift-aware estimator grounded in convex analysis and online learning theory.
result Sharp static and dynamic regret bounds for the framework, demonstrating its responsiveness to gradual drift and sudden shocks.
Paper proposes personalized climate control for driver comfort.
problem Limited research on in-vehicle climate control and driver preferences.
method IoT platform for data collection, machine learning for driver behavior recognition, and personalized preference recommendation.
result Prototype demonstrates effective and accurate climate control for driver comfort.
We develop a formalism to study linearized perturbations around the equilibria of a pure exchange economy. With the use of mean field theory techniques, we derive equations for the flow of products in an economy driven by heterogeneous preferences and probabilistic interaction between agents. We are able to show that i…
We study risk-sharing economies where heterogenous agents trade subject to quadratic transaction costs. The corresponding equilibrium asset prices and trading strategies are characterised by a system of nonlinear, fully-coupled forward-backward stochastic differential equations. We show that a unique solution generally…