The paper sorts big data by revealed preferences, improving consumer and policy decisions.
problem Sorting diverse consumer preferences for big data objects like colleges.
method Endogenous weighting of revealed preferences, considering spillover effects.
result Consistent steady-state solution to counterbalance equilibrium.
Estimates multi-attribute choice preferences using private signals and matrix factorization.
problem Modeling multi-attribute choice preferences under weak assumptions.
method Generative choice model with latent factor matrices and private signals; multi-stage matrix factorization.
result Validated estimation performance of novel algorithm through simulations.
New model reveals voter preferences from aggregate election data.
problem Infer individual-level voter preferences from aggregate election data.
method Modeling aggregate count data as Poisson binomial, relating probabilities to covariates using logistic and neural networks.
result Model predicts voter preferences at precinct and individual levels.
Activists align with large fund preferences for success.
problem Aligning with large fund preferences increases activist success.
method Analyzed previous proxy voting behavior to estimate preferences and correlated them with activist success.
result Campaigns with higher alignment receive more votes and are more successful.
Study metric learning from limited preference comparisons, showing how low-dimensional structure can still reveal metric information.
problem Learning metric from limited pairwise preference comparisons.
method Ideal point model, divide-and-conquer approach for low-dimensional structure.
result Metric can be jointly identified even with limited comparisons when items exhibit low-dimensional structure.
Paper improves parameter estimation of continuous distributions using preference feedback.
problem Improving parameter estimation of continuous distributions.
method Preference-based M-estimators and deterministic preferences.
result Preference-based estimators achieve an estimation error scaling of O(1/n), significantly faster than sample-only methods.
Improved DPO framework penalizes preference uncertainty to avoid overoptimization.
problem Aligning LLMs to human preferences is challenging due to varied, context-dependent, and ambiguous preferences.
method Developed a pessimistic framework for DPO by introducing preference uncertainty penalization schemes.
result Improved overall performance and better completions on high-uncertainty responses compared to vanilla DPO.
In applications such as recommendation systems and revenue management, it is important to predict preferences on items that have not been seen by a user or predict outcomes of comparisons among those that have never been compared. A popular discrete choice model of multinomial logit model captures the structure of the …
The paper solves portfolio selection for complex preferences in continuous time.
problem Dynamic portfolio selection for nonlinear preferences with time inconsistency.
method Stochastic maximum principle and verification theorems for equilibrium strategies.
result Equilibrium strategies derived in closed form for CRRA and CARA preferences.
Study finds equivalence between MMV and MV preferences with conic constraints.
problem Monotone mean-variance portfolio selection under conic constraints.
method Closed-form solutions for optimal strategies under MMV and MV preferences.
result Optimal strategies coincide with and without the conic constraint.
The definition of preferences assigned to individuals is a concept that concerns many disciplines, from economics, with the search of an acceptable outcome for an ensemble of individuals, to decision making an analysis of vote systems. We are concerned in the phenomena of good selection and economic fairness. In Arrow'…
Bayesian framework learns latent preference archetypes for many-objective optimization.
problem Expanding space of trade-offs and context-dependent human values.
method Dirichlet-process mixture model for latent preference archetypes, hybrid queries for efficient information.
result Mixture-aware Bayesian optimization outperforms standard methods on synthetic and real-world benchmarks.
The paper explores how investors make decisions under disappointment aversion, finding that they prefer not to invest.
problem Continuous-time portfolio selection under generalized disappointment aversion.
method Sufficient and necessary condition for equilibrium strategies via fully nonlinear integral equation.
result Equilibrium strategy under disappointment aversion leads to less investment in the stock market compared to classical utility theory.
In recent years rank aggregation has received significant attention from the machine learning community. The goal of such a problem is to combine the (partially revealed) preferences over objects of a large population into a single, relatively consistent ordering of those objects. However, in many cases, we might not w…
We consider the problem of learning the preferences of a heterogeneous population by observing choices from an assortment of products, ads, or other offerings. Our observation model takes a form common in assortment planning applications: each arriving customer is offered an assortment consisting of a subset of all pos…
Generative machine learning models reveal latent travel behavior characteristics.
problem Understanding complex travel behavior through latent variables.
method Developed a joint tri-partite Bayesian graphical network model using RBM.
result Significant improvement in model likelihood compared to traditional models.
Reduces learning regret with diverse user preferences.
problem Reducing regret in stochastic multi-armed bandit problems with diverse user preferences.
method Formulated a stochastic linear bandits model and proposed a Weighted Upper Confidence Bound (W-UCB) algorithm.
result Achieves constant regret when user preferences are sufficiently diverse.
A framework for anonymized risk sharing without revealing identities or preferences.
problem Risk sharing without revealing individual identities or preferences.
method Axiomatic framework with four key axioms: actuarial fairness, risk fairness, risk anonymity, and operational anonymity.
result The conditional mean risk sharing rule is uniquely characterized by these axioms.
Generative model reveals hidden interaction preferences in networks.
problem Separate analysis of community and hierarchy overlooks real-world network complexities.
method Generative model based on node preferences and hierarchical structures exploiting network sparsity.
result Model accurately identifies overall node preferences and discerns subsets with different behaviors.
Best-of-N sampling reveals reward targets from preference data, influencing N and base distribution choices.
problem Understanding reward extraction from Best-of-N preference data and optimal N and base distribution choices.
method Specialized analysis of preference data via induced conditional distribution, deriving reward targets and design principles.
result Reward targets are explicit functions of N and base distribution, and bounded-class minimizers approach these targets as N grows.
The UN General Debate Corpus analyzes speeches from UN member states to reveal their political positions.
problem Lack of data on state preferences in international politics.
method Text analysis of over 7,700 speeches from 1970-2016.
result Demonstrates how the UN General Debate Corpus can reveal country positions on various policy dimensions.
Meta-Router optimizes LLM selection using gold-standard and preference-based data.
problem Training a high-quality LLM router with combined data sources is challenging due to bias and scarcity.
method Developed an integrative causal router training framework to correct bias and improve routing accuracy.
result Our approach delivers more accurate routing and improves the trade-off between cost and quality.
Study shows how diverse investors' learning and preferences shape financial markets.
problem Understanding how diverse investor behaviors and preferences affect market dynamics.
method Developed a multi-agent reinforcement learning framework with heterogeneous preferences and learning mechanisms.
result Diverse investors develop differentiated strategies through interaction, leading to realistic market dynamics.
Unified framework for aligning LLMs from human feedback.
problem Lack of strong theoretical justification for RLHF and inability to compare methods.
method Reframed alignment as distribution learning from pairwise preferences, proposing three principled objectives.
result Proposed objectives achieve strong non-asymptotic convergence to target LM.
Model combines user preferences and side information for better recommendation.
problem Addressing data sparsity and improving user intent representation.
method Tensor-based model that fuses user preferences with side information.
result Demonstrates effectiveness on standard benchmark datasets.
The paper solves IRL for Bayesian stopping time problems.
problem Identifying optimal actions in Bayesian stopping time problems.
method Novel IRL framework using Bayesian revealed preferences.
result Identifies optimality and constructs cost function estimates.
This paper explores the preference-based top-K rank aggregation problem. Suppose that a collection of items is repeatedly compared in pairs, and one wishes to recover a consistent ordering that emphasizes the top-K ranked items, based on partially revealed preferences. We focus on the Bradley-Terry-Luce (BTL) model…
Study replicates reference-dependent preferences impact on risk-return trade-off in Chinese stock market.
problem Impact of reference-dependent preferences on risk-return trade-off in Chinese stock market.
method Utilized CGO proxy, econometric techniques (Dependent Double Sorting, Fama-MacBeth regressions), and data from 1995-2024.
result Reference-dependent preferences have a weaker or absent positive risk-return relationship in the Chinese market.
Neural networks combining multiple data sources can reverse preferences, affecting decision reliability.
problem Preference reversals in neural networks under pooled data.
method Formalized through Case-Based Decision Theory, analyzed Gram geometry, introduced regularization, and developed auditing methods.
result Pooled refitting can reverse shared preferences, and conditions for preserving preferences are derived.
TX-Ray analyzes and quantifies model knowledge transfer in NLP.
problem Insufficient methods for explaining and quantifying model knowledge transfer in NLP.
method Modified computer vision explainability principle to NLP, visualizing feature preference distributions.
result TX-Ray reveals how self-supervised models learn linguistic abstractions and improves generalization.
New method creates synthetic pseudo-panel data to study transport preferences.
problem Aggregation bias in cross-sectional data.
method Conditional Variational Autoencoder (CVAE) framework for probabilistic model.
result Revealed dynamics of transport preferences and categorized individuals.
Network analysis reveals hidden user preferences for platform steering.
problem Limited understanding of user preferences in data gatekeepers.
method Network science and a new measure called Network Information Patrimony.
result Platforms can use network data to better estimate user preferences.
Study optimizes insurance and investment strategies for risk-averse insurers under ambiguity.
problem Optimizing insurance and investment strategies for risk-averse insurers under ambiguity.
method Solves a coupled FBSDE to derive optimal strategies and value function.
result Optimal consumption, investment, and reinsurance strategies influenced by risk aversion and EIS.
The paper explores using set-level ratings for better user-item preference prediction in recommender systems.
problem Capturing user preferences on individual items using set-level ratings.
method Developed collaborative filtering-based methods to model user behaviors in set-level ratings.
result Collaborative filtering-based models can recover and predict user preferences on individual items using set-level ratings.
Optimizes choice sets to influence group decisions.
problem Maximizing agreement or disagreement in group decisions.
method Discrete choice modeling to develop optimization framework.
result Promoting a choice can be easier than encouraging consensus or discord.
New method learns decisions from collective preferences without individual covariates.
problem Making decisions online without individual covariates.
method Collaborative filtering, matrix completion bandit, ε-greedy policy, online gradient descent, inverse propensity weighting.
result Method outperforms benchmarks and reveals new discoveries.
Paper tackles dueling bandits with delayed feedback, revealing preference bias.
problem Real-world dueling bandit applications often face delays in feedback.
method Introduces biased dueling bandit problem with stochastic delayed feedback, presents two algorithms.
result Two algorithms achieve optimal regret bounds for dueling bandit problems with delay.
System learns user preferences to synthesize materials quickly.
problem Slow material synthesis for novice and expert users.
method Gaussian Process Regression for user preferences, neural network for real-time image predictions.
result Real-time material synthesis enables novice users to generate hundreds of models.
A new method for RLHF using proximal point Nash learning.
problem Capturing real human preferences in RLHF.
method Proximal point Nash learning, embedding self-play updates into a proximal point framework.
result High-probability last-iterate convergence for the combined method.
Proposes a method for ranking items across multiple aspects based on user feedback.
problem No principled solution exists for generating multiple item rankings over different aspects.
method Developed a directional multi-aspect ranking criterion using probabilistic multivariate tensor factorization.
result Demonstrated effectiveness of the proposed method through comprehensive experiments on real datasets.
End-to-end model predicts multiagent trajectories using game theory and neural nets.
problem Predicting trajectories of interacting agents in complex scenarios.
method Hybrid neural net with game-theoretic reasoning, using implicit layers to map preferences to Nash equilibria.
result Trains an interpretable model that predicts future trajectories and transfers to decision making.
Unsupervised model predicts facial attractiveness with high accuracy.
problem Capturing the complexity of facial attractiveness through machine learning.
method Infer probabilistic models of facial preferences using Maximum Entropy and neural networks.
result High prediction accuracy in gender classification of sculpting subjects.
Self-play fine-tuning improves diffusion models for text-to-image generation.
problem Plateauing performance of diffusion models after data saturation.
method Self-play fine-tuning (SPIN-Diffusion) using competition among model versions.
result Significantly improved model performance and human preference alignment.
Study reveals DNNs prefer easy-to-learn cues over essential ones in image recognition.
problem DNNs learn easy-to-learn features that aren't essential to the task.
method WCST-ML training setup with shortcut cues on synthetic and face datasets.
result DNNs converge to solutions focusing on preferred cues, leading to flat minima.
Optimal persuasion involves projecting state vectors onto lower-dimensional 'optimal information manifolds'.
problem Optimal persuasion of another agent observing multi-dimensional data.
method Performing non-linear dimension reduction by projecting state vectors onto the 'optimal information manifold'.
result Optimal information design splits information into 'good' and 'bad' components, revealing only the direction of good information.
Data mining reveals power structures in Bangladeshi newspapers.
problem Understanding the power dynamics and narrative structure in news reporting.
method Named entity recognition to create temporal actor networks from news statements.
result Cliquishness among powerful political leaders in news articles.
PAC Battling-Bandit tackles online learning with subset choice and Plackett-Luce feedback.
problem Identify near-best items in a PL model with subset choice and stochastic feedback.
method Introduces PAC Battling-Bandit problem, studies various feedback models, proposes algorithms with optimal sample complexity.
result Sample complexity is $O\left( \frac{n}{ε^2} \ln \frac{1}δ
ight)$ for WI feedback, Ω(mε2nlnδ1) for TR feedback. Gradient descent on normalized networks reveals sparsity preferences.
problem Understanding the inductive bias of gradient descent on normalized neural nets.
method Analysis of gradient descent on weight-normalized smooth homogeneous neural nets, focusing on SWN and EWN.
result EWN causes weights to be updated in a way that prefers asymptotic relative sparsity.