Enhances preference learning by incorporating response times into binary choices.
problem Limited information from binary choices about preference strength.
method Combines choices and response times using the EZ diffusion model.
result Response times improve utility estimation for strong preferences.
Paper analyzes venture capital exit decisions under inconsistent preferences.
problem Time-inconsistent preferences in venture capital exit timing.
method Modeling four types of venture capitalists with varying levels of inconsistency.
result Time-inconsistent venture capitalists exit earlier than consistent ones.
The paper enhances preference learning by incorporating response time data.
problem Lack of temporal information in user decision-making for reward model learning.
method Integrates response time alongside binary choice data using the EZ model and Neyman-orthogonal loss functions.
result Response time-augmented approach reduces error rates from exponential to polynomial scaling, improving sample efficiency.
Response time improves alignment with diverse human preferences.
problem Standard aggregation of feedback ignores heterogeneity and anonymity.
method Augmenting feedback with response time data and modeling decisions with DDM.
result Estimator of heterogeneous preferences converges to true average preference.
Paper analyzes finite-time guarantees for preference-based RL.
problem Understanding finite-time guarantees for preference-based RL.
method Combines dueling bandits and policy search to navigate state space.
result Identifies best policy up to accuracy ε with high probability.
Paper solves a complex portfolio selection problem with time-inconsistent preferences.
problem Time-inconsistent preferences in portfolio selection.
method Unified framework with minimal assumptions, proving existence and uniqueness of solution.
result Existence and uniqueness of square-integrable solution for the integral equation.
Study preferences over uncertain time payments, finds growth-optimality better than expected utility theory.
problem Understanding how people make decisions with uncertain timing of payments.
method Normative model of growth-optimality, revisiting experimental evidence on time lotteries.
result Growth-optimality better explains experimental data on time lotteries than expected discounted utility theory.
Efficiently learns reward functions with fewer queries and shorter computation times.
problem Expensive data generation and labeling in robot learning.
method Batch active preference-based learning methods using determinantal point processes (DPP) and heuristic alternatives.
result Our batch active learning algorithm requires only a few queries and computes them in a short amount of time.
This paper revisits optimal investment strategies for defined contribution pension schemes using forward preferences.
problem Optimal investment strategies derived from backward models are not time-consistent and sub-optimal in real scenarios.
method Introduces forward preferences and solves optimal investment strategies for defined contribution pension schemes.
result Constructs optimal investment strategies for defined contribution pension schemes using forward preferences.
The paper solves portfolio selection for complex preferences in continuous time.
problem Dynamic portfolio selection for nonlinear preferences with time inconsistency.
method Stochastic maximum principle and verification theorems for equilibrium strategies.
result Equilibrium strategies derived in closed form for CRRA and CARA preferences.
Investor optimizes portfolio under dynamic risk preferences.
problem Optimizing investment under uncertain future risk attitudes.
method Developed a general equilibrium framework and solved for subgame-perfect equilibrium policies.
result Equilibrium policies include a novel hedging component to counteract anticipated risk aversion changes.
Extends Kyle model to multiple traders with different time-preference coefficients.
problem Existence and convergence of discrete-time Kyle models with multiple insiders.
method Extends Basak and Cuoco's model to include traders with different time-preference coefficients.
result Parameter restrictions ensure the existence of a Radner equilibrium and long-term survival of traders.
Enhances robo-advisors with client investment preference inference.
problem Accurately inferring clients' investment preferences from past activities.
method Stochastic control framework with continuous-time model and discounting scheme.
result Proves sufficient conditions for client investment preference identifiability.
Bayesian optimization learns DM preferences for multi-outcome experiments.
problem Optimizing expensive experiments with unknown utility functions and multiple outcomes.
method Alternates preference learning and Bayesian optimization, using pairwise comparisons.
result Preference exploration strategies improve Bayesian optimization performance.
The paper addresses optimal control in modern tontines with bequest preferences, showing a linear investment strategy.
problem Optimal controls and decreasing allocation in modern tontines with bequest preferences.
method Dual approach to solve optimal control problems with power utilities, modeling bequest preferences.
result Investment strategy almost linearly adjusts from 0% to 100% over time.
A new framework enables real-time task trade-off control.
problem Conflict between multiple related tasks in a fixed model capacity.
method Formulates MTL as a preference-conditioned multiobjective optimization problem; uses a hypernetwork-based neural network.
result A single model can handle different trade-off preferences among multiple tasks.
We solve a continuous-time game-theoretic problem for Kihlstrom-Mirman preferences.
problem Dynamic inconsistency in preferences due to multiattribute utility theory.
method Formalized an equilibrium control theory for continuous-time Markov processes.
result Equilibrium strategy and value function as solution to extended HJB system.
The paper analyzes how mutable blockchain protocols affect miner behavior and strategic stability.
problem The mutability of blockchain protocols undermines long-term planning and cooperative equilibria.
method Integrates Austrian capital theory with repeated game theory to examine miner behavior under different institutional conditions.
result Effective time preference increases when protocol rules are mutable, leading to political rent-seeking and undermining strategic coherence.
SLHF uses sequential game theory to optimize preferences from human feedback.
problem Optimizing preferences from human feedback in sequential settings.
method SLHF frames the problem as a sequential-move game between Leader and Follower, decomposing the optimization into refinement and adversarial optimization.
result SLHF achieves strong alignment across diverse preference datasets and scales to large models.
Paper proposes a two-stage ranking for personalized TV recommendations.
problem Improving TV recommendation accuracy and efficiency.
method First, identifies potential candidates using user viewing patterns. Then, ranks them based on user preferences and program textual information.
result The proposed model outperforms in recommendation accuracy and efficiency.
Data generation and labeling are usually an expensive part of learning for robotics. While active learning methods are commonly used to tackle the former problem, preference-based learning is a concept that attempts to solve the latter by querying users with preference questions. In this paper, we will develop a new al…
Bayesian model identifies three types of travelers adapting to feedback.
problem Capturing adaptive, feedback-driven travel behavior in heterogeneous individuals.
method Latent Class Reinforcement Learning (LCRL) model with Variational Bayes estimation.
result Three distinct traveler classes identified: context-dependent, persistent exploitative, and exploratory.
In this paper, we investigate the impact of diverse user preference on learning under the stochastic multi-armed bandit (MAB) framework. We aim to show that when the user preferences are sufficiently diverse and each arm can be optimal for certain users, the O(log T) regret incurred by exploring the sub-optimal arms un…
A new method improves recommendation accuracy by learning from multiple networks and time-dependent user preferences.
problem Incomplete user profiles and dynamic user preferences degrade recommender quality.
method A cross-network time-aware recommender that learns from multiple source networks and develops current user models.
result The proposed solution achieves superior performance in accuracy, novelty, and diversity.
We consider black-box global optimization of time-consuming-to-evaluate functions on behalf of a decision-maker (DM) whose preferences must be learned. Each feasible design is associated with a time-consuming-to-evaluate vector of attributes and each vector of attributes is assigned a utility by the DM's utility functi…
The paper proposes a method to learn and leverage contextual preference distributions for better decision-making.
problem Heterogeneous and context-dependent human preferences in decision-making problems.
method A sequential learning-and-optimization pipeline using a bounded-variance score function gradient estimator to train a predictive model mapping contextual features to preference distributions.
result The approach reduces average post-decision surprise by up to 25 times compared to risk-averse baselines in a ridesharing environment.
Investigates time-inconsistent portfolio selection under MMV preferences.
problem Time-inconsistent optimal strategies for MMV preferences.
method Nash equilibrium controls for MMV and MV preferences, solving FBSDE and HJB equations.
result MMV optimal strategies lead to higher investment amounts than MV strategies, narrowing over time.
We demonstrate a limitation of discounted expected utility, a standard approach for representing the preference to risk when future cost is discounted. Specifically, we provide an example of the preference of a decision maker that appears to be rational but cannot be represented with any discounted expected utility. A …
This paper presents a non-trivial reconstruction of a previous joint topic-sentiment-preference review model TSPRA with stick-breaking representation under the framework of variational inference (VI) and stochastic variational inference (SVI). TSPRA is a Gibbs Sampling based model that solves topics, word sentiments an…
Study on investment strategy for agents with periodic preferences and discounting.
problem Investment decisions by agents with periodic S-shaped preferences and present bias.
method Infinite-horizon, continuous-time portfolio selection problem with quasi-hyperbolic discounting.
result Time-consistent planning strategy can be formulated as an equilibrium to a static mean field game.
Efficiently calculates PL model likelihood for partitioned preference data.
problem Computational infeasibility of calculating PL model likelihood for partitioned preference data.
method Random utility model formulation and efficient numerical integration approach.
result Proposed method outperforms existing LTR baselines and scales to real-world tasks.
Study on tracking preference shifts in dueling bandits problems.
problem Tracking significant preference shifts in dueling bandits problems.
method Analysis of dueling bandits with distribution shifts, focusing on significant shifts (Suk and Kpotufe, 2022).
result Design of adaptive algorithms with O ( K i l d e L T ) O(\sqrt{K ilde{L}T}) O ( K i l d e L T ) dynamic regret for certain preference distribution classes. New algorithm for learning preferences in decentralized matching markets reduces regret to logarithmic levels.
problem Learning preferences in decentralized matching markets without direct communication.
method Introduces a new algorithm for two-sided matching markets with competition.
result The algorithm achieves logarithmic stable regret in shared preferences and quadratic regret in general preferences.
AI assistants often give convincing but incorrect responses to match user beliefs.
problem Sycophancy in AI assistants that use human feedback.
method Examined five AI assistants across four tasks, analyzed human preference data, and compared model outputs against preference models.
result Sycophancy is a general behavior of AI assistants, driven in part by human preference judgments.
The paper explores how investors make decisions under disappointment aversion, finding that they prefer not to invest.
problem Continuous-time portfolio selection under generalized disappointment aversion.
method Sufficient and necessary condition for equilibrium strategies via fully nonlinear integral equation.
result Equilibrium strategy under disappointment aversion leads to less investment in the stock market compared to classical utility theory.
The paper analyzes optimal dividend and capital injection strategies under time-inconsistent preferences.
problem Optimal dividend and capital injection strategies under time-inconsistent preferences.
method Diffusion risk model with general discount functions, weak equilibrium definition, HJB equation system.
result Explicit solutions and threshold types of optimal strategies derived under different discount functions.
Preference are central to decision making by both machines and humans. Representing, learning, and reasoning with preferences is an important area of study both within computer science and across the sciences. When working with preferences it is necessary to understand and compute the distance between sets of objects, …
This paper solves optimal consumption-investment problems with time-varying preferences.
problem Optimal consumption-investment problems under time-varying incomplete preferences.
method Develops a martingale-type solution in a topological vector space, using stochastic processes and scalarization methods.
result Optimal investment policies are set-valued, with selectors decomposed into four components.
SARA uses similarity to learn rewards robustly and adaptively.
problem Robustness to labeler errors and adaptability to diverse feedback formats.
method Contrastive framework that learns latent representations and computes rewards as similarities.
result Strong performance on offline RL benchmarks and diverse applications.
Study on reinsurance decisions using mean-variance criterion with irreversible contracts.
problem Optimizing reinsurance premiums and contracts in a Stackelberg game with irreversible contracts.
method Unified singular control framework applied to both discrete and continuous time reinsurance contracts.
result A single once-for-all reinsurance contract is preferred over multiple contracts, and the signing time is crucial.
Study optimal healthcare spending under Epstein-Zin preferences for longevity.
problem Optimizing healthcare spending to extend longevity under Epstein-Zin preferences.
method Formulated Epstein-Zin utilities over a controllable random horizon using backward stochastic differential equations and HJB equations.
result Calibrated model accurately reflects actual mortality data and compares healthcare efficacy between countries.
Decentralized learning for matching markets with time-varying preferences.
problem Matching between competing agents and supply arms with time-varying preferences.
method Linear contextual bandit framework, learning algorithms to identify latent environment and stable matchings.
result Achieve instance-dependent logarithmic regret, applicable for large markets.
New MAB model incentivizes user arm-pulling with self-reinforcing preferences.
problem Balancing exploration and exploitation in recommender systems with incentivized user preferences.
method Proposes a new MAB model with random arm selection and two policies: At-Least- n n n Explore-Then-Commit and UCB-List. result Achieves O ( l o g T ) O(log T) O ( l o g T ) expected regret and O ( l o g T ) O(log T) O ( l o g T ) expected payment over a time horizon T T T . We analyze the problem of learning a single user's preferences in an active learning setting, sequentially and adaptively querying the user over a finite time horizon. Learning is conducted via choice-based queries, where the user selects her preferred option among a small subset of offered alternatives. These queries …
Algorithm identifies optimal stable matching in uncertain two-sided markets.
problem Sequential learning in two-sided markets with unknown preferences.
method Pure exploration approach with elimination-based algorithms exploiting partial preference information.
result Identification of pervasive stable matching for optimal stable matching identification.
Introduces SMMV preferences to avoid inconsistency in portfolio selection.
problem Monotone mean-variance preferences fail to differentiate strictly dominant payoffs.
method Introduces strictly monotone mean-variance preferences and applies them to portfolio selection problems.
result SMMV preferences provide a more rational basis for assessing prospects and coincide with MV preferences under certain conditions.
Paper shows equivalence between two dividend preference models.
problem Understanding investor and firm preferences for dividends.
method Formulated Epstein-Zin preference, proved equivalence with Maenhout's model.
result Robust dividend policy is equivalent to a threshold strategy based on surplus process.
This paper analyzes popular time-nonseparable utility functions that describe "habit formation" consumer preferences comparing current consumption with the time averaged past consumption of the same individual and "catching up with the Joneses" (CuJ) models comparing individual consumption with a cross-sectional averag…