Study optimal healthcare spending under Epstein-Zin preferences for longevity.
problem Optimizing healthcare spending to extend longevity under Epstein-Zin preferences.
method Formulated Epstein-Zin utilities over a controllable random horizon using backward stochastic differential equations and HJB equations.
result Calibrated model accurately reflects actual mortality data and compares healthcare efficacy between countries.
Paper shows equivalence between two dividend preference models.
problem Understanding investor and firm preferences for dividends.
method Formulated Epstein-Zin preference, proved equivalence with Maenhout's model.
result Robust dividend policy is equivalent to a threshold strategy based on surplus process.
Study Epstein-Zin preferences in mean field portfolio games, proving unique equilibria.
problem Analyzing portfolio games with Epstein-Zin preferences under non-Markovian conditions.
method Proves a one-to-one correspondence between Nash equilibria and BSDE solutions, using local stochastic maximum principle tailored to Epstein-Zin utility.
result Establishes uniqueness of equilibria in mean field portfolio games under Epstein-Zin preferences.
Study optimal consumption and investment for investors with Epstein-Zin preferences.
problem Optimal consumption and investment for investors with Epstein-Zin preferences in an incomplete market.
method Variational characterisation and direct method to prove existence of optimal policies.
result Existence and uniqueness of optimal consumption and investment policies.
Investment and consumption strategy for risk-averse agents with Epstein-Zin utility.
problem Optimal investment and consumption strategy for Epstein-Zin utility.
method Detailed introduction to Epstein-Zin utility, existence and uniqueness proof, verification argument.
result Existence and uniqueness of optimal solution for Epstein-Zin utility under certain parameter restrictions.
Study optimal consumption and investment strategies with leverage constraints using Epstein-Zin utility.
problem Optimal portfolio choice under leverage constraints and Epstein-Zin utility.
method Established viscosity solution to HJB equation, demonstrated smoothness, characterized optimal strategies, derived explicit solutions.
result Explicit solutions for optimal consumption and investment strategies under leverage constraints.
Researchers analyze optimal investment strategies for a collectivised pension fund with identical investors.
problem Optimizing investment strategies for a collectivised pension fund with identical investors.
method Analytical computation of optimal investment-consumption strategies for a fund of n identical investors with Epstein-Zin preferences.
result Constant consumption strategy is suboptimal for infinite collectives, suggesting annuities and defined benefit investments are suboptimal.
Optimizes investment strategies for retirees with longevity risk.
problem Maximizing retirement savings under longevity risk for a group of investors.
method Analytic and numerical solutions for investment strategies in both discrete and continuous time models.
result Analytic formulae for optimal investment strategies in both discrete and continuous time models.
This paper solves the consumption-investment problem under Epstein-Zin preferences on a random horizon. In an incomplete market, we take the random horizon to be a stopping time adapted to the market filtration, generated by all observable, but not necessarily tradable, state processes. Contrary to prior studies, we do…
This paper solves optimal investment-consumption problems for a risk-averse agent with special utility.
problem Optimal investment-consumption problem for a risk-averse agent with special utility.
method Introduced proper utility process and solved optimal investment-consumption problem.
result Existence and uniqueness of proper utility processes for a wide class of consumption streams.
In a market with stochastic investment opportunities, we study an optimal consumption investment problem for an agent with recursive utility of Epstein-Zin type. Focusing on the empirically relevant specification where both risk aversion and elasticity of intertemporal substitution are in excess of one, we characterize…
Investigates stability of Epstein-Zin problem under market distortions.
problem Stability of Epstein-Zin problem in incomplete markets.
method Analyzes perturbations in returns and volatility, and interest rate; proves convergence of optimal solutions.
result Proves convergence of optimal consumption streams and value functions in the limit of model perturbations.
Study optimizes insurance and investment strategies for risk-averse insurers under ambiguity.
problem Optimizing insurance and investment strategies for risk-averse insurers under ambiguity.
method Solves a coupled FBSDE to derive optimal strategies and value function.
result Optimal consumption, investment, and reinsurance strategies influenced by risk aversion and EIS.
Investigates optimal consumption and investment strategies with constraints in incomplete markets.
problem Optimal consumption and investment under constraints in incomplete markets.
method Characterizes optimal strategies via a quadratic BSDE, using martingale optimality criterion and Lyapunov functions.
result Obtains the verification theorem for optimal strategies in unbounded cases.
Investigates optimal consumption and investment strategies in non-Markovian markets with unbounded parameters.
problem Optimal consumption and investment strategies in non-Markovian markets with unbounded parameters.
method Martingale optimal principle and quadratic BSDEs with exponential moment.
result Establishes optimal strategies for consumption and investment.
Study portfolio optimization with transaction costs and recursive preferences.
problem Optimizing portfolios under transaction costs and recursive preferences.
method Recursive preferences, transaction costs, and Merton investment-consumption problem.
result Characterized all parameter combinations for well-posedness of the problem.
Study many-player investment-consumption games with power FPPs, finding market-risk preference affects consumption.
problem Investment and consumption optimization in a mean field competition setting.
method Solve many-player and mean field games using power FPPs, providing closed-form solutions.
result Market-risk relative consumption preference affects agent's consumption decisions.
Extends wealth tax neutrality framework to stochastic volatility and non-homothetic preferences.
problem Ensuring wealth taxes are neutral under various economic conditions.
method Extended Frøseth's neutrality framework to stochastic volatility and non-homothetic preferences, identified four channels of non-neutrality, and applied the framework to global minimum wealth taxes.
result Non-uniform assessment, general equilibrium effects, progressive thresholds, and endogenous labour supply can cause non-neutrality under CRRA preferences.
Paper solves investment and consumption problem with unknown risk, providing explicit solutions.
problem Solving consumption-investment problem with unknown market price of risk and terminal liability constraint.
method Introduced a coupled forward-backward stochastic differential equation (FBSDE) and provided an explicit solution.
result Explicit expressions for optimal investment strategy and value function derived.
Investors adjust spending based on a social norm, spending less during losses and more during gains.
problem Managing spending and portfolio decisions while adhering to a social norm.
method Formulated a preference ordering with two CRRA preference orderings, solved analytically and numerically.
result Annual spending should be lower than expected financial return and procyclical, with spending cuts following losses.
This memoir presents a systematic study of the utility maximization problem of an investor in a constrained and unbounded financial market. Building upon the work of Hu et al. (2005) [Ann. Appl. Probab., 15, 1691--1712] in a bounded framework, we extend our analysis to the more challenging unbounded case. Our methodolo…
This paper introduces a dual problem to study a continuous-time consumption and investment problem with incomplete markets and stochastic differential utility. For Epstein-Zin utility, duality between the primal and dual problems is established. Consequently the optimal strategy of the consumption and investment proble…
CEFOL uses deep learning for dynamic programming with recursive utility.
problem Challenges in solving dynamic programming problems with recursive utility.
method Introduces a separate neural network for certainty equivalent, uses first-order optimality conditions to learn value and policy functions.
result CEFOL achieves high accuracy in learning value and policy functions, matching VFI benchmarks.
Optimizes molecular generation for chemist preferences.
problem Models lack inherent preferences for chemist-desired structures.
method Fine-tuning with Direct Preference Optimization.
result Approach is simple, efficient, and highly effective.
New method adapts to user preferences dynamically, improving recommendation models.
problem Current recommendation models lack dynamic adaptation to changing user preferences.
method Preference Discerning with LLM-Enhanced Generative Retrieval
result Mender achieves state-of-the-art performance in adapting to evolving user preferences.
Many real-world engineering problems rely on human preferences to guide their design and optimization. We present PrefOpt, an open source package to simplify sequential optimization tasks that incorporate human preference feedback. Our approach extends an existing latent variable model for binary preferences to allow f…
Enhances preference learning by incorporating response times into binary choices.
problem Limited information from binary choices about preference strength.
method Combines choices and response times using the EZ diffusion model.
result Response times improve utility estimation for strong preferences.
Bayesian optimization learns DM preferences for multi-outcome experiments.
problem Optimizing expensive experiments with unknown utility functions and multiple outcomes.
method Alternates preference learning and Bayesian optimization, using pairwise comparisons.
result Preference exploration strategies improve Bayesian optimization performance.
New study shows personalized content recommendations can lead to polarization of user preferences.
problem Personalized content recommendations can alter user preferences, leading to polarization.
method Used a model of preference dynamics to explore how personalized content affects user preferences.
result Standard reward maximization algorithms achieve only constant regret in personalized recommendation environments.
The standard asset pricing models (the CCAPM and the Epstein-Zin non-expected utility model) counterintuitively predict that equilibrium asset prices can rise if the representative agent's risk aversion increases. If the income effect, which implies enhanced saving as a result of an increase in risk aversion, dominates…
Bayesian optimization agent learns user preferences from pairwise comparisons.
problem Learning user preferences from unknown and infinite choices.
method Sequential Bayesian optimization with pairwise comparisons.
result Optimal agent strategy minimizes remaining system uncertainty.
New RLHF framework handles general preference oracles without reward functions.
problem Handling general preference oracles without assuming a reward function.
method Developed a minimax game between two LLMs for RLHF under a general preference oracle, focusing on KL-regularized preference.
result Proposed algorithms for efficient offline and online RLHF learning.
This paper studies robust forward investment and consumption preferences within a zero-volatility context. Different from previous works, we consider an incomplete financial market model due to general investment portfolio constraints. We provide a new PDE characterization and a novel semi-explicit saddle-point constru…
In preference-based reinforcement learning (RL), an agent interacts with the environment while receiving preferences instead of absolute feedback. While there is increasing research activity in preference-based RL, the design of formal frameworks that admit tractable theoretical analysis remains an open challenge. Buil…
Study on identifying most preferred policy in bandits with vector-valued rewards.
problem Identifying the most preferred policy in bandits with vector-valued rewards.
method Derive a novel lower bound on sample complexity, design the Preference-based Track and Stop (PreTS) algorithm, and derive a new concentration inequality.
result The sample complexity of PreTS is asymptotically tight.
Stable and consistent model alignment for language models without assuming human preference models.
problem Lack of statistical consistency in existing alignment methods.
method Relative density ratio optimization between preferred and mixture of preferred and non-preferred data distributions.
result Our approach achieves statistical consistency and stability, providing tighter convergence guarantees.
Dropping a tiny fraction of preferences can significantly alter the rankings of top LLMs.
problem Robustness of LLM ranking systems to small changes in preference data.
method A computational method based on the Bradley-Terry model to evaluate robustness.
result Top LLM rankings can be highly sensitive to the removal of a small fraction of preferences.
Paper explores limits and possibilities of aligning LLMs with human preferences.
problem Aligning LLMs with diverse human preferences to ensure fairness and informed outcomes.
method Analysis of probabilistic representation of human preferences and preservation of diverse preferences.
result LLMs can't fully align with human preferences using reward-based approaches due to Condorcet cycles, but mixed strategies are statistically possible.
Paper improves parameter estimation of continuous distributions using preference feedback.
problem Improving parameter estimation of continuous distributions.
method Preference-based M-estimators and deterministic preferences.
result Preference-based estimators achieve an estimation error scaling of O(1/n), significantly faster than sample-only methods.
Paper investigates monotonicity issues in AI preference learning.
problem AI models may violate monotonicity when learning preferences.
method Investigates root causes of non-monotonicity in comparison-based preference learning.
result Proves local pairwise monotonicity under mild assumptions.
DOPL learns from preference feedback to solve RMAB problems.
problem Learning optimal decisions in RMAB with limited reward information.
method Direct online preference learning (DOPL) for Pref-RMAB.
result DOPL achieves sublinear regret for RMAB with preference feedback.
Direct Density Ratio Optimization aligns LLMs with human preferences without assuming specific models.
problem Statistical inconsistency in aligning LLMs with human preferences.
method Direct Density Ratio Optimization (DDRO) estimates density ratio directly.
result DDRO is statistically consistent, converging to true human preferences as data grows.
Bayesian optimization with preference learning identifies preferred solutions in multi-objective problems.
problem Optimizing multiple criteria with decision maker preferences in expensive functions.
method Bayesian optimization with interactive preference learning and active acquisition function.
result Identifies the most preferred solution with reduced interaction cost.
Training models to prefer certain responses can unintentionally shift probability to harmful ones.
problem Likelihood displacement in DPO models, leading to unintended unalignment.
method Characterized and mitigated likelihood displacement using CHES score.
result Training models to prefer certain responses can unintentionally shift probability mass to harmful responses.
This work proves win rate is key to understanding preference learning.
problem Understanding preference learning from generative models.
method Analyzing preference learning methods as win rate optimization or non-WRO.
result Proves win rate is the only evaluation respecting preferences and prevalences.
IDT learns human preferences from uncertain decisions, even when humans are suboptimal.
problem Learning human preferences from uncertain and suboptimal decisions.
method Inverse decision theory (IDT) framework, statistical analysis of IDT, characterizing sample complexity.
result Learning preferences is easier when decisions are more uncertain, even if humans are suboptimal.
The paper proposes a method to infer multi-objective rewards from preferences.
problem Modeling preferences based on multiple, often competing objectives.
method Modeling priorities lexicographically and inferring multi-objective rewards from observed preferences.
result Lexicographically-ordered rewards provide a better understanding of preferences and improve policies.
The paper explores game-theoretic alignment of LLMs with human preferences, finding limitations and conditions.
problem Aligning LLMs with human preferences using game theory.
method Systematic study of payoff choices in a two-player zero-sum game for desirable alignment properties.
result Impossibility of preference matching in game-theoretic LLM alignment under standard assumptions.