Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

84168251335 · Jun 202019922001200920172026
48 results for stochastic preferences

Optimal insurance and investment strategy under exponential preferences in a correlated market model.

problem Optimal investment and reinsurance strategy for an insurance company under exponential preferences.
method Stochastic control techniques to construct a forward dynamic exponential utility and characterize the optimal strategy.
result Characterization of the optimal investment and reinsurance strategy in a correlated market model.

The paper solves portfolio selection for complex preferences in continuous time.

problem Dynamic portfolio selection for nonlinear preferences with time inconsistency.
method Stochastic maximum principle and verification theorems for equilibrium strategies.
result Equilibrium strategies derived in closed form for CRRA and CARA preferences.

Study optimal healthcare spending under Epstein-Zin preferences for longevity.

problem Optimizing healthcare spending to extend longevity under Epstein-Zin preferences.
method Formulated Epstein-Zin utilities over a controllable random horizon using backward stochastic differential equations and HJB equations.
result Calibrated model accurately reflects actual mortality data and compares healthcare efficacy between countries.

Bal-PM reduces preference labeling costs for LLMs.

problem Efficiently acquiring human feedback for preference modeling in large language models.
method Bayesian Active Learning with entropy maximization in feature space.
result Bal-PM reduces the number of required preference labels by 33% to 68%.

In this paper, we investigate the impact of diverse user preference on learning under the stochastic multi-armed bandit (MAB) framework. We aim to show that when the user preferences are sufficiently diverse and each arm can be optimal for certain users, the O(log T) regret incurred by exploring the sub-optimal arms un…

2019-01-23abs ↗pdf ↗

New algorithm optimizes dueling bandits for both stochastic and adversarial preferences.

problem Optimizing decision-making in environments where only relative preferences are observed.
method Proposed a reduction from dueling bandits to multi-armed bandits, achieving optimal regret bounds.
result First best-of-both-world result for dueling bandits, optimal regret bound for Condorcet-winner benchmark.

This paper revisits optimal investment strategies for defined contribution pension schemes using forward preferences.

problem Optimal investment strategies derived from backward models are not time-consistent and sub-optimal in real scenarios.
method Introduces forward preferences and solves optimal investment strategies for defined contribution pension schemes.
result Constructs optimal investment strategies for defined contribution pension schemes using forward preferences.

The paper analyzes how behavioral investors make portfolio decisions using Markowitz Stochastic Dominance criteria.

problem Understanding how behavioral investors make portfolio decisions.
method Developed stochastic optimization problems and MILP models to capture subjective decision weights and probability weighting functions.
result The developed models can be used to formulate computationally tractable portfolio analysis problems.

Introduces SMMV preferences to avoid inconsistency in portfolio selection.

problem Monotone mean-variance preferences fail to differentiate strictly dominant payoffs.
method Introduces strictly monotone mean-variance preferences and applies them to portfolio selection problems.
result SMMV preferences provide a more rational basis for assessing prospects and coincide with MV preferences under certain conditions.

Enhances robo-advisors with client investment preference inference.

problem Accurately inferring clients' investment preferences from past activities.
method Stochastic control framework with continuous-time model and discounting scheme.
result Proves sufficient conditions for client investment preference identifiability.

New theory explains why normalization is preferred in SGD under heavy-tailed noise.

problem Understanding why normalization is preferred in stochastic gradient descent (SGD) under heavy-tailed noise.
method Developed a worst-case complexity theory for stochastically preconditioned SGD and its variants.
result Normalization guarantees convergence at optimal rates, while clipping may fail in the worst case.

Investigates optimal strategies for behavioral control problems with finite variation controls.

problem Behavioral singular stochastic control problems with finite variation controls.
method Abstract framework, applied to storage management and portfolio investment problems, using CPT preferences and Skorokhod representation theorem.
result Existence of optimal strategies for various goal functionals, including CPT preferences.

Study portfolio optimization with transaction costs and recursive preferences.

problem Optimizing portfolios under transaction costs and recursive preferences.
method Recursive preferences, transaction costs, and Merton investment-consumption problem.
result Characterized all parameter combinations for well-posedness of the problem.

Study recovers investor preferences from portfolio data using synthetic data and robust optimization.

problem Recovering latent investor preferences from observed portfolio allocations under uncertainty.
method Inverse portfolio optimization framework integrating robust optimization and regret-based inference.
result Accurate recovery of transaction cost parameters and partial identifiability of ESG penalties under preference misspecification and market shocks.

The paper develops a method to estimate consumer preferences from observed rankings.

problem Estimating consumer preferences from partial ranking information.
method Interpreting observed rankings as pairwise comparisons, modeling latent utility, and correcting for selection bias.
result The method improves recommendation performance, especially for previously unconsumed products.

New algorithm for learning preferences in decentralized matching markets reduces regret to logarithmic levels.

problem Learning preferences in decentralized matching markets without direct communication.
method Introduces a new algorithm for two-sided matching markets with competition.
result The algorithm achieves logarithmic stable regret in shared preferences and quadratic regret in general preferences.

Introduces RPU to explain randomization preference in dynamic settings.

problem Explains preference for randomization in dynamic investment problems.
method Introduces recursive perturbed utility (RPU) to incorporate randomization preference.
result Proves RPU-optimal portfolio policy is Gaussian and can be expressed in closed form.

Study forward investment performance in semimartingale markets with stochastic factors.

problem Investigate forward investment performance in incomplete semimartingale markets with power risk preferences and stochastic integrated factors.
method Develop necessary and sufficient conditions for FIPP existence, use integral representations, and solve ill-posed HJB equations.
result Explicit constructions for time-monotone FIPPs in semimartingale models, generalizing from Brownian to semimartingale markets.

Extends wealth tax neutrality framework to stochastic volatility and non-homothetic preferences.

problem Ensuring wealth taxes are neutral under various economic conditions.
method Extended Frøseth's neutrality framework to stochastic volatility and non-homothetic preferences, identified four channels of non-neutrality, and applied the framework to global minimum wealth taxes.
result Non-uniform assessment, general equilibrium effects, progressive thresholds, and endogenous labour supply can cause non-neutrality under CRRA preferences.

Paper develops methods to optimize policies directly from human feedback without reward inference.

problem Challenges in RLHF, including reward model overfitting and distribution shift.
method Develops two algorithms for RLHF without reward inference, using zeroth-order gradient approximators.
result Establishes polynomial convergence rates and outperforms existing methods in numerical experiments.

Paper shows equivalence between two dividend preference models.

problem Understanding investor and firm preferences for dividends.
method Formulated Epstein-Zin preference, proved equivalence with Maenhout's model.
result Robust dividend policy is equivalent to a threshold strategy based on surplus process.

Study optimal consumption and investment for investors with Epstein-Zin preferences.

problem Optimal consumption and investment for investors with Epstein-Zin preferences in an incomplete market.
method Variational characterisation and direct method to prove existence of optimal policies.
result Existence and uniqueness of optimal consumption and investment policies.

Study Epstein-Zin preferences in mean field portfolio games, proving unique equilibria.

problem Analyzing portfolio games with Epstein-Zin preferences under non-Markovian conditions.
method Proves a one-to-one correspondence between Nash equilibria and BSDE solutions, using local stochastic maximum principle tailored to Epstein-Zin utility.
result Establishes uniqueness of equilibria in mean field portfolio games under Epstein-Zin preferences.

This paper solves optimal consumption-investment problems with time-varying preferences.

problem Optimal consumption-investment problems under time-varying incomplete preferences.
method Develops a martingale-type solution in a topological vector space, using stochastic processes and scalarization methods.
result Optimal investment policies are set-valued, with selectors decomposed into four components.

We quantify content availability and user discovery opportunities in recommender systems.

problem Determining the maximum probability of recommending content to users.
method Stochastic reachability to compute upper bounds on recommendation likelihood.
result Reachability metrics can detect biases and diagnose user discovery limitations.

Stable matching, a classical model for two-sided markets, has long been studied with little consideration for how each side's preferences are learned. With the advent of massive online markets powered by data-driven matching platforms, it has become necessary to better understand the interplay between learning and mark…

2019-06-12abs ↗pdf ↗

Estimates users' preference for a site over others using engagement data.

problem Lack of data on users' interactions with other sites makes it hard to estimate preferences for a focal site.
method Uses Hierarchical Bayes Method with two estimation techniques: Markov Chain Monte Carlo and Stochastic Gradient with Langevin Dynamics.
result Good support found for the approach to computing personalized share of engagement.

Study asset pricing with reference-dependent preferences, finding matching equity premia.

problem Understanding asset pricing under reference-dependent preferences.
method Discrete-time consumption-based capital asset pricing model with reference-dependent preferences.
result Models can generate equity premia matching empirical estimates, showing procyclical price-dividend ratio and countercyclical equity premium.

We propose a scalable Bayesian preference learning method for jointly predicting the preferences of individuals as well as the consensus of a crowd from pairwise labels. Peoples' opinions often differ greatly, making it difficult to predict their preferences from small amounts of personal data. Individual biases also m…

2019-12-04abs ↗pdf ↗

Study optimal portfolios for many players in a market model with random coefficients.

problem Optimal portfolio selection for many players under relative performance criteria in a market model with random coefficients.
method Game theory and stochastic optimal control, focusing on CARA and CRRA risk preferences, and extending to continuum of players.
result Existence of forward Nash equilibrium and mean field equilibrium for the n-agent game and corresponding mean field stochastic optimal control problem.

This paper analyzes adaptive gradient algorithms for better performance in ill-conditioned problems.

problem Poor performance of standard stochastic gradient algorithms in ill-conditioned problems.
method Non-asymptotic analysis of adaptive gradient algorithms (Adagrad and Stochastic Newton) for strongly convex objectives.
result Theoretical analysis and adaptation to practical applications like linear regression and regularized GLM.

Investigates portfolio selection among competitive agents with mean-variance preferences.

problem Optimizing portfolios with multi-agent competition and relative wealth comparison.
method Reformulated as a constrained, non-homogeneous stochastic linear-quadratic control problem; derived optimal feedback strategies; used decoupling techniques and fixed-point theory to solve nonlinear BSDEs.
result Characterized three scenarios based on market and competition parameters: unique Nash equilibrium, no Nash equilibrium, or infinitely many Nash equilibria.