Optimal insurance and investment strategy under exponential preferences in a correlated market model.
problem Optimal investment and reinsurance strategy for an insurance company under exponential preferences.
method Stochastic control techniques to construct a forward dynamic exponential utility and characterize the optimal strategy.
result Characterization of the optimal investment and reinsurance strategy in a correlated market model.
Optimal reinsurance strategy with fixed cost and exponential preferences.
problem Maximizing expected utility of terminal wealth with fixed reinsurance cost.
method Two-step procedure: stochastic control and optimal stopping problem.
result Deterministic optimal strategy depends on model parameters.
Proves existence of equilibrium in limited participation economy.
problem Existence of an equilibrium in an economy with limited financial market access.
method Proves global existence of Radner equilibrium using BSDEs with unique solution.
result Proves existence of Radner equilibrium with limited participation.
The paper analyzes optimal dividend and capital injection strategies under time-inconsistent preferences.
problem Optimal dividend and capital injection strategies under time-inconsistent preferences.
method Diffusion risk model with general discount functions, weak equilibrium definition, HJB equation system.
result Explicit solutions and threshold types of optimal strategies derived under different discount functions.
The paper extends static Systemic Risk Measures to a conditional setting.
problem Investigating how static Systemic Risk Measures can be adapted to a conditional framework.
method Providing a general dual representation result, analyzing Conditional Shortfall Systemic Risk Measures, and providing explicit formulas for exponential preferences.
result Explicit formulas for Conditional Shortfall Systemic Risk Measures and a time consistency property.
In a collectivised pension fund, investors agree that any money remaining in the fund when they die can be shared among the survivors. We give a numerical algorithm to compute the optimal investment-consumption strategy for an infinite collective of identical investors with exponential Kihlstrom--Mirman preferences, in…
Paper explores limits and possibilities of aligning LLMs with human preferences.
problem Aligning LLMs with diverse human preferences to ensure fairness and informed outcomes.
method Analysis of probabilistic representation of human preferences and preservation of diverse preferences.
result LLMs can't fully align with human preferences using reward-based approaches due to Condorcet cycles, but mixed strategies are statistically possible.
This paper revisits optimal investment strategies for defined contribution pension schemes using forward preferences.
problem Optimal investment strategies derived from backward models are not time-consistent and sub-optimal in real scenarios.
method Introduces forward preferences and solves optimal investment strategies for defined contribution pension schemes.
result Constructs optimal investment strategies for defined contribution pension schemes using forward preferences.
We tackle the problem of constructive preference elicitation, that is the problem of learning user preferences over very large decision problems, involving a combinatorial space of possible outcomes. In this setting, the suggested configuration is synthesized on-the-fly by solving a constrained optimization problem, wh…
The paper enhances preference learning by incorporating response time data.
problem Lack of temporal information in user decision-making for reward model learning.
method Integrates response time alongside binary choice data using the EZ model and Neyman-orthogonal loss functions.
result Response time-augmented approach reduces error rates from exponential to polynomial scaling, improving sample efficiency.
Optimal insurance policy for exponential utility maximization with convex premium calculation.
problem Maximizing terminal wealth utility with exponential utility function and convex premium formula.
method Necessary condition for optimal indemnity, numerical algorithm to compute it, convergence proof.
result Numerical algorithm converges to unique optimal indemnity.
Study on MMV in jump-diffusion models resolves MV's non-monotonicity issues.
problem Non-monotonicity and free cash flow stream problems in MV preferences.
method Explicit solution for MMV preferences in jump-diffusion models, proving non-negative potential measures.
result MMV resolves MV's non-monotonicity and free cash flow stream issues.
Unique optimal strategy identified for state-dependent risk aversion.
problem Consistency of optimal portfolio choice for varying risk aversion.
method Analysis of state-dependent exponential utilities in arbitrage-free markets.
result Uniqueness of optimal strategy across any time horizon.
This paper analyzes popular time-nonseparable utility functions that describe "habit formation" consumer preferences comparing current consumption with the time averaged past consumption of the same individual and "catching up with the Joneses" (CuJ) models comparing individual consumption with a cross-sectional averag…
This paper studies stability of the exponential utility maximization when there are small variations on agent's utility function. Two settings are considered. First, in a general semimartingale model where random endowments are present, a sequence of utilities defined on R converges to the exponential utility. Under a …
We analyze the generalized Mallows model, a popular exponential model over rankings. Estimating the central (or consensus) ranking from data is NP-hard. We obtain the following new results: (1) We show that search methods can estimate both the central ranking pi0 and the model parameters theta exactly. The search is n!…
In rank aggregation (RA), a collection of preferences from different users are summarized into a total order under the assumption of homogeneity of users. Model misspecification in RA arises since the homogeneity assumption fails to be satisfied in the complex real-world situation. Existing robust RAs usually resort to…
Improves RLHF sample efficiency by scaling reward complexity polynomially.
problem Exponential sample complexity in RLHF algorithms for skewed preferences.
method SE-POPO, an online RLHF algorithm that achieves polynomial sample complexity.
result SE-POPO outperforms existing algorithms in sample efficiency.
New method for efficient online exploration in RLHF reduces regret.
problem Efficiently collecting new preference data in RLHF to refine reward model and policy.
method Proposes a new exploration scheme that directs preference queries toward reducing uncertainty in reward differences most relevant to policy improvement.
result Establishes regret bounds of order T(β+1)/(β+2) for online RLHF, with polynomial scaling in all model parameters. Gradient descent on normalized networks reveals sparsity preferences.
problem Understanding the inductive bias of gradient descent on normalized neural nets.
method Analysis of gradient descent on weight-normalized smooth homogeneous neural nets, focusing on SWN and EWN.
result EWN causes weights to be updated in a way that prefers asymptotic relative sparsity.
We propose a mathematical framework for the study of a family of random fields--called forward performances--which arise as numerical representation of certain rational preference relations in mathematical finance. Their spatial structure corresponds to that of utility functions, while the temporal one reflects a Nisio…
Study growth rates of subgroups in groups with a constricting element.
problem Understanding growth rates of subgroups in groups with a constricting element.
method Examining the spectrum of relative and quotient exponential growth rates of quasi-convex subgroups.
result Determine when growth rates of subgroups are strictly smaller or coincide with the group's growth rate.
Study optimal investment and reinsurance for insurance companies in a dynamic market model.
problem Optimal investment and reinsurance strategies for insurance companies in a regime-switching market model.
method Forward dynamic exponential utility, value function construction, proportional reinsurance optimization.
result Characterization of optimal investment strategy and proportional reinsurance level.
New RL approach handles non-exponential discounting for sequential decisions.
problem Modeling human discounting in sequential decision-making tasks.
method Generalized model-based reinforcement learning with arbitrary discount functions, using Hamilton-Jacobi-Bellman equation and collocation method.
result Validated approach on simulated problems, showing applicability to human discounting.
Compound interest as well as inflation grows exponentially with time, whereas other means to repay debt grow polynomially. For this and other, mostly political, reasons, debt without inflation is unsustainable. We suggest a discontinuous way to eliminate debt by nullifying it. This scenario is preferable to current cen…
Paper uses deep learning for systemic risk measures.
problem Computing optimal capital allocations for systemic risk.
method Deep learning algorithms to solve primal and dual problems.
result Deep learning provides fair risk allocations.
We study an optimization problem for a portfolio with a risk-free, a liquid, and an illiquid risky asset. The illiquid risky asset is sold in an exogenous random moment with a prescribed liquidation time distribution. The investor prefers a negative or a positive exponential utility function. We prove that both cases a…
In this paper, we consider the problem of optimal investment by an insurer. The insurer invests in a market consisting of a bank account and m risky assets. The mean returns and volatilities of the risky assets depend nonlinearly on economic factors that are formulated as the solutions of general stochastic different…
We consider the problem of utility maximization with exponential preferences in a market where the traded stock/risky asset price is modelled as a Lévy-driven pure jump process (i.e. the driving Lévy process has no Brownian component). In this setting, we study the terminal utility optimization problem in the presence …
This paper studies the optimal risk-averse timing to sell a risky asset. The investor's risk preference is described by the exponential, power, or log utility. Two stochastic models are considered for the asset price -- the geometric Brownian motion and exponential Ornstein-Uhlenbeck models -- to account for, respectiv…
Mirror flow optimizes separable data problems, converging to a maximum margin classifier.
problem Optimizing classification problems with separable data using mirror flow.
method Examine mirror flow on linearly separable classification problems, focusing on the horizon function of the mirror potential.
result Mirror flow converges to a maximum margin classifier for separable data under certain conditions.
Dropout improves regularization in flexible models for rare features.
problem Understanding theoretical properties of dropout in generalized linear models.
method Theoretical analysis and application to adaptive smoothing with B-splines.
result Dropout prefers rare features in mean and dispersion parameters.
Study dynamic equilibrium with insider and general uninformed agent preferences.
problem Analyzing asymmetric information and general utility functions in a continuous-time economy.
method Introducing a new method to prove existence of a partial communication equilibrium (PCE) for agents with general utility functions.
result Identify the equilibrium price in the small and large risk aversion limits for agents with power utility.
This work extends implicit bias analysis to multiclass classification using a new loss framework.
problem The implicit bias of gradient descent on multiclass data without explicit regularization.
method Employing the PERM framework to introduce a multiclass extension of the exponential tail property.
result Extended implicit bias result to multiclass classification using a new loss framework.
This paper considers the Merton portfolio management problem. We are concerned with non-exponential discounting of time and this leads to time inconsistencies of the decision maker. Following Ekeland and Pirvu 2006, we introduce the notion of equilibrium policies and we characterize them by an integral equation. The ma…
NPO method improves LLM unlearning without catastrophic collapse.
problem Efficiently unlearning undesirable data from LLMs without losing model utility.
method Negative Preference Optimization (NPO) method based on alignment.
result NPO-based methods achieve better unlearning results and maintain model utility.
New decision-theoretic characterization separates belief and decision posteriors.
problem Understanding the conditions under which loss-based updating coincides with Bayesian updating.
method Decision-theoretic approach to distinguish belief and decision posteriors.
result Generalized Bayes coincides with ordinary Bayesian updating only if the loss is proportional to negative log-likelihood.
Improved sample efficiency in preference-based RL with multiple comparisons.
problem Sample inefficiency in preference-based reinforcement learning with pairwise comparisons.
method Proposes M-AUPO, an algorithm that selects multiple actions by maximizing average uncertainty within subsets.
result Achieves a suboptimality gap of $O\left( \frac{d}{T} \sqrt{ \sum_{t=1}^T \frac{1}{|S_t|}}
ight)$, improving performance with larger subsets.
In this paper, we propose an equilibrium pricing model in a dynamic multi-period stochastic framework with uncertain income streams. In an incomplete market, there exist two traded risky assets (e.g. stock/commodity and weather derivative) and a non-traded underlying (e.g. temperature). The risk preferences are of expo…
The paper studies automorphisms of free groups with a North-South dynamics.
problem Characterizing automorphisms of free groups with dynamical properties.
method Analyzes the action of outer automorphisms on a relative space of currents.
result Proves the existence of a preferred compact space on which automorphisms act with North-South dynamics.
This paper considers the optimal portfolio selection problem in a dynamic multi-period stochastic framework with regime switching. The risk preferences are of exponential (CARA) type with an absolute coefficient of risk aversion which changes with the regime. The market model is incomplete and there are two risky asset…
This paper extends the classical consumption and portfolio rules model in continuous time (Merton 1969, 1971) to the framework of decision-makers with time-inconsistent preferences. The model is solved for different utility functions for both, naive and sophisticated agents, and the results are compared. In order to so…
The Turaev-Viro invariants are a powerful family of topological invariants for distinguishing between different 3-manifolds. They are invaluable for mathematical software, but current algorithms to compute them require exponential time. The invariants are parameterised by an integer r≥3. We resolve the question …
Optimizes molecular generation for chemist preferences.
problem Models lack inherent preferences for chemist-desired structures.
method Fine-tuning with Direct Preference Optimization.
result Approach is simple, efficient, and highly effective.
Develops asset pricing models with mean field game theory for heterogeneous agents.
problem Tackles equilibrium asset pricing in incomplete markets with heterogeneous agents.
method Uses mean field game theory and mean field backward stochastic differential equations (BSDEs).
result Derives equilibrium risk premium and shows market clearing in the large population limit.
New method adapts to user preferences dynamically, improving recommendation models.
problem Current recommendation models lack dynamic adaptation to changing user preferences.
method Preference Discerning with LLM-Enhanced Generative Retrieval
result Mender achieves state-of-the-art performance in adapting to evolving user preferences.
This paper improves recommender systems by handling dynamic user preferences and item popularity.
problem Dynamic user preferences and changing item popularity in recommender systems.
method Developed a Thompson sampling-based policy for a high-dimensional linear bandit problem, reducing feature vector dimensionality and using exponentially increasing weights.
result Proved a regret bound that scales with the reduced dimension, demonstrating effectiveness in trade-off between computational complexity and regret performance.
Many real-world engineering problems rely on human preferences to guide their design and optimization. We present PrefOpt, an open source package to simplify sequential optimization tasks that incorporate human preference feedback. Our approach extends an existing latent variable model for binary preferences to allow f…