Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

3570105140 · May 202619922001200920182026
48 results for relative preferences

Stable and consistent model alignment for language models without assuming human preference models.

problem Lack of statistical consistency in existing alignment methods.
method Relative density ratio optimization between preferred and mixture of preferred and non-preferred data distributions.
result Our approach achieves statistical consistency and stability, providing tighter convergence guarantees.

Investment and consumption strategies optimized with uncertain parameters.

problem Investment and consumption preferences in an incomplete financial market with uncertain parameters.
method PDE characterization and semi-explicit saddle-point construction of forward preferences and optimal strategies.
result A specific relationship between initial investment preference and forward consumption preference is necessary.

Study optimal portfolios for many players in a market model with random coefficients.

problem Optimal portfolio selection for many players under relative performance criteria in a market model with random coefficients.
method Game theory and stochastic optimal control, focusing on CARA and CRRA risk preferences, and extending to continuum of players.
result Existence of forward Nash equilibrium and mean field equilibrium for the n-agent game and corresponding mean field stochastic optimal control problem.

Proposes a method to select fair performance metrics through metric elicitation.

problem Choosing fair performance metrics in multiclass classification with multiple sensitive groups.
method Metric elicitation strategy that requires only relative preference feedback and is robust to noise.
result Elicits group-fair performance metrics for multiclass classification problems.

Study many-player investment-consumption games with power FPPs, finding market-risk preference affects consumption.

problem Investment and consumption optimization in a mean field competition setting.
method Solve many-player and mean field games using power FPPs, providing closed-form solutions.
result Market-risk relative consumption preference affects agent's consumption decisions.

Extends Kyle model to multiple traders with different time-preference coefficients.

problem Existence and convergence of discrete-time Kyle models with multiple insiders.
method Extends Basak and Cuoco's model to include traders with different time-preference coefficients.
result Parameter restrictions ensure the existence of a Radner equilibrium and long-term survival of traders.

Paper improves parameter estimation of continuous distributions using preference feedback.

problem Improving parameter estimation of continuous distributions.
method Preference-based M-estimators and deterministic preferences.
result Preference-based estimators achieve an estimation error scaling of O(1/n), significantly faster than sample-only methods.

DRO-REBEL improves LLM alignment by robustly updating models online.

problem Overfitting and drifting of LLMs during RLHF.
method DRO-REBEL uses type-pp Wasserstein, KL, and χ2χ^2 ambiguity sets for robust online updates.
result DRO-REBEL achieves faster convergence and better performance than prior methods.

The paper analyzes investment and consumption strategies under uncertain market conditions.

problem Investment and consumption under drift and volatility uncertainties.
method Randomization approach to construct robust preferences and strategies.
result Developed optimal and robust investment and consumption strategies remain valid in the physical market.

We introduce a number of new tools for the study of relatively hyperbolic groups. First, given a relatively hyperbolic group G, we construct a nice combinatorial Gromov hyperbolic model space acted on properly by G, which reflects the relative hyperbolicity of G in many natural ways. Second, we construct two useful bic…

2006-01-13abs ↗pdf ↗

New algorithms minimize regret in combinatorial online learning with relative feedback.

problem Minimizing regret in online learning with subset-wise relative preference feedback.
method Instance-dependent and order-optimal regret algorithms for two settings: bounded size subsets and fixed size subsets.
result Regret bounds of O(nmlnT)O(\frac{n}{m} \ln T) and O(nklnT)O(\frac{n}{k} \ln T) for respective settings.

Training models to prefer certain responses can unintentionally shift probability to harmful ones.

problem Likelihood displacement in DPO models, leading to unintended unalignment.
method Characterized and mitigated likelihood displacement using CHES score.
result Training models to prefer certain responses can unintentionally shift probability mass to harmful responses.

Optimizes portfolio growth rate for a behavioral investor considering terminal relative growth rate.

problem Optimizing a behavioral investor's portfolio growth rate under relative growth criterion.
method Martingale method, concavification, and quantile optimization techniques.
result Derives closed-form optimal growth rate and finds significant impact of benchmark growth rate.

This paper develops a new method for eliciting more flexible metrics, improving fairness and applicability.

problem Limited flexibility in existing metric elicitation strategies for reflecting user preferences.
method Develops a strategy for eliciting quadratic metrics based on predictive rates, requiring only relative preference feedback.
result Achieves near-optimal query complexity and broadens the use cases for metric elicitation.

Investor optimizes portfolio under dynamic risk preferences.

problem Optimizing investment under uncertain future risk attitudes.
method Developed a general equilibrium framework and solved for subgame-perfect equilibrium policies.
result Equilibrium policies include a novel hedging component to counteract anticipated risk aversion changes.

The paper sorts big data by revealed preferences, improving consumer and policy decisions.

problem Sorting diverse consumer preferences for big data objects like colleges.
method Endogenous weighting of revealed preferences, considering spillover effects.
result Consistent steady-state solution to counterbalance equilibrium.

Extended model ensures long-term survival of traders in limited stock market participation.

problem Limited stock market participation and survival of traders over long periods.
method Extended Basak and Cuoco (1998) model with different time-preference coefficients.
result Parameter restrictions ensure long-term survival of traders.

The paper explores how investors make decisions under disappointment aversion, finding that they prefer not to invest.

problem Continuous-time portfolio selection under generalized disappointment aversion.
method Sufficient and necessary condition for equilibrium strategies via fully nonlinear integral equation.
result Equilibrium strategy under disappointment aversion leads to less investment in the stock market compared to classical utility theory.

In recent years rank aggregation has received significant attention from the machine learning community. The goal of such a problem is to combine the (partially revealed) preferences over objects of a large population into a single, relatively consistent ordering of those objects. However, in many cases, we might not w…

2014-10-03abs ↗pdf ↗

We provide an axiomatic foundation for the representation of numéraire-invariant preferences of economic agents acting in a financial market. In a static environment, the simple axioms turn out to be equivalent to the following choice rule: the agent prefers one outcome over another if and only if the expected (under t…

2009-03-22abs ↗pdf ↗

Improved model for analyzing topics, sentiments, and user preferences in online reviews.

problem Inefficient processing of large-scale online review datasets.
method Developed variational inference models (vTSPRA, svTSPRA, ovTSPRA) for faster and more efficient processing of large datasets.
result The new models (svTSPRA, ovTSPRA) achieve better performance and faster convergence compared to the original TSPRA model.

Extends reinforcement learning alignment to scalar rewards, improving math reasoning.

problem Designing reinforcement learning algorithms for general LLM alignment.
method Introduces f-GRPO and f-HAL, estimating f-divergences between reward-aligned and unaligned distributions.
result Improves math reasoning RLVR tasks and mitigates reward hacking.

Study optimal investment decisions for diverse risk-tolerant agents.

problem Optimizing investment choices for agents with varying risk preferences.
method Characterizes optimal behavior using certainty equivalents and lognormal risks.
result Derives optimal decision menus under known and uncertain preference distributions.

A new method uses preference relations to reconcile contradictory trading signals from multiple securities.

problem Difficulty in exploiting multiple pairs trading signals due to contradictions.
method Proposes a portfolio construction method based on preference relation graphs to reconcile contradictory signals.
result Portfolios based on preference relations exhibit robust returns even with high transaction costs and improve with more securities considered.

Study shows LLM-advisors match human performance in eliciting preferences but struggle with conflicting needs and trust.

problem How do LLM-advisors perform in complex financial domains where domain expertise is crucial?
method Lab-based user study with 64 participants, focusing on three challenges: preference elicitation, personalized guidance, and relationship building.
result LLM-advisors can match human performance in preference elicitation but struggle with conflicting needs and trust issues.

The paper studies automorphisms of free groups with a North-South dynamics.

problem Characterizing automorphisms of free groups with dynamical properties.
method Analyzes the action of outer automorphisms on a relative space of currents.
result Proves the existence of a preferred compact space on which automorphisms act with North-South dynamics.

The algebra of transactions as fundamental measurements is constructed on the basis of the analysis of their properties and represents an expansion of the Boolean algebra. The notion of the generalized economic measurements of the economic quantity and quality of objects of transactions is introduced. It has been shown…

2014-12-18abs ↗pdf ↗

The paper develops a method to estimate consumer preferences from observed rankings.

problem Estimating consumer preferences from partial ranking information.
method Interpreting observed rankings as pairwise comparisons, modeling latent utility, and correcting for selection bias.
result The method improves recommendation performance, especially for previously unconsumed products.

Introduces RPU to explain randomization preference in dynamic settings.

problem Explains preference for randomization in dynamic investment problems.
method Introduces recursive perturbed utility (RPU) to incorporate randomization preference.
result Proves RPU-optimal portfolio policy is Gaussian and can be expressed in closed form.

A new algorithm for conversational recommendation systems using dueling bandits in GLMs.

problem Limited user feedback in existing conversational bandit methods.
method Integrates dueling bandits with relative feedback in generalized linear models.
result Theoretical and empirical validation of ConDuel's efficacy.

Optimizes a portfolio for an investor preferring accepted securities over a reference security.

problem Investor preference for a set of securities over a reference security with constraints.
method Mean-variance optimization with Sharpe Ratio performance measurement.
result Derives an optimal portfolio that maximizes returns while minimizing risk.

This paper explores the preference-based top-KK rank aggregation problem. Suppose that a collection of items is repeatedly compared in pairs, and one wishes to recover a consistent ordering that emphasizes the top-KK ranked items, based on partially revealed preferences. We focus on the Bradley-Terry-Luce (BTL) model…

2015-04-27abs ↗pdf ↗

ADPO optimizes relative advantage in reinforcement learning from human feedback.

problem Optimizing policy alignment in reinforcement learning from human preferences.
method ADPO explicitly parameterizes the optimal structure through anchored logits, decoupling response quality from prior popularity.
result Empirically, ADPO achieves state-of-the-art performance on reasoning tasks, outperforming GRPO by 30.9 percent.

New algorithm optimizes dueling bandits for both stochastic and adversarial preferences.

problem Optimizing decision-making in environments where only relative preferences are observed.
method Proposed a reduction from dueling bandits to multi-armed bandits, achieving optimal regret bounds.
result First best-of-both-world result for dueling bandits, optimal regret bound for Condorcet-winner benchmark.

The paper analyzes financial market equilibrium with heterogeneous risk preferences and convex constraints.

problem Characterizing equilibrium in a market with heterogeneous risk preferences and convex constraints.
method Continuous-time financial market model with heterogeneous agents and convex portfolio constraints.
result Margin constraints increase market price of risk and decrease interest rates, leading to higher equity risk premium and pro-cyclical leverage cycles.

Neural networks combining multiple data sources can reverse preferences, affecting decision reliability.

problem Preference reversals in neural networks under pooled data.
method Formalized through Case-Based Decision Theory, analyzed Gram geometry, introduced regularization, and developed auditing methods.
result Pooled refitting can reverse shared preferences, and conditions for preserving preferences are derived.

We study the top-KK ranking problem where the goal is to recover the set of top-KK ranked items out of a large collection of items based on partially revealed preferences. We consider an adversarial crowdsourced setting where there are two population sets, and pairwise comparison samples drawn from one of the populat…

2016-02-15abs ↗pdf ↗

We study the statistical regularities of opening call auction using the ultra-high-frequency data of 22 liquid stocks traded on the Shenzhen Stock Exchange in 2003. The distribution of the relative price, defined as the relative difference between the order price in opening call auction and the closing price of last tr…

2009-05-05abs ↗pdf ↗

This paper solves optimal consumption-investment problems with time-varying preferences.

problem Optimal consumption-investment problems under time-varying incomplete preferences.
method Develops a martingale-type solution in a topological vector space, using stochastic processes and scalarization methods.
result Optimal investment policies are set-valued, with selectors decomposed into four components.

Proposes a bond portfolio solution for managing interest rate risk.

problem Managing long-term assets and liabilities under interest rate risk.
method Proposes a bond portfolio solution based on ambiguity-averse preferences, accommodating various constraints and interest rate perturbations.
result Optimal portfolio can be computed as a simple generalized least squares problem, enhancing out-of-sample performance.

This paper improves fairness in recommendation systems by learning individual preferences across multiple dimensions.

problem Fairness in recommender systems, especially in areas with social impact.
method Opportunistic multi-aspect re-ranking approach that learns individual preferences and enhances provider fairness.
result Achieves a better trade-off between accuracy and fairness across multiple fairness dimensions.

Investment strategy in uncertain markets improved by learning and risk-ambiguity preferences.

problem Investment in financial markets with unknown drift coefficients.
method Optimization under KMM approach, considering risk and ambiguity preferences.
result Optimal investment strategy can be adjusted based on prior drift distribution.

Investigates optimal pension policies in PAYG systems with forward utility and ageing population.

problem Optimal investment and pension policies in PAYG systems with sustainability and adequacy constraints.
method Non-zero volatility forward CRRA utilities, closed-form optimal policies, detailed numerical analysis.
result Characterization of optimal policies and detailed impact analysis under various scenarios.