Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

83166248331 · May 202619922001200920182026
48 results for constructive preference elicitation

We propose a decomposition technique to reduce user cognitive load in constructive preference elicitation.

problem Learning user preferences in large combinatorial decision problems.
method Part-wise inference and feedback over partial configurations.
result Significantly reduced user cognitive load and up to exponentially less computational demand.

This paper develops a new method for eliciting more flexible metrics, improving fairness and applicability.

problem Limited flexibility in existing metric elicitation strategies for reflecting user preferences.
method Develops a strategy for eliciting quadratic metrics based on predictive rates, requiring only relative preference feedback.
result Achieves near-optimal query complexity and broadens the use cases for metric elicitation.

In this paper we propose an approach to preference elicitation that is suitable to large configuration spaces beyond the reach of existing state-of-the-art approaches. Our setwise max-margin method can be viewed as a generalization of max-margin learning to sets, and can produce a set of "diverse" items that can be use…

2016-04-20abs ↗pdf ↗

Proposes a method to select fair performance metrics through metric elicitation.

problem Choosing fair performance metrics in multiclass classification with multiple sensitive groups.
method Metric elicitation strategy that requires only relative preference feedback and is robust to noise.
result Elicits group-fair performance metrics for multiclass classification problems.

Cost-effective framework for eliciting and aggregating preferences.

problem Eliciting preferences efficiently under budget constraints.
method Iterative computation of cost-effective questions using Plackett-Luce model and various information criteria.
result Carefully designed information criteria lead to more accurate predictions with fewer questions.

A framework for eliciting utility functions from investor preferences.

problem Hard elicitation of specific utility functions in portfolio selection.
method Preference-fitting method using probability-wealth pairs and PHARA approximation.
result Fitted utility function converges to the optimal one as more data is used.

Gradient optimization improves preference elicitation for large item spaces.

problem Computational infeasibility of EVOI for large item spaces in recommender systems.
method Continuous formulation of EVOI as a differentiable network, optimized using gradient methods.
result Gradient-based EVOI optimization achieves state-of-the-art performance and scalability.

This paper improves decision support in multi-objective planning by better eliciting user preferences.

problem Determining optimal policies from user preference profiles in multi-objective decision making.
method Extending Gaussian process and pairwise comparison methods to multi-objective scenarios, proposing new ordered preference elicitation strategies.
result Proposed elicitation strategies outperform existing methods and users prefer ranking.

We state the problem of inverse reinforcement learning in terms of preference elicitation, resulting in a principled (Bayesian) statistical formulation. This generalises previous work on Bayesian inverse reinforcement learning and allows us to obtain a posterior distribution on the agent's preferences, policy and optio…

2011-04-29abs ↗pdf ↗

Study shows LLM-advisors match human performance in eliciting preferences but struggle with conflicting needs and trust.

problem How do LLM-advisors perform in complex financial domains where domain expertise is crucial?
method Lab-based user study with 64 participants, focusing on three challenges: preference elicitation, personalized guidance, and relationship building.
result LLM-advisors can match human performance in preference elicitation but struggle with conflicting needs and trust issues.

Framework uses IRL and RL to elicit and optimize risk preferences robustly to noise.

problem Eliciting and optimizing risk preferences in noisy environments.
method Adaptive Bayesian IRL for elicitation, model-free RL for optimization, using quantile networks.
result Framework achieves convergence rate of O(exp(cm+O(mlogm)))O(\exp(-cm+O(\sqrt{m\log m}))).

Paper develops a robust preference model for multi-attribute choices.

problem Ambiguity in multi-attribute choice functions.
method Pairwise comparisons for preference elicitation, robust optimization model based on worst-case choice function.
result Developed tractable formulations for robust preference optimization.

Platform uses queries to elicit investor preferences for portfolio trades, improving allocation efficiency.

problem Hidden-information problem in institutional crossing markets where investors value trades as portfolios but liquidity discovery is organized by individual securities.
method Modeling portfolio crossing as preference elicitation, using price-directed demand queries and value queries to verify selected packages.
result Hybrid procedure using demand and value queries recovers 88-95% of full-information welfare with a limited query budget.

Elicit performance metrics from classifier comparisons.

problem Discover the performance metric a practitioner prefers for binary classification.
method Formalize and exploit geometric properties of confusion matrices for efficient metric elicitation.
result Provably efficient algorithms for eliciting linear and linear-fractional metrics from pairwise feedback.

Bayesian method helps decision-makers find preferred solutions in multi-objective optimization.

problem Identifying preferred solutions from the Pareto set in multi-objective optimization problems.
method Bayesian model to estimate decision-maker's utility function based on pairwise comparisons, guided by a principled elicitation strategy.
result Superior performance in finding high-utility solutions with a small number of queries.

GBS uses machine learning to design products based on consumer preferences.

problem Designing products to meet consumer preferences.
method GBS is a discrete choice experiment that uses machine learning to adaptively construct paired comparison questions.
result GBS outperforms existing methods in accuracy and sample efficiency.

Develops a statistical framework to measure uncertainty in model rankings based on human preferences.

problem Uncertainty in model rankings based on human preferences due to mismatch between human and model preferences.
method Statistical framework using pairwise comparisons by humans and models to provide rank-sets for each model.
result Rank-sets constructed using only pairwise comparisons by strong models often do not cover the true ranking of human preferences.

New algorithm speeds up user preference learning in conversational contexts.

problem Limited performance of existing conversational contextual bandit approaches.
method Proposes ConLinUCB framework and two algorithms, ConLinUCB-BS and ConLinUCB-MCR, with explorative key-term selection.
result Proves tighter regret bounds and achieves significant computational efficiency improvements.

Study on identifying most preferred policy in bandits with vector-valued rewards.

problem Identifying the most preferred policy in bandits with vector-valued rewards.
method Derive a novel lower bound on sample complexity, design the Preference-based Track and Stop (PreTS) algorithm, and derive a new concentration inequality.
result The sample complexity of PreTS is asymptotically tight.

Duel-Evolve uses LLM self-preferences for test-time optimization of discrete outputs.

problem Optimizing LLM outputs at test time with limited or unreliable scalar rewards.
method Duel-Evolve uses pairwise comparisons from the LLM to guide optimization, aggregating them via a Bayesian Bradley-Terry model.
result Achieves significant improvement over existing methods in accuracy.

This study measures price risk aversion using indirect utility functions in a lab experiment.

problem Measuring risk aversion with uncertain prices in experimental economics.
method Using indirect utility functions and a multiple price list method in a lab experiment.
result Price risk aversion is statistically greater than payoff risk aversion.

Develops methods for eliciting multiple, continuously valued treatment policies using causal inverse classification.

problem Tackles the problem of eliciting multiple, continuously valued treatment policies.
method Adopts a causal approach to inverse classification, developing the inverse classification potential outcomes framework (ICPOF) and approximate propensity score (APS).
result Demonstrates the viability of the methods on student performance.

The paper analyzes elicitability of return risk measures and their scoring functions.

problem Elicitability of return risk measures and their scoring functions.
method Dual representation results for convex and geometrically convex return risk measures, axiomatic characterizations of Orlicz premia, and construction of strictly consistent scoring functions.
result Orlicz premia are the only elicitable return risk measures under different sets of conditions.

The paper enhances preference learning by incorporating response time data.

problem Lack of temporal information in user decision-making for reward model learning.
method Integrates response time alongside binary choice data using the EZ model and Neyman-orthogonal loss functions.
result Response time-augmented approach reduces error rates from exponential to polynomial scaling, improving sample efficiency.

The paper tackles PAC ranking with subset-wise feedback, achieving optimal sample complexity.

problem Probably Approximately Correct (PAC) ranking of items with subset-wise preference feedback.
method Adaptive subset-wise preference feedback, Plackett-Luce model, pivot trick for score estimates.
result Achieves optimal sample complexity for PAC ranking with subset-wise feedback.

LLMs can be influenced by unseen dataset subtexts, revealing new ways to select data subsets.

problem Understanding how datasets subtly influence LLMs and their properties.
method Logit-Linear-Selection (LLS) method to select subsets of datasets.
result LLS reveals hidden effects in LLMs that persist across different models and architectures.

FedConPE improves conversational recommender systems efficiency and privacy.

problem Efficiently eliciting user preferences in interactive systems with heterogeneous clients.
method Phase elimination-based federated conversational bandit algorithm with adaptive key term construction.
result Minimizes uncertainty across all dimensions in feature space and offers improved efficiency and privacy.

ELECTRE Tree infers ELECTRE Tri-B parameters using a machine learning approach.

problem Infer ELECTRE Tri-B parameters from decision-maker inputs.
method Random Forest inspired algorithm: generate models, optimize parameters, merge or vote.
result ELECTRE Tree generates non-linear decision boundaries for voting, linear for merged model.

New algorithm learns human preferences from few comparisons efficiently.

problem Learning human preferences from limited comparison feedback.
method Formulated as D-optimal design for Plackett-Luce model, solved using randomized Frank-Wolfe algorithm.
result Proposed algorithm efficiently solves D-optimal design problem for Plackett-Luce objective.

A property, or statistical functional, is said to be elicitable if it minimizes expected loss for some loss function. The study of which properties are elicitable sheds light on the capabilities and limitations of point estimation and empirical risk minimization. While recent work asks which properties are elicitable, …

2015-06-23abs ↗pdf ↗

We discuss equivalent axiomatic characterizations of distortion risk measures, and give a novel and concise proof of the characterization of elicitable distortion risk measures. Elicitability has recently been discussed as a desirable criterion for risk measures, motivated by statistical considerations of forecasting. …

2014-05-15abs ↗pdf ↗

This paper tackles sample elicitation for learning systems, introducing a method to incentivize truthful samples.

problem Eliciting credible training samples for complex distributions from humans is challenging.
method Introduces a deep learning aided method to incentivize truthful samples from self-interested and rational agents.
result Achieves approximate incentive compatibility in eliciting truthful samples via accurate estimation of ff-divergence function.

Robustifies elicitable functionals to handle small distribution misspecifications.

problem Determining uniquely optimal forecasts under distributional misspecification.
method Integrates statistical robustness into elicitable functionals using Kullback-Leibler divergence.
result Robust elicitable functionals admit unique solutions at the boundary of uncertainty regions.

A statistical functional, such as the mean or the median, is called elicitable if there is a scoring function or loss function such that the correct forecast of the functional is the unique minimizer of the expected score. Such scoring functions are called strictly consistent for the functional. The elicitability of a …

2015-03-27abs ↗pdf ↗

Off-policy reinforcement learning has many applications including: learning from demonstration, learning multiple goal seeking policies in parallel, and representing predictive knowledge. Recently there has been an proliferation of new policy-evaluation algorithms that fill a longstanding algorithmic void in reinforcem…

2016-02-28abs ↗pdf ↗

Active Inverse Reward Design improves AI agent training by querying users for reward function preferences.

problem Iterative reward function tuning in AI agents is inefficient and may not generalize well.
method Structured queries to the user to compare reward functions, updating posterior with IRD.
result Substantially outperforms IRD in test environments, inferring non-linear rewards.