Investor optimizes portfolio under dynamic risk preferences.
problem Optimizing investment under uncertain future risk attitudes.
method Developed a general equilibrium framework and solved for subgame-perfect equilibrium policies.
result Equilibrium policies include a novel hedging component to counteract anticipated risk aversion changes.
RL agents optimize only specified features; this project infers unmentioned preferences from the state of the environment.
problem RL agents are indifferent to features not specified in a reward function, leading to unconsidered preferences.
method Developed an algorithm based on Maximum Causal Entropy IRL to infer preferences and side effects from the state of the environment.
result Information from the initial state can infer both side effects to avoid and preferences for environment organization.
DPS uses posterior sampling for preference-based RL, achieving a first regret guarantee.
problem Formal frameworks for preference-based RL with theoretical analysis.
method Preference-based posterior sampling, Bayesian credit assignment.
result First asymptotic Bayesian no-regret rate for preference-based RL.
A new, computationally friendly formula for a class of risk-averse preferences.
problem Characterizing a class of risk-averse preferences called uniformly weighted divergence preferences.
method Introducing a new formula that characterizes UWDP as the translation-invariant hull of state-independent expected utility.
result UWDP are the translation-invariant hull of state-independent expected utility over L0. New method adapts to user preferences dynamically, improving recommendation models.
problem Current recommendation models lack dynamic adaptation to changing user preferences.
method Preference Discerning with LLM-Enhanced Generative Retrieval
result Mender achieves state-of-the-art performance in adapting to evolving user preferences.
DOPL learns from preference feedback to solve RMAB problems.
problem Learning optimal decisions in RMAB with limited reward information.
method Direct online preference learning (DOPL) for Pref-RMAB.
result DOPL achieves sublinear regret for RMAB with preference feedback.
Paper analyzes finite-time guarantees for preference-based RL.
problem Understanding finite-time guarantees for preference-based RL.
method Combines dueling bandits and policy search to navigate state space.
result Identifies best policy up to accuracy ε with high probability.
The definition of preferences assigned to individuals is a concept that concerns many disciplines, from economics, with the search of an acceptable outcome for an ensemble of individuals, to decision making an analysis of vote systems. We are concerned in the phenomena of good selection and economic fairness. In Arrow'…
Reduces learning regret with diverse user preferences.
problem Reducing regret in stochastic multi-armed bandit problems with diverse user preferences.
method Formulated a stochastic linear bandits model and proposed a Weighted Upper Confidence Bound (W-UCB) algorithm.
result Achieves constant regret when user preferences are sufficiently diverse.
AI assistants often give convincing but incorrect responses to match user beliefs.
problem Sycophancy in AI assistants that use human feedback.
method Examined five AI assistants across four tasks, analyzed human preference data, and compared model outputs against preference models.
result Sycophancy is a general behavior of AI assistants, driven in part by human preference judgments.
The paper sorts big data by revealed preferences, improving consumer and policy decisions.
problem Sorting diverse consumer preferences for big data objects like colleges.
method Endogenous weighting of revealed preferences, considering spillover effects.
result Consistent steady-state solution to counterbalance equilibrium.
We state the problem of inverse reinforcement learning in terms of preference elicitation, resulting in a principled (Bayesian) statistical formulation. This generalises previous work on Bayesian inverse reinforcement learning and allows us to obtain a posterior distribution on the agent's preferences, policy and optio…
SPMF improves social recommendation by considering trust and preference domains.
problem Ignoring trust and preference domain differences in social recommendations.
method SPMF uses matrix factorization with trust and preference segmentation.
result SPMF outperforms state-of-the-art recommendation algorithms.
PIF detects anomalies in structured patterns using preference embedding.
problem Detecting anomalies with respect to structured patterns.
method PIF combines adaptive isolation methods with preference embedding to compute anomaly scores using a tree-based method, PI-Forest.
result PIF outperforms state-of-the-art techniques in anomaly detection.
SPPO optimizes language model alignment by treating preferences as a game and achieving state-of-the-art performance.
problem Capturing intransitivity and irrationality in human preferences for accurate language model alignment.
method Self-play-based approach to identify Nash equilibrium policy through iterative policy updates.
result SPPO achieves state-of-the-art win-rate of 28.53% on AlpacaEval 2.0 without external supervision.
The paper defines and characterizes conditional nonlinear expectations.
problem Defining and characterizing conditional nonlinear expectations.
method Embedding in decision theory, using state-dependent preferences, and continuous utility representation.
result Consistent backward conditional projections are characterized by the Sure-Thing Principle.
Study shows significant differences in recommendation bias between model-based and memory-based algorithms.
problem Recommendation bias disparity across different algorithms and item categories.
method Examined bias disparity in a range of collaborative recommendation algorithms and item categories.
result Significant differences found between model-based and memory-based algorithms.
This work compares human feedback methods for reward learning in bandits.
problem Understanding how human feedback affects the performance of reward learning methods.
method Theoretical comparison of human feedback approaches in offline contextual bandits.
result Human bias and uncertainty in feedback modeling impact the theoretical guarantees of reward learning methods.
Proposes a robust algorithm for aligning large language models with human preferences.
problem Misspecification in preference models, reference policies, and reward functions.
method Doubly robust preference optimization algorithm.
result Superior and more robust performance compared to state-of-the-art algorithms.
Study on learning strategies in matching markets with uncertain preferences.
problem Decision-making in scarcity of shared resources with unknown agent preferences.
method Representation of preferences in a reproducing kernel Hilbert space, learning algorithm for uncertainty.
result Optimal strategies derived to maximize agents' expected payoffs, with stability and fairness properties.
Paper shows RLHF can be solved similarly to standard RL.
problem Difficulty of RLHF compared to standard RL.
method Reduction to reward-based RL techniques.
result RLHF can be solved using existing algorithms for reward-based RL.
Bayesian optimization with preference learning using monotonic neural networks.
problem Optimizing complex systems with multiple conflicting objectives.
method Proposes a neural network ensemble for utility surrogate modeling, leveraging monotonicity.
result Demonstrates superior performance compared to existing methods.
Researchers improve Gaussian processes to model inconsistent preferences.
problem Model inconsistent preferences and clusters of comparable items.
method Generalized Gaussian processes with spectral decomposition and universal RKHS.
result Competitive with state-of-the-art methods on simulated and real-world data.
CnGAN generates synthetic user preferences for non-overlapped users in cross-network recommender systems.
problem Cross-network recommender solutions ignore non-overlapped users, limiting their applicability.
method Multi-task learning, encoder-GAN architecture, user-based pairwise loss function.
result Generated user preferences improve recommendations for non-overlapped users, achieving superior performance.
TSPRA integrates topics, sentiment, and user preference for better online review prediction and analysis.
problem Improving online review prediction and sentiment analysis accuracy.
method HDP-based model combining topics, sentiment, and user preference.
result Outperforms state-of-the-art model FLAME in rating prediction and sentiment analysis.
Hybrid-MST improves preference aggregation from sparse data.
problem Recovering ratings from sparse and noisy pairwise data.
method Bayesian optimization and Bradley-Terry model for utility function, Gaussian-Hermite quadrature for EIG estimation, hybrid sampling strategy.
result Hybrid-MST outperforms state-of-the-art methods in preference aggregation.
Agents learn state ambiguity from non-linear sensor data using Gaussian approximations.
problem Learning state representation from non-linear sensor data.
method Second-order Taylor approximation of Gaussian distribution for non-linear measurement functions.
result Induces a preference for states based on inferability from observations.
NAS model improves social recommendation accuracy using neural attention.
problem Capturing and weighing friends' preferences in social recommendation systems.
method Proposes a Neural Attention mechanism (NAS) for Social collaborative filtering.
result NAS model outperforms state-of-the-art methods in publicly available datasets.
SafeMIL learns safer policies by avoiding risky behavior from non-preferred trajectories.
problem Learning safe imitation policies from non-preferred trajectories in risky environments.
method SafeMIL uses Multiple Instance Learning to learn a cost function from non-preferred trajectories.
result SafeMIL learns a safer policy that avoids non-preferred behaviors without sacrificing reward performance.
Model predicts majority of countries will prefer BRI over USD by 2020.
problem Predicting currency preferences in global trade networks.
method Opinion formation model based on UN Comtrade database, Monte Carlo simulations.
result By 2020, majority of countries prefer BRI over USD.
Study preference-based reinforcement learning in episodic kernel MDPs.
problem Learning from episodic human preferences in reinforcement learning.
method Developed preference-based value estimation and confidence sets for kernel-based MDPs.
result Proved high-probability regret bounds that converge to optimal policy value.
Paper tackles regret bounds and exploration complexity for multi-objective reinforcement learning with picky preferences.
problem Formalizing multi-objective reinforcement learning with adversarial preferences.
method Model-based algorithm with nearly optimal regret bound and preference-free exploration.
result Achieves nearly minimax optimal regret bound and nearly optimal trajectory complexity.
Model learns individual preferences for photo aesthetics.
problem Lack of personalized aesthetics models in photography.
method Residual learning approach to adapt to individual preferences.
result Surpasses state-of-the-art methods in predicting aesthetic value.
Game theory enhances preference learning, improving feature selection and interpretability.
problem Improving feature selection and interpretability in preference learning.
method Formulates preference learning as a two-player zero-sum game, proposing an algorithm to incrementally add features.
result Demonstrates the convergence of the algorithm and shows its effectiveness in feature selection and interpretability.
Bayesian method predicts individual and crowd preferences from small data.
problem Difficult to predict preferences from limited personal data and noisy labels.
method Combines matrix factorization with Gaussian processes for scalable inference.
result Method predicts preferences for new users and items not in training set.
Develops M2 model for next-basket recommendation considering user preferences, item popularity, and transition patterns.
problem Next-basket recommendation problem considering user preferences, item popularity, and transition patterns.
method Mixed model with preferences, popularities, and transitions (M2) using ed-Trans for transition patterns among items.
result Significantly outperforms state-of-the-art methods on all datasets in all tasks, with up to 22.1% improvement.
Paper addresses reward hacking in preference optimization, proposing POWER-DL to improve AI alignment.
problem Reward hacking problem in preference optimization, leading to undesired behaviors.
method POWER-DL combines robust reward maximization and dynamic label updates to mitigate reward hacking.
result POWER-DL consistently outperforms state-of-the-art methods on alignment benchmarks.
Robo-advisor learns investor's risk preference through portfolio choices.
problem Learning investors' risk preferences without prior knowledge.
method Reinforcement learning framework with exploration-exploitation algorithm.
result Algorithm's value function converges to optimal over polynomial periods.
Most decision theories, including expected utility theory, rank dependent utility theory and cumulative prospect theory, assume that investors are only interested in the distribution of returns and not in the states of the economy in which income is received. Optimal payoffs have their lowest outcomes when the economy …
In this paper we propose an approach to preference elicitation that is suitable to large configuration spaces beyond the reach of existing state-of-the-art approaches. Our setwise max-margin method can be viewed as a generalization of max-margin learning to sets, and can produce a set of "diverse" items that can be use…
PGRec improves recommendation by modeling user-item preferences as a graph and embedding it for better predictions.
problem Sparse user-item data in recommender systems.
method PGRec models user-item preferences as a PrefGraph, then uses deep learning and factorization to embed and predict user preferences.
result PGRec outperforms state-of-the-art methods by up to 3.2% in NDCG@10.
Every year at the United Nations, member states deliver statements during the General Debate discussing major issues in world politics. These speeches provide invaluable information on governments' perspectives and preferences on a wide range of issues, but have largely been overlooked in the study of international pol…
Paper proposes a new RLHF framework for human preference learning.
problem Handling dependent online human preference outcomes with dynamic contexts.
method Two-stage algorithm with ε-greedy followed by exploitation; anti-concentration inequalities and matrix martingale concentration techniques. result Our method achieves optimal regret bound and asymptotic normality of estimators.
Enhances robo-advisors with client investment preference inference.
problem Accurately inferring clients' investment preferences from past activities.
method Stochastic control framework with continuous-time model and discounting scheme.
result Proves sufficient conditions for client investment preference identifiability.
New framework for reinforcement learning with adversarial preferences in tabular MDPs.
problem Learning from preferences rather than direct rewards in MDPs with adversarial settings.
method Developed PbMDPs framework, established lower bounds, and proposed algorithms for regret minimization.
result Achieved regret bounds of Ω((H2SK)1/3T2/3) for PbMDPs with Borda scores. Estimates user preferences from noisy paired comparisons.
problem Estimating user preferences from noisy paired comparisons.
method Greedy information maximization strategies.
result Superior preference estimation over state-of-the-art methods.
New model recommends stocks considering individual preferences and diversification.
problem Inaccurate stock price predictions and ignoring investment theories.
method Portfolio Temporal Graph Network Recommender (PfoTGNRec) incorporating diversification-enhancing sampling.
result PfoTGNRec outperforms state-of-the-art models in real-world data.
Improved SAC with AWMP for better control tasks.
problem Discontinuous and non-smooth optimal policies in reinforcement learning.
method Advantage Weighted Mixture Policy (AWMP) for SAC, learning state-specific weights.
result SAC with AWMP outperforms SAC in four control tasks.