Paper extends RUMs with features to handle incomplete preferences and proves identifiability.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Preference learning (PL) is a core area of machine learning that handles datasets with ordinal relations. As the number of generated data of ordinal nature is increasing, the importance and role of the PL field becomes central within machine learning research and practice. This paper introduces an open source, scalable…
A large body of research is currently investigating on the connection between machine learning and game theory. In this work, game theory notions are injected into a preference learning framework. Specifically, a preference learning problem is seen as a two-players zero-sum game. An algorithm is proposed to incremental…
Adaptive reward models capture individual preferences from human feedback.
Modeling preference rankings with salient features to explain irrational choices.
We study the problem of ranking a set of items from nonactively chosen pairwise preferences where each item has feature information with it. We propose and characterize a very broad class of preference matrices giving rise to the Feature Low Rank (FLR) model, which subsumes several models ranging from the classic Bradl…
Recently, product images have gained increasing attention in clothing recommendation since the visual appearance of clothing products has a significant impact on consumers' decision. Most existing methods rely on conventional features to represent an image, such as the visual features extracted by convolutional neural …
Recommending new items to existing users has remained a challenging problem due to absence of user's past preferences for these items. The user personalized non-collaborative methods based on item features can be used to address this item cold-start problem. These methods rely on similarities between the target item an…
The paper proposes a method to learn and leverage contextual preference distributions for better decision-making.
This paper learns user preferences from comparisons using Mahalanobis metrics.
New method detects inconsistencies in AHP matrices using triadic preference reversals.
LLMs prefer Bitcoin under crisis frames, affecting financial decisions.
Bal-PM reduces preference labeling costs for LLMs.
MallowsPO enhances LLM fine-tuning with a dispersion index of human preferences.
New method proves fast regret bounds for online RLHF with generalized preferences.
Model learns metrics and preferences from user comparisons.
A neuroscience method to understanding the brain is to find and study the preferred stimuli that highly activate an individual cell or groups of cells. Recent advances in machine learning enable a family of methods to synthesize preferred stimuli that cause a neuron in an artificial or biological brain to fire strongly…
With the proliferation of social media platforms and e-commerce sites, several cross-domain collaborative filtering strategies have been recently introduced to transfer the knowledge of user preferences across domains. The main challenge of cross-domain recommendation is to weigh and learn users' different behaviors in…
Improved ranking method for scarce data with feature info.
We propose a set of conservative models in which agents exchange wealth with a preference in the choice of interacting agents in different ways. The common feature in all the models is that the temporary values of financial status of agents is a deciding factor for interaction. Other factors which may play important ro…
Investor optimizes portfolio under dynamic risk preferences.
Paper proposes using pairwise feature comparisons to infer modification costs for user recourse.
Paper proposes personalized climate control for driver comfort.
The aggregation of k-ary preferences is a historical and important problem, since it has many real-world applications, such as peer grading, presidential elections and restaurant ranking. Meanwhile, variants of Plackett-Luce model has been applied to aggregate k-ary preferences. However, there are two urgent issues sti…
Survival trees exhibit end-cut preference, leading to biased splits.
Reinforcement learning (RL) agents optimize only the features specified in a reward function and are indifferent to anything left out inadvertently. This means that we must not only specify what to do, but also the much larger space of what not to do. It is easy to forget these preferences, since these preferences are …
Proposes a new framework for resource-limited recommendation.
In this paper we propose an approach to preference elicitation that is suitable to large configuration spaces beyond the reach of existing state-of-the-art approaches. Our setwise max-margin method can be viewed as a generalization of max-margin learning to sets, and can produce a set of "diverse" items that can be use…
Bayesian method predicts individual and crowd preferences from small data.
We solve a continuous-time game-theoretic problem for Kihlstrom-Mirman preferences.
Optimizes explanations for better listener understanding.
Optimal insurance and investment strategy under exponential preferences in a correlated market model.
Recommender systems aim to find an accurate and efficient mapping from historic data of user-preferred items to a new item that is to be liked by a user. Towards this goal, energy-based sequence generative adversarial nets (EB-SeqGANs) are adopted for recommendation by learning a generative model for the time series of…
RGAM builds more accurate models by preferring linear features over non-linear ones.
TX-Ray analyzes and quantifies model knowledge transfer in NLP.
Traditional approaches to ranking in web search follow the paradigm of rank-by-score: a learned function gives each query-URL combination an absolute score and URLs are ranked according to this score. This paradigm ensures that if the score of one URL is better than another then one will always be ranked higher than th…
Generates low-dimensional node vectors for graphs with privacy while preserving structural preferences.
Stable matching, a classical model for two-sided markets, has long been studied with little consideration for how each side's preferences are learned. With the advent of massive online markets powered by data-driven matching platforms, it has become necessary to better understand the interplay between learning and mark…
Decentralized learning for matching markets with time-varying preferences.
Contextual bandit learning is an increasingly popular approach to optimizing recommender systems via user feedback, but can be slow to converge in practice due to the need for exploring a large feature space. In this paper, we propose a coarse-to-fine hierarchical approach for encoding prior knowledge that drastically …
Study optimal healthcare spending under Epstein-Zin preferences for longevity.
A fast method estimates stability of ensemble feature selectors.
Mitigates overoptimization in RLHF by reformulating SFT loss as a preference optimization loss.
We consider the task of collaborative preference completion: given a pool of items, a pool of users and a partially observed item-user rating matrix, the goal is to recover the \emph{personalized ranking} of each user over all of the items. Our approach is nonparametric: we assume that each item and each user h…
The first step in constructing a machine learning model is defining the features of the data set that can be used for optimal learning. In this work we discuss feature selection methods, which can be used to build better models, as well as achieve model interpretability. We applied these methods in the context of stres…
New MAB model incentivizes user arm-pulling with self-reinforcing preferences.
A generalized gamification framework is introduced as a form of smart infrastructure with potential to improve sustainability and energy efficiency by leveraging humans-in-the-loop strategy. The proposed framework enables a Human-Centric Cyber-Physical System using an interface to allow building managers to interact wi…
Recently, how to expand data transmission to reduce cell data and repeated cell transmission has received more and more research attention. In mobile social networks, content popularity prediction has always been an important part of traffic offloading and expanding data dissemination. However, current mainstream conte…