New models ensure monotonicity in preference learning, improving accuracy especially with limited data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper investigates monotonicity issues in AI preference learning.
Proposes a revenue function to evaluate dendrograms from comparisons.
There is increasing interest in learning algorithms that involve interaction between human and machine. Comparison-based queries are among the most natural ways to get feedback from humans. A challenge in designing comparison-based interactive learning algorithms is coping with noisy answers. The most common fix is to …
Traditionally, psychophysical experiments are conducted by repeated measurements on a few well-trained participants under well-controlled conditions, often resulting in, if done properly, high quality data. In recent years, however, crowdsourcing platforms are becoming increasingly popular means of data collection, mea…
We address the classical problem of hierarchical clustering, but in a framework where one does not have access to a representation of the objects or their pairwise similarities. Instead, we assume that only a set of comparisons between objects is available, that is, statements of the form "objects and are more …
Contextual policy search (CPS) is a class of multi-task reinforcement learning algorithms that is particularly useful for robotic applications. A recent state-of-the-art method is Contextual Covariance Matrix Adaptation Evolution Strategies (C-CMA-ES). It is based on the standard black-box optimization algorithm CMA-ES…
We consider the problem of classification in a comparison-based setting: given a set of objects, we only have access to triplet comparisons of the form "object is closer to object than to object ." In this paper we introduce TripletBoost, a new method that can learn a classifier just from such triplet …
We consider machine learning in a comparison-based setting where we are given a set of points in a metric space, but we have no access to the actual distances between the points. Instead, we can only ask an oracle whether the distance between two points and is smaller than the distance between the points an…
We consider the problem of finding a target object using pairwise comparisons, by asking an oracle questions of the form \emph{"Which object from the pair is more similar to ?"}. Objects live in a space of latent features, from which the oracle generates noisy answers. First, we consider the {\em non-bli…
Experiments used in current continual learning research do not faithfully assess fundamental challenges of learning continually. Instead of assessing performance on challenging and representative experiment designs, recent research has focused on increased dataset difficulty, while still using flawed experiment set-ups…
We consider the problem of search through comparisons, where a user is presented with two candidate objects and reveals which is closer to her intended target. We study adaptive strategies for finding the target, that require knowledge of rank relationships but not actual distances between objects. We propose a new str…
Deep neuroevolution, that is evolutionary policy search methods based on deep neural networks, have recently emerged as a competitor to deep reinforcement learning algorithms due to their better parallelization capabilities. However, these methods still suffer from a far worse sample efficiency. In this paper we invest…
The problem of content search through comparisons has recently received considerable attention. In short, a user searching for a target object navigates through a database in the following manner: the user is asked to select the object most similar to her target from a small list of objects. A new object list is then p…
In artificial intelligence, we often specify tasks through a reward function. While this works well in some settings, many tasks are hard to specify this way. In deep reinforcement learning, for example, directly specifying a reward as a function of a high-dimensional observation is challenging. Instead, we present an …
In the context of post-hoc interpretability, this paper addresses the task of explaining the prediction of a classifier, considering the case where no information is available, neither on the classifier itself, nor on the processed data (neither the training nor the test data). It proposes an instance-based approach wh…
New framework estimates treatment effects based on preferences.
Test log-likelihood comparisons can be misleading.
We propose a new yet natural algorithm for learning the graph structure of general discrete graphical models (a.k.a. Markov random fields) from samples. Our algorithm finds the neighborhood of a node by sequentially adding nodes that produce the largest reduction in empirical conditional entropy; it is greedy in the se…
Paper compares feature selection methods using GCM and LOCO, showing GCM methods generally outperform LOCO.
Assume we are given a set of items from a general metric space, but we neither have access to the representation of the data nor to the distances between data points. Instead, suppose that we can actively choose a triplet of items (A,B,C) and ask an oracle whether item A is closer to item B or to item C. In this paper,…
Paper optimizes PCA for fairness using MMD and Stiefel manifold optimization.
Paper tackles clustering with ordinal comparisons, achieving near-optimal results.
We aim at developing and improving the imbalanced business risk modeling via jointly using proper evaluation criteria, resampling, cross-validation, classifier regularization, and ensembling techniques. Area Under the Receiver Operating Characteristic Curve (AUC of ROC) is used for model comparison based on 10-fold cro…
Benchmark evaluates financial misinformation detection models, revealing weaknesses without external context.
We consider a novel setting of zeroth order non-convex optimization, where in addition to querying the function value at a given point, we can also duel two points and get the point with the larger function value. We refer to this setting as optimization with dueling-choice bandits since both direct queries and duels a…
Paper proposes a method to verify PINN fidelity using Fisher information from dynamical systems.
Physics-informed neural networks improve baryonic predictions from dark matter simulations.
Statistical framework improves LLM chatbot ranking.
This paper shows neural networks can solve complex graph problems efficiently.
New method improves ABC for Bayesian model comparison.
We consider the problem of approximate -means clustering with outliers and side information provided by same-cluster queries and possibly noisy answers. Our solution shows that, under some mild assumptions on the smallest cluster size, one can obtain an -approximation for the optimal potential with probabilit…
Meta-learning adapts models for unseen tasks across AI, robotics, and NLP.
Meta-learning improves neural networks by adapting learning algorithms.
Survey explores how transfer learning improves deep reinforcement learning.
Machine learning models adapt to motor learning but face challenges.
New method uses bi-level optimization to learn useful representations for imitation learning.
Study Whittle index learning algorithms for restless bandits with constant stepsizes.
AI learns to learn sequentially without forgetting.
Poisson learning doesn't solve graph semi-supervised learning issues.
New unsupervised learning technique learns independent kernels for better machine learning tasks.
Meta-learning helps models learn quickly from few samples.
New self-imitation learning method improves performance in continuous control tasks.
Deep reinforcement learning finds optimal learning policies for adaptive systems.
Unified framework explains all types of learning, including brain.
Cyclical learning rates improve DRL performance without manual tuning.
Study batch reinforcement learning methods for personalized medical treatments.
A new meta-meta classification method tackles few-shot learning tasks.