Paper extends top-k Mallows model for better user preference analysis.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper recovers top-two answers and confusion probability in multi-choice crowdsourcing.
The paper ranks items based on top choices in multiway comparisons.
Mixtures of ranking models have been widely used for heterogeneous preferences. However, learning a mixture model is highly nontrivial, especially when the dataset consists of partial orders. In such cases, the parameter of the model may not be even identifiable. In this paper, we focus on three popular structures of p…
Motivated by applications in recommender systems, web search, social choice and crowdsourcing, we consider the problem of identifying the set of top items from noisy pairwise comparisons. In our setting, we are non-actively given pairwise comparisons between each pair of items, where each comparison has noi…
The paper proposes a new method to learn choice functions using Pareto-embeddings.
Paper analyzes trade-offs in top-k classification accuracies and proposes a new loss function.
Improved theoretical guarantees for Top Two algorithms.
New statistical models for predicting ranked preferences from partial orders.
A new method for virtual drug screening detects top treatments.
Ranking data arises in a wide variety of application areas but remains difficult to model, learn from, and predict. Datasets often exhibit multimodality, intransitivity, or incomplete rankings---particularly when generated by humans---yet popular probabilistic models are often too rigid to capture such complexities. In…
The large asymptotics (perturbation series) for integrals of the form , where is a smooth top form and is a smooth function on a manifold , both of which are invariant under the action of a symmetry group , may be computed using the stationary phase approximation…
We consider combinatorial online learning with subset choices when only relative feedback information from subsets is available, instead of bandit or semi-bandit feedback which is absolute. Specifically, we study two regret minimisation problems over subsets of a finite ground set , with subset-wise relative prefe…
RCPO uses ranked choice modeling for better LLM alignment.
The paper provides robustness guarantees for mode estimation in bandits.
Real-valued word representations have transformed NLP applications; popular examples are word2vec and GloVe, recognized for their ability to capture linguistic regularities. In this paper, we demonstrate a {\em very simple}, and yet counter-intuitive, postprocessing technique -- eliminate the common mean vector and a f…
A widely applied diversification paradigm is the naive diversification choice heuristic. It stipulates that an economic agent allocates equal decision weights to given choice alternatives independent of their individual characteristics. This article provides mathematically and economically sound choice theoretic founda…
Study competition in OTC CDS market through CCP and interdealer choice models.
Ensemble methods are arguably the most trustworthy techniques for boosting the performance of machine learning models. Popular independent ensembles (IE) relying on naive averaging/voting scheme have been of typical choice for most applications involving deep neural networks, but they do not consider advanced collabora…
We define a hierarchy of special classes of constrained Willmore surfaces by means of the existence of a polynomial conserved quantity of some type, filtered by an integer. Type 1 with parallel top term characterises parallel mean curvature surfaces and, in codimension 1, type 1 characterises constant mean curvature su…
This work provides improved guarantees for streaming principle component analysis (PCA). Given sampled independently from distributions satisfying for , this work provides an -space linear-time single-pass streaming algorithm …
FPL allows users to control their data in federated top-N recommendation.
Paper shows faster core identification in matching markets.
Future predictions on sequence data (e.g., videos or audios) require the algorithms to capture non-Markovian and compositional properties of high-level semantics. Context-free grammars are natural choices to capture such properties, but traditional grammar parsers (e.g., Earley parser) only take symbolic sentences as i…
Within its traditional range of perversity parameters, intersection cohomology is a topological invariant of pseudomanifolds. This is no longer true once one allows superperversities, in which case intersection cohomology may depend on the choice of the stratification by which it is defined. Topological invariance also…
We introduce the probably approximately correct (PAC) \emph{Battling-Bandit} problem with the Plackett-Luce (PL) subset choice model--an online learning framework where at each trial the learner chooses a subset of arms from a fixed set of arms, and subsequently observes a stochastic feedback indicating prefere…
We study the problem of using low computational cost to automate the choices of learners and hyperparameters for an ad-hoc training dataset and error metric, by conducting trials of different configurations on the given training data. We investigate the joint impact of multiple factors on both trial cost and model erro…
Despite the recent success of deep transfer learning approaches in NLP, there is a lack of quantitative studies demonstrating the gains these models offer in low-shot text classification tasks over existing paradigms. Deep transfer learning approaches such as BERT and ULMFiT demonstrate that they can beat state-of-the-…
Ensembling DNNs improves minority group performance, leading to fairness.
Simple models are preferred over complex models, but over-simplistic models could lead to erroneous interpretations. The classical approach is to start with a simple model, whose shortcomings are assessed in residual-based model diagnostics. Eventually, one increases the complexity of this initial overly simple model a…
Nonlinear conjugate gradient (NLCG) based optimizers have shown superior loss convergence properties compared to gradient descent based optimizers for traditional optimization problems. However, in Deep Neural Network (DNN) training, the dominant optimization algorithm of choice is still Stochastic Gradient Descent (SG…
Fuzzy Forests reduces feature space in high-dimensional survey data.
We consider whether algorithmic choices in over-parameterized linear matrix factorization introduce implicit regularization. We focus on noiseless matrix sensing over rank- positive semi-definite (PSD) matrices in , with a sensing mechanism that satisfies restricted isometry properties (RIP)…
Top/O's first two k-invariants are zero.
Class ambiguity is typical in image classification problems with a large number of classes. When classes are difficult to discriminate, it makes sense to allow k guesses and evaluate classifiers based on the top-k error instead of the standard zero-one loss. We propose top-k multiclass SVM as a direct method to optimiz…
The top- error is often employed to evaluate performance for challenging classification tasks in computer vision as it is designed to compensate for ambiguity in ground truth labels. This practical success motivates our theoretical analysis of consistent top- classification. Surprisingly, it is not rigorously und…
We consider a notion of balanced metrics for triples (X,L,E) which depend on a parameter α, where X is smooth complex manifold with an ample line bundle L and E is a holomorphic vector bundle over X. For generic choice of α, we prove that the limit of a convergent sequence of balanced metrics leads to a Hermitian-Einst…
Paper introduces a new loss function for deep imbalanced classification.
Unified model for prediction and deferral selects top-k entities efficiently.
Work on making classifiers robust against adversarial attacks for top-k predictions.
Proposes top-label calibration and M2B framework for multiclass to binary calibration.
In order to push the performance on realistic computer vision tasks, the number of classes in modern benchmark datasets has significantly increased in recent years. This increase in the number of classes comes along with increased ambiguity between the class labels, raising the question if top-1 error is the right perf…
Financial time series forecasting is, without a doubt, the top choice of computational intelligence for finance researchers from both academia and financial industry due to its broad implementation areas and substantial impact. Machine Learning (ML) researchers came up with various models and a vast number of studies h…
Smoothed top-k operator improves model training efficiency.
In this paper, we introduce a geometric structure called top, which is a trivialized bundle of plane pencils over a Riemannian 3-manifold, defined as the set of kernels of a circle of 1-forms (e.g. of contact and integrable forms) with particular properties with respect to the metric. We classify the manifolds which ad…
New algorithm reduces sample complexity for Top Two method.
TyXe enables flexible Bayesian neural networks in Pytorch.
Paper introduces efficient top-k selection with differential privacy.