Choppy optimizes ranked list truncation using Transformer architecture.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New statistical models for predicting ranked preferences from partial orders.
We propose Top-N-Rank, a novel family of list-wise Learning-to-Rank models for reliably recommending the N top-ranked items. The proposed models optimize a variant of the widely used discounted cumulative gain (DCG) objective function which differs from DCG in two important aspects: (i) It limits the evaluation of DCG …
Learning the true ordering between objects by aggregating a set of expert opinion rank order lists is an important and ubiquitous problem in many applications ranging from social choice theory to natural language processing and search aggregation. We study the problem of unsupervised rank aggregation where no ground tr…
We study the problem of rank aggregation: given a set of ranked lists, we want to form a consensus ranking. Furthermore, we consider the case of extreme lists: i.e., only the rank of the best or worst elements are known. We impute missing ranks by the average value and generalise Spearman's ρto extreme ranks. Our main …
FUJI scores similarity of ranked lists more robustly.
Relevance ranking and result diversification are two core areas in modern recommender systems. Relevance ranking aims at building a ranked list sorted in decreasing order of item relevance, while result diversification focuses on generating a ranked list of items that covers a broad range of topics. In this paper, we s…
Proposes a method to handle sparse multiway count data with false zeros using zero-truncated Poisson regression.
Truncated Singular Value Decomposition (SVD) calculates the closest rank- approximation of a given input matrix. Selecting the appropriate rank defines a critical model order choice in most applications of SVD. To obtain a principled cut-off criterion for the spectrum, we convert the underlying optimization prob…
In this paper, we study the problem of safe online learning to re-rank, where user feedback is used to improve the quality of displayed lists. Learning to rank has traditionally been studied in two settings. In the offline setting, rankers are typically learned from relevance labels created by judges. This approach has…
Non-negative matrix factorization (NMF) minimizes the Euclidean distance between the data matrix and its low rank approximation, and it fails when applied to corrupted data because the loss function is sensitive to outliers. In this paper, we propose a Truncated CauchyNMF loss that handle outliers by truncating large e…
The conventional solution to the recommendation problem greedily ranks individual document candidates by prediction scores. However, this method fails to optimize the slate as a whole, and hence, often struggles to capture biases caused by the page layout and document interdepedencies. The slate recommendation problem …
Study ranking in generalized linear bandits with position and item dependencies.
Many web systems rank and present a list of items to users, from recommender systems to search and advertising. An important problem in practice is to evaluate new ranking policies offline and optimize them before they are deployed. We address this problem by proposing evaluation algorithms for estimating the expected …
Unsupervised ranking faces one critical challenge in evaluation applications, that is, no ground truth is available. When PageRank and its variants show a good solution in related subjects, they are applicable only for ranking from link-structure data. In this work, we focus on unsupervised ranking from multi-attribute…
We show that given an estimate that is close to a general high-rank positive semi-definite (PSD) matrix in spectral norm (i.e., ), the simple truncated SVD of produces a multiplicative approximation of in Frobenius norm. This observation leads to many inte…
Algorithm learns diverse rankings for search engines.
New algorithm predicts ranked stock lists for long-short portfolios.
Recent work has demonstrated the effectiveness of gradient descent for directly recovering the factors of low-rank matrices from random linear measurements in a globally convergent manner when initialized properly. However, the performance of existing algorithms is highly sensitive in the presence of outliers that may …
Efficiently estimate Boolean product distribution parameters from truncated samples.
List-wise learning to rank methods are considered to be the state-of-the-art. One of the major problems with these methods is that the ambiguous nature of relevance labels in learning to rank data is ignored. Ambiguity of relevance labels refers to the phenomenon that multiple documents may be assigned the same relevan…
The outcome of a functional genomics pipeline is usually a partial list of genomic features, ranked by their relevance in modelling biological phenotype in terms of a classification or regression model. Due to resampling protocols or just within a meta-analysis comparison, instead of one list it is often the case that …
New methods recover best rank-r approximations from few entries.
The paper tackles targeted attacks on rank aggregation methods, proving the fixed point of adversarial game.
We study the rank distribution, the cumulative probability, and the probability density of returns of stock prices of listed firms traded in four stock markets. We find that the rank distribution and the cumulative probability of stock prices traded in are consistent approximately with the Zipf's law or a power law. It…
Learning to rank is an important problem in machine learning and recommender systems. In a recommender system, a user is typically recommended a list of items. Since the user is unlikely to examine the entire recommended list, partial feedback arises naturally. At the same time, diverse recommendations are important be…
Paper proposes a new model for imputing missing spatiotemporal traffic data.
Many latent (factorized) models have been proposed for recommendation tasks like collaborative filtering and for ranking tasks like document or image retrieval and annotation. Common to all those methods is that during inference the items are scored independently by their similarity to the query in the latent embedding…
Study identifies and analyzes three types of errors in learning Fourier operators.
Recovering a large matrix from limited measurements is a challenging task arising in many real applications, such as image inpainting, compressive sensing and medical imaging, and this kind of problems are mostly formulated as low-rank matrix approximation problems. Due to the rank operator being non-convex and discont…
We show that an infinite family of contractible 4-manifolds have the same boundary as a special type of plumbing. Consequently their Ozsvath--Szabo invariants can be calculated algorithmically. We run this algorithm for the first few members of the family and list the resulting Heegaard--Floer homologies. We also show …
New model improves website ranking by considering user choices as a whole.
Smoothed analysis shows that many classes become learnable from positive-only samples.
In a stock market, the price fluctuations are interactive, that is, one listed company can influence others. In this paper, we seek to study the influence relationships among listed companies by constructing a directed network on the basis of Chinese stock market. This influence network shows distinct topological prope…
Paper proposes a new optimization framework for learning eigenfunctions of operators.
Due to the iterative nature of most nonnegative matrix factorization (\textsc{NMF}) algorithms, initialization is a key aspect as it significantly influences both the convergence and the final solution obtained. Many initialization schemes have been proposed for NMF, among which one of the most popular class of methods…
Sparse PCA is a widely used technique for high-dimensional data analysis. In this paper, we propose a new method called low-rank principal eigenmatrix analysis. Different from sparse PCA, the dominant eigenvectors are allowed to be dense but are assumed to have a low-rank structure when matricized appropriately. Such a…
New methods provide stable ranking without assumptions on data distributions.
We prove for a large class of knots that the meridional rank coincides with the bridge number. This class contains all knots whose exterior is a graph manifold. This gives a partial answer to a question of S. Cappell and J. Shaneson, see problem 1.11 on Kirby's list.
For many internet businesses, presenting a given list of items in an order that maximizes a certain metric of interest (e.g., click-through-rate, average engagement time etc.) is crucial. We approach the aforementioned task from a learning-to-rank perspective which reveals a new problem setup. In traditional learning-t…
By proving precisely which singularity index lists arise from the pair of invariant foliations for a pseudo-Anosov surface homeomorphism, Masur and Smillie determined a Teichmueller flow invariant stratification of the space of quadratic differentials. In this paper we determine an analog to the theorem for .…
The problem of searching for experts in a given academic field is hugely important in both industry and academia. We study exactly this issue with respect to a database of authors and their publications. The idea is to use Latent Semantic Indexing (LSI) and Latent Dirichlet Allocation (LDA) to perform topic modelling i…
Improves medication name inference for telemedicine and conversational agents.
Algorithm for low-rank matrix bandits with heavy-tailed rewards, achieving nearly optimal regret bound.
Recommender systems are widely used to recommend the most appealing items to users. These recommendations can be generated by applying collaborative filtering methods. The low-rank matrix completion method is the state-of-the-art collaborative filtering method. In this work, we show that the skewed distribution of rati…
The study assesses low-rank approximations in Gaussian Process regression.
The study assesses low-rank approximations in Gaussian Process regression.
Physics-inspired methods optimize SVD compression of LLMs.