The paper introduces a framework to select efficient datasets for preserving model rankings.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper tackles fair low-rank approximation and column subset selection.
Optimizes tensor rank selection for neural network compression.
MARS automatically selects tensor decomposition ranks, improving performance in neural network tasks.
The paper tackles learning true rankings from noisy, incomplete data.
New algorithms select and rank features from MTS without feature extraction.
This work analyzes tree-based methods from a ranking perspective, providing insights and new statistics.
We consider the problem of constructing a reduced-rank regression model whose coefficient parameter is represented as a singular value decomposition with sparse singular vectors. The traditional estimation procedure for the coefficient parameter often fails when the true rank of the parameter is high. To overcome this …
Selecting the right drugs for the right patients is a primary goal of precision medicine. In this manuscript, we consider the problem of cancer drug selection in a learning-to-rank framework. We have formulated the cancer drug selection problem as to accurately predicting 1). the ranking positions of sensitive drugs an…
Truncated Singular Value Decomposition (SVD) calculates the closest rank- approximation of a given input matrix. Selecting the appropriate rank defines a critical model order choice in most applications of SVD. To obtain a principled cut-off criterion for the spectrum, we convert the underlying optimization prob…
We study the model selection problem in conditional average treatment effect (CATE) prediction. Unlike previous works on this topic, we focus on preserving the rank order of the performance of candidate CATE predictors to enable accurate and stable model selection. To this end, we analyze the model performance ranking …
Unified framework for ranking-and-selection with multiple correct answers and non-answerable estimates
We develop an efficient algorithm for low-rank approximation with improved approximation guarantees.
This paper studies simultaneous feature selection and extraction in supervised and unsupervised learning. We propose and investigate selective reduced rank regression for constructing optimal explanatory factors from a parsimonious subset of input features. The proposed estimators enjoy sharp oracle inequalities, and w…
RI-based variable ranking and selection outperforms lasso in high-dimensional datasets.
Rank-based Bayesian Optimization improves molecule selection in chemical systems.
Investigates portfolio selection for rank-dependent utilities in incomplete markets.
This manuscript presents the following: (1) an improved version of the Binary Simultaneous Perturbation Stochastic Approximation (SPSA) Method for feature selection in machine learning (Aksakalli and Malekipirbazari, Pattern Recognition Letters, Vol. 75, 2016) based on non-monotone iteration gains computed via the Barz…
New unsupervised feature selection method for imbalanced datasets.
An ensemble technique is characterized by the mechanism that generates the components and by the mechanism that combines them. A common way to achieve the consensus is to enable each component to equally participate in the aggregation process. A problem with this approach is that poor components are likely to negativel…
This paper evaluates various loss functions for Transformer models in stock ranking.
The paper improves PCS approximation for ranking and selection under limited simulation budgets.
Myopic procedures are shown to be asymptotically optimal in ranking and selection problems.
Under a Bayesian framework, we formulate the fully sequential sampling and selection decision in statistical ranking and selection as a stochastic control problem, and derive the associated Bellman equation. Using value function approximation, we derive an approximately optimal allocation policy. We show that this poli…
The most popular approach for analyzing survival data is the Cox regression model. The Cox model may, however, be misspecified, and its proportionality assumption may not always be fulfilled. An alternative approach for survival prediction is random forests for survival outcomes. The standard split criterion for random…
Introduces greedy feature selection for classifier-dependent feature ranking.
New method selects features via tensor decomposition and submodular optimization.
Null-Calibrated Conformal Selection via Target-Membership Scores
New method compresses neural networks up to 14x with minimal performance loss.
Annealed Entropic Allocation improves ranking and selection by mitigating hard switching and improving finite-budget discrimination.
Feature selection methods are widely used in order to solve the 'curse of dimensionality' problem. Many proposed feature selection frameworks, treat all data points equally; neglecting their different representation power and importance. In this paper, we propose an unsupervised hypergraph feature selection method via …
We empirically test predictability on asset price by using stock selection rules based on maximum drawdown and its consecutive recovery. In various equity markets, monthly momentum- and weekly contrarian-style portfolios constructed from these alternative selection criteria are superior not only in forecasting directio…
New OCBA procedures minimize PICS in robust R&S.
In tensor completion tasks, the traditional low-rank tensor decomposition models suffer from the laborious model selection problem due to their high model sensitivity. In particular, for tensor ring (TR) decomposition, the number of model possibilities grows exponentially with the tensor order, which makes it rather ch…
Physics-inspired methods optimize SVD compression of LLMs.
Ranking data arises in a wide variety of application areas but remains difficult to model, learn from, and predict. Datasets often exhibit multimodality, intransitivity, or incomplete rankings---particularly when generated by humans---yet popular probabilistic models are often too rigid to capture such complexities. In…
Hutter (2007) recently introduced the loss rank principle (LoRP) as a generalpurpose principle for model selection. The LoRP enjoys many attractive properties and deserves further investigations. The LoRP has been well-studied for regression framework in Hutter and Tran (2010). In this paper, we study the LoRP for clas…
Develops efficient method for updating models with small data changes.
This paper examines the problem of ranking a collection of objects using pairwise comparisons (rankings of two objects). In general, the ranking of objects can be identified by standard sorting methods using pairwise comparisons. We are interested in natural situations in which relationships among the o…
In this paper, we propose new listwise learning-to-rank models that mitigate the shortcomings of existing ones. Existing listwise learning-to-rank models are generally derived from the classical Plackett-Luce model, which has three major limitations. (1) Its permutation probabilities overlook ties, i.e., a situation wh…
A new method ranks and selects features without model fitting.
Proposes a new allocation method for distributionally robust ranking and selection.
GeLoRA optimizes LoRA fine-tuning by dynamically adjusting ranks based on intrinsic dimensionality.
Graph Neural Networks (GNNs) have been a latest hot research topic in data science, due to the fact that they use the ubiquitous data structure graphs as the underlying elements for constructing and training neural networks. In a GNN, each node has numerous features associated with it. The entire task (for example, cla…
Meta-AAD uses deep reinforcement learning to improve anomaly detection by selecting the most informative instances.
The paper ranks items based on top choices in multiway comparisons.
Multivariate binary data is becoming abundant in current biological research. Logistic principal component analysis (PCA) is one of the commonly used tools to explore the relationships inside a multivariate binary data set by exploiting the underlying low rank structure. We re-expressed the logistic PCA model based on …
Prototype selection improved using topological data analysis.