A new screening method for high-dimensional data reduces computational cost.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A new screening rule improves lasso solving speed.
RaSE screens variables via random subspaces, identifying joint effects.
New Bayesian optimization models for efficient material screening.
In data sets with many more features than observations, independent screening based on all univariate regression models leads to a computationally convenient variable selection method. Recent efforts have shown that in the case of generalized linear models, independent screening may suffice to capture all relevant feat…
This paper introduces LR-FFS for robust feature screening in federated learning under label shift.
A variable screening procedure via correlation learning was proposed Fan and Lv (2008) to reduce dimensionality in sparse ultra-high dimensional models. Even when the true model is linear, the marginal regression can be highly nonlinear. To address this issue, we further extend the correlation learning to marginal nonp…
New AI platform screens portfolios for desirable firms and news.
This paper treats the problem of screening for variables with high correlations in high dimensional data in which there can be many fewer samples than variables. We focus on threshold-based correlation screening methods for three related applications: screening for variables with large correlations within a single trea…
Recent computational strategies based on screening tests have been proposed to accelerate algorithms addressing penalized sparse regression problems such as the Lasso. Such approaches build upon the idea that it is worth dedicating some small computational effort to locate inactive atoms and remove them from the dictio…
New algorithms speed up learning from large screens of proteins.
Safe screening rules reduce -regression computation by fixing 76% of variables.
New rules reduce SLOPE model fitting time by screening out irrelevant variables.
New online feature selection method handles streaming data with concept drift.
This paper proposes a model-free and data-adaptive feature screening method for ultra-high dimensional datasets. The proposed method is based on the projection correlation which measures the dependence between two random vectors. This projection correlation based method does not require specifying a regression model an…
Variable screening is a fast dimension reduction technique for assisting high dimensional feature selection. As a preselection method, it selects a moderate size subset of candidate variables for further refining via feature selection to produce the final model. The performance of variable screening depends on both com…
Deep learning predicts breast cancer with high accuracy from patient data.
PS^2 selects assets then weights for high-dimensional investing.
We propose a flexible nonparametric regression method for ultrahigh-dimensional data. As a first step, we propose a fast screening method based on the favored smoothing bandwidth of the marginal local constant regression. Then, an iterative procedure is developed to recover both the important covariates and the regress…
GIDS reduces high-dimensional response and predictor spaces, improving interpretability and computational efficiency.
Understanding how features interact with each other is of paramount importance in many scientific discoveries and contemporary applications. Yet interaction identification becomes challenging even for a moderate number of covariates. In this paper, we suggest an efficient and flexible procedure, called the interaction …
New method detects biomarker-treatment interactions in clinical trials.
Sparse classifiers such as the support vector machines (SVM) are efficient in test-phases because the classifier is characterized only by a subset of the samples called support vectors (SVs), and the rest of the samples (non SVs) have no influence on the classification result. However, the advantage of the sparsity has…
Safe screening rules reduce computation time in logistic regression with regularization.
New method aggregates GDS analyses of randomly selected interaction models to identify important factors in screening experiments.
CAT framework improves AI medical screening fairness and reliability.
We present safe active incremental feature selection~(SAIF) to scale up the computation of LASSO solutions. SAIF does not require a solution from a heavier penalty parameter as in sequential screening or updating the full model for each iteration as in dynamic screening. Different from these existing screening methods,…
Paper develops a new algorithm to improve screening processes.
A new distributed method speeds up sparse model training.
Sparse learning techniques have been routinely used for feature selection as the resulting model usually has a small number of non-zero entries. Safe screening, which eliminates the features that are guaranteed to have zero coefficients for a certain value of the regularization parameter, is a technique for improving t…
We propose a novel transfer learning approach for orphan screening called corresponding projections. In orphan screening the learning task is to predict the binding affinities of compounds to an orphan protein, i.e., one for which no training data is available. The identification of compounds with high affinity is a ce…
We design simple screening tests to automatically discard data samples in empirical risk minimization without losing optimization guarantees. We derive loss functions that produce dual objectives with a sparse solution. We also show how to regularize convex losses to ensure such a dual sparsity-inducing property, and p…
The problem of learning a sparse model is conceptually interpreted as the process of identifying active features/samples and then optimizing the model over them. Recently introduced safe screening allows us to identify a part of non-active features/samples. So far, safe screening has been individually studied either fo…
Deep neural network identifies potential SARS-CoV-2 inhibitors.
Statistical inference can be computationally prohibitive in ultrahigh-dimensional linear models. Correlation-based variable screening, in which one leverages marginal correlations for removal of irrelevant variables from the model prior to statistical inference, can be used to overcome this challenge. Prior works on co…
New screening rules improve lasso model fitting efficiency.
A new screening rule 'dynamic Sasvi' improves sparse optimization speed.
Study improves reliability of neural models for virtual screening.
In the present paper, we introduce screen transversal lightlike submanifolds of metallic semi-Riemannian manifolds with its subclasses, namely screen transversal anti-invariant, radical screen transversal and isotropic screen transversal lightlike submanifolds, and give an example. We show that there do not exist co-is…
We study safe screening for metric learning. Distance metric learning can optimize a metric over a set of triplets, each one of which is defined by a pair of same class instances and an instance in a different class. However, the number of possible triplets is quite huge even for a small dataset. Our safe triplet scree…
In this paper we develop the notion of screen isoparametric hypersurface for null hypersurfaces of Robertson-Walker spacetimes. Using this formalism we derive Cartan identities for the screen principal curvatures of null screen hypersurfaces in Lorentzian space forms and provide a local characterization of such hypersu…
A new method for virtual drug screening detects top treatments.
Variable selection is a challenging issue in statistical applications when the number of predictors far exceeds the number of observations . In this ultra-high dimensional setting, the sure independence screening (SIS) procedure was introduced to significantly reduce the dimensionality by preserving the true mod…
Modern bio-technologies have produced a vast amount of high-throughput data with the number of predictors far greater than the sample size. In order to identify more novel biomarkers and understand biological mechanisms, it is vital to detect signals weakly associated with outcomes among ultrahigh-dimensional predictor…
Recently, to solve large-scale lasso and group lasso problems, screening rules have been developed, the goal of which is to reduce the problem size by efficiently discarding zero coefficients using simple rules independently of the others. However, screening for overlapping group lasso remains an open challenge because…
We propose {graphical sure screening}, or GRASS, a very simple and computationally-efficient screening procedure for recovering the structure of a Gaussian graphical model in the high-dimensional setting. The GRASS estimate of the conditional dependence graph is obtained by thresholding the elements of the sample covar…
DeepFS uses deep neural networks to select significant features in ultra high-dimensional data.
The main purpose of the present paper is to study the geometry of screen transversal lightlike submanifolds and radical screen transversal lightlike submanifolds and screen transversal anti-invariant lightlike submanifolds of Golden Semi-Riemannian manifolds. We investigate the geometry of distributions and obtain nece…