The article compares predictor importance in classification problems with categorical outcomes.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper proposes a sparse synthetic control method to select important predictors.
Efficient oblique RSF method improves prediction and interpretability.
Paper introduces new importance metrics for machine learning models, linking them to CATE.
Defines a new metric to measure importance of predictors in complex machine learning models.
We introduce a variable importance measure to quantify the impact of individual input variables to a black box function. Our measure is based on the Shapley value from cooperative game theory. Many measures of variable importance operate by changing some predictor values with others held fixed, potentially creating unl…
DynForest R package predicts outcomes with time-dependent predictors.
In the era of "big data", it is becoming more of a challenge to not only build state-of-the-art predictive models, but also gain an understanding of what's really going on in the data. For example, it is often of interest to know which, if any, of the predictors in a fitted model are relatively influential on the predi…
Improves transfer learning by weighting importance based on test-over-training density.
RI-based variable ranking and selection outperforms lasso in high-dimensional datasets.
We consider selection of random predictors for high-dimensional regression problem with binary response for a general loss function. Important special case is when the binary model is semiparametric and the response function is misspecified under parametric model fit. Selection for such a scenario aims at recovering th…
In this work, we study the problem of aggregating a finite number of predictors for nonstationary sub-linear processes. We provide oracle inequalities relying essentially on three ingredients: (1) a uniform bound of the norm of the time varying sub-linear coefficients, (2) a Lipschitz assumption on the predict…
Study uses deep learning to predict mycotoxin levels in Irish oats.
Group model selection is the problem of determining a small subset of groups of predictors (e.g., the expression data of genes) that are responsible for majority of the variation in a response variable (e.g., the malignancy of a tumor). This paper focuses on group model selection in high-dimensional linear models, in w…
Researchers predict butt rot volume using harvester data and remote sensing.
Global Sensitivity Analysis improves feature importance ranking in Random Forests.
Technology and collaboration enable dramatic increases in the size of psychological and psychiatric data collections, but finding structure in these large data sets with many collected variables is challenging. Decision tree ensembles like random forests (Strobl, Malley, and Tutz, 2009) are a useful tool for finding st…
Study examines challenges in variable importance ranking due to feature correlation.
This paper discusses a counterpart of conformal prediction for e-values, conformal e-prediction. Conformal e-prediction is conceptually simpler and had been developed in the 1990s as a precursor of conformal prediction. When conformal prediction emerged as result of replacing e-values by p-values, it seemed to have imp…
New method converts LVAs into linear projections for better understanding of complex models.
DynForest predicts event probabilities from longitudinal data, handling endogenous predictors.
We consider new formulations and methods for sparse quantile regression in the high-dimensional setting. Quantile regression plays an important role in many applications, including outlier-robust exploratory analysis in gene selection. In addition, the sparsity consideration in quantile regression enables the explorati…
Detection of protein-protein interactions (PPIs) plays a vital role in molecular biology. Particularly, infections are caused by the interactions of host and pathogen proteins. It is important to identify host-pathogen interactions (HPIs) to discover new drugs to counter infectious diseases. Conventional wet lab PPI pr…
Predict accuracy of neural architectures using non-neural models.
We analyze the local Rademacher complexity of empirical risk minimization (ERM)-based multi-label learning algorithms, and in doing so propose a new algorithm for multi-label learning. Rather than using the trace norm to regularize the multi-label predictor, we instead minimize the tail sum of the singular values of th…
DFNNs predict non-Euclidean responses from Euclidean predictors.
Study uses machine learning and survival analysis to predict CKD progression.
Despite their ability to memorize large datasets, deep neural networks often achieve good generalization performance. However, the differences between the learned solutions of networks which generalize and those which do not remain unclear. Additionally, the tuning properties of single directions (defined as the activa…
A novel feature selection method using noise-based hypothesis testing improves feature selection accuracy.
We give examples of data-generating models under which Breiman's random forest may be extremely slow to converge to the optimal predictor or even fail to be consistent. The evidence provided for these properties is based on mostly intuitive arguments, similar to those used earlier with simpler examples, and on numerica…
CDPs visualize causal dependencies in AI models.
Optical Wireless Communication (OWC) propagation channel characterization plays a key role on the design and performance analysis of Vehicular Visible Light Communication (VVLC) systems. Current OWC channel models based on deterministic and stochastic methods, fail to address mobility induced ambient light, optical tur…
New method learns to encode predictions within interpretations, improving evaluation.
Flexible framework for bounding high-loss predictions using quantiles.
Unified multitask learning framework for mixed-type outcomes.
Study identifies key trades predicting market movements.
Proposes using prior variable importance information in high-dimensional regression.
For nonlinear supervised learning models, assessing the importance of predictor variables or their interactions is not straightforward because it can vary in the domain of the variables. Importance can be assessed locally with sensitivity analysis using general methods that rely on the model's predictions or their deri…
New method minimizes regret in AMDP with high probability.
Electronic health records are an increasingly important resource for understanding the interactions between patient health, environment, and clinical decisions. In this paper we report an empirical study of predictive modeling of several patient outcomes using three state-of-the-art machine learning methods. Our primar…
Learning linear predictors with the logistic loss---both in stochastic and online settings---is a fundamental task in machine learning and statistics, with direct connections to classification and boosting. Existing "fast rates" for this setting exhibit exponential dependence on the predictor norm, and Hazan et al. (20…
Forecast dam inflow using sea surface feature weights.
New method corrects correlation bias in feature importance.
NACT improves tensor regression predictions with regularization.
As data collections become larger, exploratory regression analysis becomes more important but more challenging. When observations are hierarchically clustered the problem is even more challenging because model selection with mixed effect models can produce misleading results when nonlinear effects are not included into…
This paper tackles continuous covariate shift by adaptively training predictors.
Method improves treatment effect prediction robust to unknown covariate shifts.
New method for tensor classification with missing data.