Bayesian approach selects subsets of variables for interpretable prediction and identifies key factors in educational outcomes.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
In this article, we propose a new algorithm for supervised learning methods, by which one can both capture the non-linearity in data and also find the best subset model. To produce an enhanced subset of the original variables, an ideal selection method should have the potential of adding a supplementary level of regres…
A/B testing improves marketing decisions by selecting effective stratification variables.
Bayesian method selects subsets for LMMs with structured dependence.
New sampler reduces MCMC complexity for Bayesian variable selection.
Algorithm selects variables and bandwidths for geographically weighted regression.
To find efficient screening methods for high dimensional linear regression models, this paper studies the relationship between model fitting and screening performance. Under a sparsity assumption, we show that a subset that includes the true submodel always yields smaller residual sum of squares (i.e., has better model…
A fast algorithm selects best subsets in high-dimensional models.
LESS combines local predictors for subsets to learn from heterogeneous input-output pairs.
The paper explores how agents can generalize to new environments with unseen variables.
Statistical detection of a rare class of objects in a two-class classification problem can pose several challenges. Because the class of interest is rare in the training data, there is relatively little information in the known class response labels for model building. At the same time the available explanatory variabl…
In life sciences, the experts generally use empirical knowledge to recode variables, choose interactions and perform selection by classical approach. The aim of this work is to perform automatic learning algorithm for variables selection which can lead to know if experts can be help in they decision or simply replaced …
Algorithm estimates common mean from Gaussian variables with unknown variances.
A novel non-supervised method detects anomalies in multivariate time series.
Mutual information has been successfully adopted in filter feature-selection methods to assess both the relevancy of a subset of features in predicting the target variable and the redundancy with respect to other variables. However, existing algorithms are mostly heuristic and do not offer any guarantee on the proposed…
Paper proposes a uniqueness Shapley measure to compare variable importance.
Subset selection for multiple linear regression aims to construct a regression model that minimizes errors by selecting a small number of explanatory variables. Once a model is built, various statistical tests and diagnostics are conducted to validate the model and to determine whether the regression assumptions are me…
Given a Gaussian Markov random field, we consider the problem of selecting a subset of variables to observe which minimizes the total expected squared prediction error of the unobserved variables. We first show that finding an exact solution is NP-hard even for a restricted class of Gaussian Markov random fields, calle…
Extends linear structural causal models to include deterministic relations and latent confounders for causal discovery.
Efficiently estimates variable importance in prediction tasks using Shapley values.
Designs efficient algorithms to maximize the expectation of Gaussian random variables.
Subset selection in multiple linear regression aims to choose a subset of candidate explanatory variables that tradeoff fitting error (explanatory power) and model complexity (number of variables selected). We build mathematical programming models for regression subset selection based on mean square and absolute errors…
Hotelling's -test for the mean of a multivariate normal distribution is one of the triumphs of classical multivariate analysis. It is uniformly most powerful among invariant tests, and admissible, proper Bayes, and locally and asymptotically minimax among all tests. Nonetheless, investigators often prefer non-inva…
Optimal kernel learning improves GP regression for high-dimensional inputs.
We consider the empirical risk minimization problem for linear supervised learning, with regularization by structured sparsity-inducing norms. These are defined as sums of Euclidean norms on certain subsets of variables, extending the usual -norm and the group -norm by allowing the subsets to overlap. T…
A new method selects features for clustering without labels.
We consider the solid angle that a planar compact subset subtends at a point in a level set of height h and study two extremal problems for the solid angle. One of the variables is a point in such a plane, that is, we study the properties of the solid angle maximizer. The other is the pair of a planar compact subset an…
Paper introduces DP methods for high-dimensional variable selection.
New method identifies causal variables from partially observed data.
New method falsifies causal discovery results without ground truth.
In many cases, feature selection is often more complicated than identifying a single subset of input variables that would together explain the output. There may be interactions that depend on contextual information, i.e., variables that reveal to be relevant only in some specific circumstances. In this setting, the con…
Methods of transfer learning try to combine knowledge from several related tasks (or domains) to improve performance on a test task. Inspired by causal methodology, we relax the usual covariate shift assumption and assume that it holds true for a subset of predictor variables: the conditional distribution of the target…
Abstract: Determines thermoelastic coefficients from boundary data.
This paper deals with prediction of anopheles number, the main vector of malaria risk, using environmental and climate variables. The variables selection is based on an automatic machine learning method using regression trees, and random forests combined with stratified two levels cross validation. The minimum threshol…
Paper generalizes bipolar theorems for non-negative random variables.
Fitting statistical models is computationally challenging when the sample size or the dimension of the dataset is huge. An attractive approach for down-scaling the problem size is to first partition the dataset into subsets and then fit using distributed algorithms. The dataset can be partitioned either horizontally (i…
In this paper we discuss the variable selection method from \ell0-norm constrained regression, which is equivalent to the problem of finding the best subset of a fixed size. Our study focuses on two aspects, consistency and computation. We prove that the sparse estimator from such a method can retain all of the importa…
Feature selection aims to select the smallest feature subset that yields the minimum generalization error. In the rich literature in feature selection, information theory-based approaches seek a subset of features such that the mutual information between the selected features and the class labels is maximized. Despite …
Proposes a few-shot learning method for feature selection without labeled data.
New algorithm groups variables by ancestral relationships to improve causal graph estimation accuracy.
This paper improves volatility forecasting using dynamic subset selection in genetic programming.
A new framework for time series analysis using state-space learning.
Delta-AI speeds up inference in sparse PGMs by local credit assignment.
We derive fundamental sample complexity bounds for recovering sparse and structured signals for linear and nonlinear observation models including sparse regression, group testing, multivariate regression and problems with missing features. In general, sparse signal processing problems can be characterized in terms of t…
We study the problem of learning a tree Ising model from samples such that subsequent predictions made using the model are accurate. The prediction task considered in this paper is that of predicting the values of a subset of variables given values of some other subset of variables. Virtually all previous work on graph…
Turbiner's conjecture posits that a Lie-algebraic Hamiltonian operator whose domain is a subset of the Euclidean plane admits a separation of variables. A proof of this conjecture is given in those cases where the generating Lie-algebra acts imprimitively. The general form of the conjecture is false. A counter-example …
abess efficiently solves various machine learning problems quickly.
Study predicts P2P lending platform failures using machine learning.