New score-based methods identify causal structures with latent variables.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Proposes a new method using Copula Entropy for variable selection.
The causal discovery of Bayesian networks is an active and important research area, and it is based upon searching the space of causal models for those which can best explain a pattern of probabilistic dependencies shown in the data. However, some of those dependencies are generated by causal structures involving varia…
A new variable importance measure for DRFs detects broader impacts on output distributions.
We present a new variable selection method based on model-based gradient boosting and randomly permuted variables. Model-based boosting is a tool to fit a statistical model while performing variable selection at the same time. A drawback of the fitting lies in the need of multiple model fits on slightly altered data (e…
New method models complex dynamics using a base variable.
This article describes the R package varrank. It has a flexible implementation of heuristic approaches which perform variable ranking based on mutual information. The package is particularly suitable for exploring multivariate datasets requiring a holistic analysis. The core functionality is a general implementation of…
This work explores variably scaled kernels to improve non-stationary Gaussian processes.
Proposes a new variable grouping approach to improve Bayesian additive regression tree (BART) performance.
We propose a technique for increasing the efficiency of gradient-based inference and learning in Bayesian networks with multiple layers of continuous latent vari- ables. We show that, in many cases, it is possible to express such models in an auxiliary form, where continuous latent variables are conditionally determini…
New method identifies latent variables in cognitive models using neural networks.
The study examines how market trade randomness influences price and return volatility.
The paper shows how neural networks with less decision boundary variability generalize better.
Proposes a gradient-based variable selection method for binary classification in RKHS.
A new method reduces CI tests for causal structure learning.
MVRSM optimizes expensive functions with mixed variables, outperforming state-of-the-art methods.
RI-based variable ranking and selection outperforms lasso in high-dimensional datasets.
Model-based clustering is a popular approach for clustering multivariate data which has seen applications in numerous fields. Nowadays, high-dimensional data are more and more common and the model-based clustering approach has adapted to deal with the increasing dimensionality. In particular, the development of variabl…
In this study, we propose an automatic learning method for variables selection based on Lasso in epidemiology context. One of the aim of this approach is to overcome the pretreatment of experts in medicine and epidemiology on collected data. These pretreatment consist in recoding some variables and to choose some inter…
Clustering is an essential technique for discovering patterns in data. The steady increase in amount and complexity of data over the years led to improvements and development of new clustering algorithms. However, algorithms that can cluster data with mixed variable types (continuous and categorical) remain limited, de…
This paper deals with prediction of anopheles number, the main vector of malaria risk, using environmental and climate variables. The variables selection is based on an automatic machine learning method using regression trees, and random forests combined with stratified two levels cross validation. The minimum threshol…
Variable selection for optimal treatment regime in a clinical trial or an observational study is getting more attention. Most existing variable selection techniques focused on selecting variables that are important for prediction, therefore some variables that are poor in prediction but are critical for decision-making…
Improved Shapley Value method for better model interpretation.
Hierarchical statistical models are widely employed in information science and data engineering. The models consist of two types of variables: observable variables that represent the given data and latent variables for the unobservable labels. An asymptotic analysis of the models plays an important role in evaluating t…
In this paper, we propose a variable selection method for general nonparametric kernel-based estimation. The proposed method consists of two-stage estimation: (1) construct a consistent estimator of the target function, (2) approximate the estimator using a few variables by l1-type penalized estimation. We see that the…
Unified framework for variable selection in model-based clustering with missing data.
Deep Bayesian neural networks effectively select variables with rigorous uncertainty quantification.
SADCBO optimizes contextual variables by balancing relevance and cost.
VC-PCR improves prediction by clustering correlated variables.
ABM automates feature engineering and variable selection for loss-based models.
New method reduces high-dimensional data to key features.
We propose a method that performs anomaly detection and localisation within heterogeneous data using a pairwise undirected mixed graphical model. The data are a mixture of categorical and quantitative variables, and the model is learned over a dataset that is supposed not to contain any anomaly. We then use the model o…
CARMS improves gradient estimation for categorical variables.
The mixture models have become widely used in clustering, given its probabilistic framework in which its based, however, for modern databases that are characterized by their large size, these models behave disappointingly in setting out the model, making essential the selection of relevant variables for this type of cl…
Variable selection for Gaussian process models is often done using automatic relevance determination, which uses the inverse length-scale parameter of each input variable as a proxy for variable relevance. This implicitly determined relevance has several drawbacks that prevent the selection of optimal input variables i…
In this paper, we propose multi-variable LSTM capable of accurate forecasting and variable importance interpretation for time series with exogenous variables. Current attention mechanism in recurrent neural networks mostly focuses on the temporal aspect of data and falls short of characterizing variable importance. To …
Develops a new method for decision trees using categorical variable structure.
Random forest hyperparameters affect variable selection in omics studies.
Two Bayesian optimization methods tackle dynamic design spaces with mixed variables.
CEBMs learn flexible latent mappings from data.
Study on NNs for forecasting time series with novel control variable combinations.
New gradient estimators for discrete variables improve model training.
Proposes SVI for covariate-shift generalization with sparse variable independence.
Statistical inference can be computationally prohibitive in ultrahigh-dimensional linear models. Correlation-based variable screening, in which one leverages marginal correlations for removal of irrelevant variables from the model prior to statistical inference, can be used to overcome this challenge. Prior works on co…
SCORE improves tree-based predictions with boosted residual extraTrees.
DualIV simplifies non-linear IV regression via dual formulation.
Proposes a non-convex optimization method for a parsimonious weighted naive Bayes classifier.
In this paper, we introduce Adaptive Cluster Lasso(ACL) method for variable selection in high dimensional sparse regression models with strongly correlated variables. To handle correlated variables, the concept of clustering or grouping variables and then pursuing model fitting is widely accepted. When the dimension is…