Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

5.0%10.0%15.0%20.0% · Jan 199519922001200920182026
48 results for relevance variable

New methods rank variables for Gaussian processes better than automatic relevance determination.

problem Variable selection for Gaussian process models using inverse length-scale parameters has limitations.
method Two novel methods rank variables based on their predictive relevance using posterior predictive distribution predictions.
result Improved variable selection compared to automatic relevance determination in terms of variability and predictive performance.

SADCBO optimizes contextual variables by balancing relevance and cost.

problem Optimizing contextual variables with varying costs and unknown relevance.
method Adaptive selection of relevant contextual variables using sensitivity analysis and early stopping.
result Consistent improvement in optimization across various examples.

Identifies feature relevance bounds for ordinal regression models.

problem Interpreting ordinal regression models is challenging due to variable dependencies.
method Identifies feature relevance bounds explicitly differentiating between strongly and weakly relevant features.
result Identification of feature relevance bounds for ordinal regression models.

A method to assess variable importance in complex predictive models.

problem Assessing the importance of variables in complex predictive models.
method Assigning relevance measures to each variable by comparing predictions with a ghost variable and analyzing joint effects.
result The method provides insights into variable importance and joint effects not available with other methods.

The multi-layer IB problem optimizes relevance and compression rates.

problem Optimizing relevance and compression rates in multi-layer information propagation.
method Single-letter characterization of the rate-relevance region, conditions for successive refinability, and counterexamples.
result Successive refinability of binary and Gaussian models, counterexample provided.

New local MDI variable importances derived from global scores match Shapley values.

problem Local feature relevance in tree-based models.
method Deriving local MDI importance measure from global scores and linking it to Shapley values.
result Local MDI importances have a natural connection with Shapley values.

Study identifies key ESG variables for assessing financial risk.

problem Assessing financial risk from ESG data with many variables.
method Proposed framework for hierarchical ESG data, selecting relevant variables.
result Selected ESG variables are more relevant to financial risk than aggregated scores.

XGBoost fails to accurately identify relevant features, while interpretable methods do.

problem Accurately identifying relevant features in black-box models like XGBoost.
method Comparison of variable importance methods (CART, Optimal Trees, XGBoost, SHAP) across various experiments.
result Interpretable methods outperform black-box models in feature selection accuracy.

Methodology to measure lag relevance in time series models.

problem Measuring lag relevance in machine learning models for univariate time series.
method Ghost variables, Shapley values, additive importance measures, auto-relevance and partial auto-relevance functions, one-step forecast.
result Calculated relevance measures successfully demonstrate expected lag structure in almost all cases.

New algorithm selects relevant variables in high-dimensional graphical models.

problem Automatic selection of relevant variables in high-dimensional graphical models.
method Extends Chow and Liu's algorithm using mutual information and entropy coefficient of determination.
result Outperforms existing methods in selecting variables with explanatory power.

Bayesian neural network improves feature selection and prediction.

problem Improving feature selection and prediction accuracy in neural networks.
method BNN-ARD with l2-norm feature importance measure.
result Improves variable selection and predictive performance on real-world data.

Paper proposes learning causal graphs with only relevant variables.

problem Discovering causal relationships in large-scale graphs often includes irrelevant variables.
method Developed NSCSL algorithm to learn necessary and sufficient causal graphs (NSCG).
result NSCSL algorithm identifies relevant causal features for specific outcomes.

TCMI assesses mutual dependence of continuous variables without parametric assumptions.

problem Estimating mutual information from continuous distributions.
method TCMI extends mutual information to continuous variables using cumulative distributions.
result TCMI facilitates feature selection and ranking of variable sets.

Sparse GEMINI selects relevant features for clustering without assumptions.

problem Feature selection in clustering with relevant clusters and variables.
method Discriminative clustering model maximizing GEMINI with l1 penalty.
result Sparse GEMINI selects relevant subsets of variables without prior hypotheses.

Tree-based method selects features from high-dimensional datasets with memory constraints.

problem Feature selection in high-dimensional datasets with limited memory.
method Randomized trees on subsamples of variables, mixing relevant and randomly selected variables.
result The method provides theoretical analysis and convergence speed under various scenarios.

The study finds that low frequency macroeconomic variables are more important for short-term electricity price forecasting.

problem Improving short-term forecasting of daily electricity prices using macroeconomic variables.
method Developed a Bayesian reverse unrestricted MIDAS model to account for frequency mismatch.
result Inclusion of macroeconomic low frequency variables improves short-term forecasts more than using only surveys or industrial production data.

A recurring problem when building probabilistic latent variable models is regularization and model selection, for instance, the choice of the dimensionality of the latent space. In the context of belief networks with latent variables, this problem has been adressed with Automatic Relevance Determination (ARD) employing…

2015-05-28abs ↗pdf ↗

Improving the detection of relevant variables using a new bivariate measure could importantly impact variable selection and large network inference methods. In this paper, we propose a new statistical coefficient that we call the rank minrelation coefficient. We define a minrelation of X to Y (or equivalently a majrela…

2013-05-09abs ↗pdf ↗

A fast and scalable method for variable selection in high-dimensional Gaussian processes.

problem Inefficient variable selection in high-dimensional Gaussian processes.
method Developed a fast and scalable variational inference algorithm for spike and slab Gaussian processes.
result Consistently outperforms vanilla and sparse variational GPs while retaining similar runtimes.

ML helps select variables for minimum-variance portfolios, reducing risk and improving performance.

problem Optimizing minimum-variance portfolios with relevant predictors.
method Parameterized minimum-variance portfolio weights using a large pool of firm-level characteristics and their transformations.
result ML-selected predictors lead to lower risk and better performance in minimum-variance portfolios.

Reduces selection bias in estimating individual treatment effects.

problem Selection bias in counterfactual reasoning.
method Auto-encoder with regularized loss based on Pearson Correlation Coefficient.
result Improves performance in estimating individual treatment effects.

New fuzzy clustering method for distribution-valued data using adaptive Wasserstein distances.

problem Clustering distribution-valued data with adaptive weights.
method Fuzzy c-means algorithms using adaptive L2L2 Wasserstein distances.
result Adaptive distances improve clustering of distribution-valued data.

FRI identifies relevant features in high-dimensional data for biomedical experiments.

problem Spurious feature selection in high-dimensional data.
method Feature relevance method for identifying all-relevant variables in linear classification and regression.
result FRI can identify causal features in biomedical experiments.

Proposes a few-shot learning method for feature selection without labeled data.

problem Feature selection in unlabeled data with limited instances.
method Uses Concrete random variables and permutation-invariant neural networks to select features from multiple source tasks.
result Outperforms existing methods in feature selection performance.

This paper examines variable selection for clustering using Gaussian mixture models.

problem Modern databases require efficient variable selection for clustering models.
method Recalls basics of clustering, examines variable selection methods for model-based clustering.
result Opportunities for improving variable selection methods are presented.

New method selects variables for GP regression using sparse projection.

problem Identifying environmental factors affecting metal corrosion.
method Sparse projection of input variables, gradient descent optimization, non-convex marginal likelihood.
result Proposed method outperforms benchmarks in variable selection accuracy.

We provide a classification of graphical models according to their representation as subfamilies of exponential families. Undirected graphical models with no hidden variables are linear exponential families (LEFs), directed acyclic graphical models and chain graphs with no hidden variables, including Bayesian networks …

2013-01-30abs ↗pdf ↗

Bayes-Factor-VAE models improve disentanglement of latent factors in data.

problem Disentangling latent factors in data using standard Gaussian priors is suboptimal.
method Introduced hierarchical Bayesian deep auto-encoder models with hyper-priors on latent variances.
result Bayes-Factor-VAEs outperform existing methods in latent disentanglement.

R package varrank ranks variables based on mutual information for multivariate data analysis.

problem Selecting and ranking variables for multivariate datasets.
method Minimum redundancy maximum relevance (mRMRe) model based on information theory.
result Flexible implementation for discrete and continuous data.