Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

90181271361 · Jun 202019922001200920172026
48 results for covariate importance

PAMA learns covariate importance for better matching in observational studies.

problem Poor performance of conventional matching methods when covariates differ in relevance.
method PAMA is a semi-supervised framework that learns covariate importance from paired data and optimizes a weighted quadratic score.
result PAMA outperforms standard methods, particularly in high-dimensional settings and under model misspecification.

This paper improves error estimation in covariate shift by incorporating target information.

problem Error estimation is inaccurate in covariate shift scenarios.
method Proposes a redefinition of importance using target information for better error estimation.
result Incorporating target information leads to more accurate error estimation, especially with KLIEP.

The paper proves a new method to improve generalization in covariate-shift scenarios.

problem Improving performance on test distributions that differ from training distributions.
method Independence-driven importance weighting algorithms for feature selection.
result Theoretical proof that these algorithms can identify optimal variables for covariate-shift generalization.

A new method identifies class-specific covariates in multi-class prediction tasks.

problem Identifying covariates specifically associated with one or more outcome classes in multi-class prediction tasks.
method Introducing multi forests (MuFs) with multi-way and binary splits to measure class-associated discriminatory ability.
result The multi-class VIM specifically ranks class-associated covariates highly, unlike conventional VIMs.

The paper proposes AIS for Bayesian inversion of multioutput signals with covariance estimation.

problem Performing uncertainty analysis of covariance matrices in Bayesian inversion problems for multioutput signals.
method Adaptive Importance Sampling (AIS) scheme, split variables, frequentist approach for noise covariance, prior density over covariance matrix.
result Estimation of model parameters and covariance matrix of noise.

Algorithm calibrates predictions for covariate shift using domain adaptation.

problem Uncertainty estimates overestimate certainty when real-world data differs from training data.
method Uses importance weighting and learns a feature map to equalize distributions.
result Outperforms existing approaches in calibrated prediction when covariate shift occurs.

Proposes a new method to estimate variable importance in black box models, mitigating correlation effects.

problem Correlation between covariates affects the interpretation of variable importance parameters.
method Develops a modified LOCO (Leave Out COvariates) method and uses semiparametric models for estimation.
result Shows how to estimate a modified LOCO method that mitigates correlation effects.

Paper tackles unbounded density ratio estimation for covariate shift adaptation.

problem Understudied challenge in statistical learning: unbounded density ratios.
method Three-step estimation method: relative density ratio, truncation, and transformation.
result Established rigorous convergence guarantees for density ratio and regression estimators.

Proposes a robust method for predicting missing outcomes in covariate shift adaptation.

problem Predicting missing outcomes in test data with covariate shift.
method Doubly robust estimator for covariate shift adaptation via importance weighting, incorporating an additional estimator for the regression function.
result Shows robustness against density-ratio estimation errors, maintaining consistency if either estimator is consistent.

The paper examines when importance weighting is needed for nonparametric and misspecified models.

problem When is importance weighting correction needed for covariate shift adaptation?
method Analysis of IW-corrected kernel ridge regression in various settings.
result The importance weighting correction is needed for nonparametric and misspecified models to obtain the best approximation of the true unknown function.

Constructs covariant derivatives for Ehresmann connections.

problem Developing a method for covariant derivatives in fibre bundles.
method Introducing a vertical endomorphism to construct covariant derivatives on vertical and horizontal distributions.
result Covariant derivatives can be constructed separately on vertical and horizontal distributions and then glued together.

The paper introduces Shapley curves for measuring variable importance in nonparametric settings.

problem Limited statistical understanding of Shapley values as variable importance measures.
method Introduces Shapley curves based on conditional expectation and covariate distribution; derives convergence rates and normality; proposes a novel bootstrap procedure.
result Validates theoretical findings with numerical studies and analyzes vehicle prices determinants.

A new one-step method for covariate shift adaptation.

problem Real-world data often violates the assumption of same distribution for training and test samples.
method Proposes a one-step optimization approach to jointly learn the model and weights.
result The proposed method achieves a generalization error bound and is empirically effective.

The paper addresses portfolio allocation with uncertain covariance matrices, finding a logarithmic risk dependence.

problem Portfolio allocation with uncertain covariance matrices.
method Calculates the expected value of CARA utility function over a distribution of covariance matrices, considering uncertainty in future returns and covariances.
result Marginalization introduces a logarithmic dependence on risk, leading to lower allocation levels for higher uncertainties.

New method optimizes model selection in high-dimensional regression models.

problem Model selection in high-dimensional misspecified regression models with covariate shift.
method Importance-weighted orthogonal greedy algorithm (IWOGA) and high-dimensional importance-weighted information criterion (HDIWIC).
result IWOGA + HDIWIC achieves optimal convergence rates in terms of prediction error.

The paper derives theoretical foundations for two common machine learning variable importance measures.

problem Understanding variable importance in machine learning problems.
method The paper derives closed-form expressions for Permute-and-Predict (PaP) and Leave-One-Covariate-Out (LOCO) methods.
result Theoretical derivations explain the behavior of PaP and LOCO under collinearity, linking them to coefficients and predictor variability.

CPI overcomes limitations of permutation importance by providing accurate variable selection.

problem Misidentification of unimportant variables in complex models due to covariate correlations.
method Developed a model agnostic and computationally lean Conditional Permutation Importance (CPI) approach.
result CPI provides accurate type-I error control and more parsimonious variable selection.

The paper compares LOCO and Shapley values for feature importance, highlighting their limitations and suggesting improvements.

problem Quantifying feature importance in the presence of feature correlation.
method LOCO and Shapley Values, critiquing their axioms and proposing new measures.
result Shapley values do not eliminate feature correlation, and a modified LOCO is recommended.

In the covariate shift learning scenario, the training and test covariate distributions differ, so that a predictor's average loss over the training and test distributions also differ. In this work, we explore the potential of extreme dimension reduction, i.e. to very low dimensions, in improving the performance of imp…

2017-11-29abs ↗pdf ↗

Study forecasts volatility and risk in electricity markets using matrix-HAR models.

problem Forecasting volatility and risk in electricity markets.
method Constructed a parsimonious matrix-HAR type model to estimate realized covariation and risk premia in electricity markets.
result Inclusion of longer time horizons and renewable generation information improves forecasts.

Network-assisted regression uses conformal prediction for valid inference.

problem Predicting node attributes using network and conventional covariates with valid statistical inference.
method Network analog of conformal prediction under mild joint exchangeability assumption.
result Achieves finite sample validity and asymptotic conditional validity for various network covariates.

New estimator handles covariate shift with closed-form solution and super-efficiency.

problem Handling covariate shift in missing data and causal inference problems.
method Minimum Wasserstein distance estimation framework.
result Closed-form expression and super-efficiency relative to semiparametric efficient estimator.

A new algorithm speeds up rerandomization for better experiment balance.

problem Achieving optimal covariate balance in randomized experiments.
method Metropolis-Hastings framework with sampling-importance resampling.
result PSRSRR achieves significant speedups while maintaining statistical guarantees.

We propose a novel calibration method for computer simulators, dealing with the problem of covariate shift. Covariate shift is the situation where input distributions for training and test are different, and ubiquitous in applications of simulations. Our approach is based on Bayesian inference with kernel mean embeddin…

2018-09-21abs ↗pdf ↗

Anomalies and outliers are common in real-world data, and they can arise from many sources, such as sensor faults. Accordingly, anomaly detection is important both for analyzing the anomalies themselves and for cleaning the data for further analysis of its ambient structure. Nonetheless, a precise definition of anomali…

2018-11-10abs ↗pdf ↗

Study tackles variable selection with missing covariates and outcomes using machine learning and imputation.

problem Missing data in both covariates and outcomes complicates variable selection in health studies.
method Exploits machine learning flexibility and bootstrap imputation for variable selection, comparing multiple methods.
result XGBoost and BART perform best in variable selection with bootstrap imputation, achieving high F1F_1 scores and low Type I errors.

Paper analyzes spectral algorithms under covariate shift, providing convergence rates.

problem Addressing distributional mismatch in regression models.
method Incorporates importance weights into spectral algorithms in RKHS.
result Establishes minimax-optimal convergence rates for misspecified cases.

Method tackles missing covariates in large-scale datasets.

problem Cross-population missing data problem in large-scale datasets.
method Augmented transfer regression learning method combining importance-weighted estimating equations and imputation terms.
result Estimator is n1/2n^{1/2}-consistent and asymptotically normal, attaining semiparametric efficiency bound under correct specification.

Enhanced Transformer models predict ETF portfolio performance by optimizing covariance and semi-covariance matrices.

problem Static covariance estimates fail to capture dynamic market fluctuations and non-linear correlations.
method Transformer-based models for real-time covariance and semi-covariance predictions.
result Portfolios optimized with semi-covariance matrix outperform those with standard covariance matrix, especially in volatile conditions.

There has been a lot of work fitting Ising models to multivariate binary data in order to understand the conditional dependency relationships between the variables. However, additional covariates are frequently recorded together with the binary data, and may influence the dependence relationships. Motivated by such a d…

2012-09-27abs ↗pdf ↗

This paper explores how effective sample size, dimensionality, and model performance are related in covariate shift adaptation.

problem Understanding the relationship between effective sample size, dimensionality, and generalization in covariate shift adaptation.
method Building a unified theory connecting effective sample size, data dimensionality, and generalization in the context of covariate shift adaptation.
result Dimensionality reduction or feature selection can increase effective sample size, supporting the practice of reducing dimensionality before covariate shift adaptation.

This paper rethinks confidence calibration under covariate shifts.

problem Calibration methods struggle with covariate shifts and unstable importance weighting.
method Derives Expectation consistency condition and proposes Expectation consistency loss (ECL).
result ECL loss is compatible with various types of calibration and has the same sample complexity as ECE.