New method estimates treatment effects from high dimensional data.
problem Estimating treatment effects from high dimensional data with confounders.
method Generative modeling approach to backdoor adjustment in variational inference.
result Empirically, estimates interventional likelihood in high dimensional settings.
We study the problem of treatment effect estimation in randomized experiments with high-dimensional covariate information, and show that essentially any risk-consistent regression adjustment can be used to obtain efficient estimates of the average treatment effect. Our results considerably extend the range of settings …
High-dimensional adjustment reduces bias in estimating peer effects from observational data.
problem Estimating peer effects from observational data is challenging due to confounding variables and high bias.
method Used high-dimensional adjustment with propensity score models to estimate peer effects.
result High-dimensional adjustment produces estimates of peer effects statistically indistinguishable from randomized experiments.
New method addresses crowding in high-dimensional data visualization.
problem Crowding issue in visualizing high-dimensional data.
method Adjusting capacity of high-dimensional balls and estimating correlation dimension.
result Mitigates crowding in various distance metrics.
A new perceptual adjustment query for metric learning reduces complexity in high-dimensional data.
problem Metric learning in high-dimensional data with limited human feedback.
method Inverted measurement scheme and two-stage estimator for PAQs.
result Sample complexity guarantees for the two-stage estimator of metric learning from PAQs.
The paper provides PAC bounds for estimating causal effects using covariate adjustment with a valid set.
problem Estimating causal effects in high-dimensional settings without randomized experiments.
method PAC learning perspective, valid adjustment set, $\eps$ -Markov blanket, constraint-based algorithms.
result PAC-bounds the estimation error of covariate adjustment by a term exponential in the size of the adjustment set.
SEDA improves RLDA for high-dimensional data.
problem Inconsistent performance of RLDA in high-dimensional scenarios.
method Developed a non-asymptotic approximation of misclassification rate, derived new theoretical results on eigenvectors, and proposed SEDA algorithm.
result SEDA achieves higher classification accuracy and dimensionality reduction compared to existing LDA methods.
This paper tackles robust factor models for high-dimensional data.
problem Challenges in high-dimensional, dependent data from various fields.
method Robust high-dimensional factor analysis.
result Classical PCA can be adapted for modern statistical challenges.
New method improves sampling from high-dimensional target densities.
problem Sampling from high-dimensional target densities using Monte Carlo algorithms.
method Extends Metropolis-Adjusted Langevin Diffusion algorithm with random precondition matrix modeling.
result Significantly improves performance and computational efficiency over standard MCMC methods.
A new sampler for complex discrete distributions efficiently updates all variables in parallel.
problem Sampling complex high-dimensional discrete distributions efficiently and accurately.
method Discrete Langevin proposal (DLP) for parallel coordinate updates with controlled stepsize.
result DLP efficiently explores high-dimensional and strongly correlated variables with asymptotic bias of zero for log-quadratic distributions.
Proposes a convex method to estimate GGMs with covariates.
problem Improving conditional independence structure estimation with covariates.
method Convex optimization framework for joint estimation of mean and precision matrix.
result Improved theoretical guarantees and practical utility demonstrated.
Proposes AAA for efficient association estimation with confounders.
problem Summarizing log odds ratio as a function of confounders.
method Develops efficient DML estimators for AAA.
result Demonstrates practicality and effectiveness of AAA estimators.
SP-SPCA improves sparse PCA by adaptively adjusting variable penalties, enhancing interpretability and stability.
problem Poor interpretability and variable redundancy in PCA for high-dimensional data.
method Introduces a single equilibrium parameter to adaptively adjust variable penalties in the L2 regularization framework.
result Consistently outperforms standard sparse PCA methods in identifying sparse loading patterns and preserving cumulative variance.
Enhances forecasting of complex systems using FKMD.
problem Forecasting high-dimensional dynamical systems with unknown features.
method Featurized Koopman Mode Decomposition (FKMD) using delay embedding and learned Mahalanobis distance.
result Improves prediction accuracy for various complex systems.
A new explicit scheme calculates XVA adjustments using neural networks and conditional expectations.
problem Calculating cross valuation adjustments (XVA) in realistic financial scenarios.
method Simulation/regression scheme for BSDEs, using neural networks and quantile regressions.
result The scheme outperforms Picard iterations in high-dimensional and hybrid market risks.
DOPE efficiently estimates ATE with complex covariates.
problem Efficient estimation of ATE from complex covariates.
method Proposed DOPE framework for efficient adjustment.
result DOPE retains efficiency even with highly predictive covariates.
High-dimensional unimodal distributions can cause MCMC methods to fail.
problem Failure of MCMC methods in high-dimensional unimodal distributions.
method Examples and theoretical analysis of MCMC methods, including Metropolis-Hastings adjusted methods.
result MCMC methods can take an exponential run-time for high-dimensional unimodal distributions.
Robust and reliable covariance estimates play a decisive role in financial and many other applications. An important class of estimators is based on Factor models. Here, we show by extensive Monte Carlo simulations that covariance matrices derived from the statistical Factor Analysis model exhibit a systematic error, w…
With the development of high-throughput technologies, principal component analysis (PCA) in the high-dimensional regime is of great interest. Most of the existing theoretical and methodological results for high-dimensional PCA are based on the spiked population model in which all the population eigenvalues are equal ex…
MOCA uses modular attention to estimate causal effects from complex data.
problem Estimating causal effects from observational data with complex, non-linear, and high-dimensional treatment and outcome mechanisms.
method MOCA is a transformer-based framework that separates treatment and outcome modeling through modular design and one-way attention mechanism, with cutting-feedback to prevent outcome influence on treatment representations.
result MOCA outperforms classical estimators and machine learning approaches across various simulated and real-world scenarios.
Proposes GAGA algorithm for automatic hyperparameter learning in signal recovery.
problem Difficulty in selecting hyperparameters in traditional signal recovery methods.
method Global Adaptive Generative Adjustment (GAGA) algorithm for automatic hyperparameter learning and signal estimate.
result Consistency of model selection and signal estimate output.
The study uses pre-trained neural networks to adjust for confounding in non-tabular data.
problem Neglecting non-tabular data sources can lead to biased ATE estimates.
method Leverages latent features from pre-trained neural networks to adjust for confounding.
result Neural networks can achieve fast convergence rates for ATE estimation with latent features.
This paper addresses credit valuation adjustment with a new closeout convention.
problem Accurate estimation of financial claim value considering counterparty credit risk.
method Theoretical and computational analysis of a nonlinear valuation system using neural networks.
result A neural network-based algorithm effectively solves the high-dimensional nonlinear valuation system.
FASC clusters data with latent factors, improving on naive methods.
problem Clustering high-dimensional data with correlated variables.
method Factor Adjusted Spectral Clustering (FASC) algorithm.
result FASC achieves an exponentially low mislabeling rate under general assumptions.
New method controls bias in unadjusted Hamiltonian Monte Carlo and underdamped Langevin.
problem Bias in unadjusted Hamiltonian Monte Carlo and underdamped Langevin samplers.
method Delocalization of bias technique applied to these samplers.
result Control W 2 W_2 W 2 bias with O ( K ) O(\sqrt{K}) O ( K ) integration steps for high-dimensional distributions. Self-Distilled Disentanglement improves counterfactual predictions by separating variables.
problem Improving counterfactual predictions in the presence of confounders and unobserved variables.
method Self-Distilled Disentanglement framework based on information theory.
result Effective counterfactual inference in synthetic and real-world datasets.
The paper models and predicts co-occurrence counts using Gamma regression.
problem Predicting relevance between items or users from high-dimensional sparse co-occurrence count data.
method Shared parameter alternating zero-inflated Gamma regression models (SA-ZIG) with Fisher scoring and learning rate adjustment.
result SA-ZIG with learning rate adjustment performs satisfactorily in predicting relevance.
A new method for clustering heterogeneous data using likelihood-adjusted SDP.
problem Clustering heterogeneous data with different cluster shapes and sizes.
method Iterative likelihood-adjusted semidefinite programming (iLA-SDP) method.
result iLA-SDP achieves lower mis-clustering errors compared to other methods.
RL accelerates portfolio optimization and option pricing by dynamically adjusting preconditioner sizes.
problem Large linear systems in portfolio optimization and option pricing lead to slow convergence.
method Reinforcement Learning (RL) dynamically adjusts block-preconditioner sizes to accelerate convergence.
result RL-driven solver significantly reduces computational cost and accelerates convergence.
Analysis of momentum methods on quadratic models, showing SGD's superiority.
problem Analysis of stochastic gradient algorithms with momentum on quadratic models.
method Inspired by random matrix theory, exact characterization of loss values.
result Stochastic heavy-ball momentum does not improve over SGD in the strongly convex setting.
AutoML enhances clinical metabolic profiling by adjusting for confounders.
problem Identifying and adjusting for clinical confounders in AutoML for metabolic profiling.
method Tandem rank-accuracy measure for feature selection, residual training adjustment for confounders.
result Increased homocysteine concentration associated with long-term metformin exposure.
A winning method for day-ahead electricity demand forecasting during and after the COVID-19 pandemic.
problem Day-ahead electricity demand forecasting during and after the COVID-19 pandemic.
method Online forecast combination of multiple point prediction models with a holiday adjustment procedure and smoothed Bernstein Online Aggregation (BOA).
result Excellent forecasting performance, particularly due to the holiday adjustment procedure and fully adaptive smoothed BOA approach.
Develops a gradient-enhanced approach for online estimation in high-dimensional generalized linear models with streaming data.
problem Online estimation for high-dimensional generalized linear models with streaming data.
method Proposes a gradient-enhanced surrogate loss for non-distributed setting and extends to distributed streaming data.
result Derives non-asymptotic error bounds under high-dimensional scaling without batch-number constraint.
High-dimensional VAEs inevitably collapse to prior, requiring large datasets for good performance.
problem Posterior collapse in VAEs leads to poor representation learning quality.
method Analyzed a minimal VAE in a high-dimensional limit, evaluating conditions for posterior collapse with respect to beta and dataset size.
result VAEs face 'inevitable posterior collapse' beyond a certain beta threshold, regardless of dataset size.
Develops a dynamic latent-factor model for high-dimensional asset characteristics.
problem Estimating asset pricing tests with high-dimensional data.
method Dynamic latent-factor model with Double Selection Lasso regularization.
result The inflation-mimicking portfolio in the crypto asset class has positive risk compensation.
Optimal preconditioning improves Langevin sampling efficiency.
problem Improving sampling efficiency in high-dimensional target distributions.
method Optimal preconditioning using Fisher information, applied to MALA.
result Adaptive MCMC scheme significantly outperforms other methods.
Study EM and GD for clustering with penalties for misspecification and high dimensions.
problem Clustering with misspecification and high-dimensional data.
method Model-based Gaussian Mixture Models, EM algorithm, GD optimization with AD, penalized likelihood.
result GD outperforms EM on high-dimensional data but both have poor cluster interpretation.
A deep BSDE approach tackles multi-layered xVA calculations for portfolio valuation.
problem Computational intractability in nested simulations for multi-layered xVA calculations.
method Iterative deep BSDE approach, change-of-measure method, quantile regression for margin computation.
result Reduces computational demands and successfully scales to high-dimensional portfolios.
Generalizes causal inference to high-dimensional outcomes.
problem Limited causal inference methods for multivariate outcomes.
method Formulates causal discrepancy tests for nominal variables, uses conditional independence tests.
result Causal CDcorr method improves finite sample validity and power.
The paper tackles counterfactual inference with multioutput deep kernels in high-dimensional settings.
problem Performing counterfactual inference with observational data in high-dimensional settings with multiple actions and outcomes.
method The paper presents a general class of counterfactual multi-task deep kernels models based on Structural Causal Models (SCM) and Gaussian Processes.
result The models estimate causal effects and learn policies efficiently, scaling well with high dimensions.
PROBE algorithm efficiently solves sparse high-dimensional linear regression.
problem Sparse high-dimensional linear regression models with complex parameter spaces.
method Partitioned empirical Bayes ECM algorithm for computationally efficient MAP estimation.
result PROBE algorithm provides robust and efficient coordinate-wise optimization.
SPPCSO addresses multicollinearity in high-dimensional data, improving model stability and predictive accuracy.
problem Multicollinearity in high-dimensional data leads to unstable estimation and reduced predictive accuracy.
method SPPCSO integrates principal component regression and L1 regularization to adaptively adjust shrinkage factors.
result SPPCSO achieves stable and reliable estimation in high-noise settings, distinguishing signal variables from noise.
Sparse Canonical Correlation Analysis (CCA) has received considerable attention in high-dimensional data analysis to study the relationship between two sets of random variables. However, there has been remarkably little theoretical statistical foundation on sparse CCA in high-dimensional settings despite active methodo…
Robust variable selection for high-dimensional data with missing and measurement errors.
problem Missing data and measurement errors confound data distribution.
method Exponential loss function with inverse probability weighting and additive error models.
result The Atan punishment method improves robust variable selection.
Bayesian DDR models complex multivariate distributions.
problem Modeling relationships between multivariate distributions with differing dimensions.
method Generalized Bayesian framework using sliced Wasserstein distance and MALA for inference.
result Posterior consistency and robust fits demonstrated in simulations and real data.
This paper finds a linear relationship between t-SNE perplexity and data set size.
problem Choosing the right perplexity for t-SNE embeddings.
method Analyzed the relationship between perplexity and data set size.
result Embeddings remain structurally consistent when perplexity is adjusted accordingly.
A new PCA-based imputation method for high-dimensional data.
problem Missing data in high-dimensional datasets.
method Principal Component Analysis Imputation (PCAI) framework.
result PCAI significantly speeds up imputation and maintains high accuracy.
New method improves counterfactual distribution learning for high-dimensional outcomes.
problem Counterfactual distribution learning for high-dimensional outcomes with concentrated structure.
method Geometry-adaptive diffusion-guided smoothing estimators combining causal nuisance adjustment and local outcome geometry.
result Geometry-adaptive methods show steeper error decay in semi-synthetic experiments.