Randomly selected factors preserve correlation structure in high-dimensional data.
problem Preserving correlation structure in high-dimensional data.
method Random projection method to select factors, preserving covariance matrix and time-series accuracy.
result Randomly selected factors accurately represent time-series and their cross-correlations.
Novel neural GP kernels learn stable, flexible covariance structures.
problem Scalable and flexible covariance kernels for Gaussian processes.
method Directly learn kriging coefficients and conditional standard deviations using deep neural architectures exploiting permutation-equivariant structure.
result Improved training stability and data efficiency with expressive, non-stationary kernels.
Diagonal transformations preserve independence structures in non-Gaussian distributions.
problem Preserving independence structures in non-Gaussian distributions.
method Diagonal nonlinear transformations of multivariate normal variables.
result Independence structures are preserved in non-Gaussian distributions under diagonal transformations.
New federated method preserves privacy and estimates treatment effects.
problem Privacy-preserving causal inference for multi-site studies.
method Multiply robust nuisance function estimation, transfer learning.
result Efficient and optimal treatment effect estimation under different scenarios.
Improves low-shot learning with novel GAN for diverse example generation.
problem Overfitting and forgetting in small data settings.
method Covariance-Preserving Adversarial Augmentation Networks (CPGANs).
result Significant improvement on ImageNet benchmark.
New method needed for class prior estimation when covariates are reduced.
problem Class prior estimation fails under covariate shift when covariates are reduced.
method Propose a probing algorithm for class prior estimation.
result Provable transformations preserving covariate shift are necessary for class prior estimation.
Generative models often fail to preserve joint structure despite matching marginals.
problem Generative models fail to capture complex dependencies beyond univariate marginals.
method Introduced D_Sigma(P,Q) = ||Sigma_P - Sigma_Q||_F to measure covariance-level dependence fidelity.
result Covariance-level divergence can lead to structural instability in downstream inference.
Paper tackles imbalanced time series classification with a novel oversampling method.
problem Imbalanced time series classification challenges due to high dimensionality and correlation.
method Density-ratio based clustering followed by shrinkage technique for covariance estimation, then generating synthetic samples.
result OHIT outperforms state-of-the-art methods in F1, G-mean, and AUC metrics.
A new method preserves useful information in data rows with outlying cells.
problem Preserving useful information in data rows with outlying cells.
method Cellwise robust Minimum Covariance Determinant (cellMCD) method using observed likelihood and a penalty term on cellwise outliers.
result The cellMCD method performs well in simulations and on real data.
Proposes a network to predict structured uncertainty distributions for images.
problem Previous methods only predicted diagonal covariance matrices, limiting reconstruction accuracy.
method Learns to predict full Gaussian covariance matrices for efficient sampling and likelihood evaluation.
result Accurately reconstructs ground truth correlated residual distributions and generates plausible high frequency samples.
A new method to improve image restoration by re-fitting standard methods.
problem Systematic errors in popular restoration algorithms for image processing.
method Developing a re-fitting approach that preserves covariant information and Jacobian of original estimators.
result Improved image restoration results on numerical simulations.
We show that if a Finsler space is conformally automorphic to a Riemannian space and the automorphism is positively homogeneous with respect to tangent vectors, then the indicatrix of the Finsler space is a space of constant curvature. In this case, the Finslerian two-vector angle can explicitly be found, which gives r…
SDR outperforms IDR in multimodal data analysis, especially with fewer samples.
problem Understanding and optimizing data efficiency in multimodal representation learning.
method Generative linear model to synthesize multimodal data, comparing IDR and SDR methods.
result Linear SDR methods yield higher-quality, more succinct reduced-dimensional representations with smaller datasets.
The Finsler spaces in which the tangent Riemannian spaces are conformally flat prove to be characterized by the condition that the indicatrix is a space of constant curvature. In such spaces the Finslerian normalized two-vector angle can be explicated from the respective two-vector angle of the associated Riemannian sp…
optHSIC tests independence between covariates and censored lifetimes using optimal transport.
problem Testing independence between a covariate and right-censored lifetimes.
method optHSIC uses optimal transport to transform censored data into uncensored data, then applies a permutation test with a kernel-based dependence measure.
result optHSIC has power against a wider class of alternatives than Cox regression and maintains type 1 error control even when censoring depends on the covariate.
New neural network captures spatial correlations in wind speed predictions.
problem Uncertainty quantification in neural network predictions for high-dimensional, correlated data.
method Training neural networks with multidimensional Gaussian loss, preserving spatial correlation and computational tractability.
result Demonstrated super-resolution of surface wind speed with explicit correlation modeling.
DARTS optimizes covariate selection in trials with limited data.
problem Limited budget for high-dimensional pretreatment data.
method Dynamic Adaptive Rerandomization via Thompson Sampling (DARTS).
result DARTS efficiently concentrates budget on informative features.
Generative synthetic data can preserve predictive accuracy but distort causal inference.
problem Distortion of average treatment effect estimates in synthetic data.
method Hybrid synthetic-data framework that generates covariates while modeling treatment and outcome mechanisms separately.
result Hybrid synthesis improves causal fidelity compared to fully generative baselines.
Kalman filtering and smoothing algorithms are used in many areas, including tracking and navigation, medical applications, and financial trend filtering. One of the basic assumptions required to apply the Kalman smoothing framework is that error covariance matrices are known and given. In this paper, we study a general…
Paper optimizes federated PCA for covariance estimation under privacy constraints.
problem Privacy-preserving covariance estimation in federated learning.
method Federated PCA, matrix version of van Trees' inequality, three-layer spectral decomposition.
result Optimal rates of convergence for central server's estimation, robust to inconsistent local estimators.
Study on linear regression with dependent covariates, proving universality and error characterization.
problem Linear regression with dependent covariates in high-dimensional settings.
method Analysis of ridge regression performance, Gaussian universality theorem, spectral properties of covariance matrices.
result Asymptotic performance of ridge regression is invariant under non-Gaussian covariates with preserved mean and covariance.
NICE learns a representation to avoid bad controls in causal inference.
problem Avoiding bad controls in causal inference from observational data.
method Uses invariant risk minimization (IRM) to learn a representation of covariates that avoids bad controls.
result NICE outperforms adjusting for all covariates in cases with unknown collider variables and bad controls.
This work preserves linear invariants in ensemble filters for non-Gaussian data assimilation.
problem Maintaining critical invariants like mass, stoichiometric balance, and charge in non-Gaussian data assimilation.
method Introducing a novel class of nonlinear ensemble filters using measure transport theory.
result Recovery of a constrained Kalman filter for Gaussian settings and combination with regularization techniques.
A classification scheme of the conformal almost contact metric manifolds with respect to the covariant derivative of the Lee form is given. The subclasses of one basic class and their exact characterizations by the maximal subgroups of the contact conformal group preserving itself are found.
Preserves sensitive data distribution for privacy while maintaining utility.
problem Protecting individual privacy while preserving data utility for specific analyses.
method Combines distribution-preserving quantization and k-member clustering.
result Demonstrates improved privacy and utility in real-world applications.
Paper tackles domain generalization by minimizing domain-based covariance.
problem Training data and test data have different distributions, leading to poor generalization.
method Find a central subspace minimizing domain-based covariance while preserving functional relationships.
result The proposed method achieves better generalization performance on unseen test datasets.
Paper optimizes private PCA for covariance estimation in statistics.
problem Private estimation of covariance matrices and principal components.
method Developed differentially private estimators for spiked covariance model.
result Established minimax rates of convergence for principal components and covariance matrix estimation.
A new framework for efficient sequence maps using Bayesian filtering and covariance.
problem Designing efficient recurrent sequence maps from explicit memory assumptions.
method Design-model framework, exact Bayesian filtering, query-dependent readout, linear-Gaussian instantiation.
result Improved robustness and retrieval performance across various benchmarks.
Meta learns low-rank covariance factors for better uncertainty estimation.
problem Sub-optimal covariance matrices in multi-task settings.
method Meta learns diagonal or diagonal plus low-rank factors using an attentive set encoder.
result Efficiently constructed task-specific covariance matrices improve uncertainty estimation.
Paper proposes a new method to protect model information in multi-task learning.
problem Protecting model information in multi-task learning from adversaries.
method Proposes a privacy-preserving MTL framework using perturbation of the covariance matrix.
result Our algorithms outperform existing privacy-preserving MTL methods and STL methods.
CIR method preserves relation for case-control studies.
problem Learning low-dimensional structure in case-control studies.
method Contrastive inverse regression (CIR) on Stiefel manifold.
result CIR outperforms other methods for high-dimensional data.
ITSPACE improves covariance alignment faster than other methods.
problem Optimizing covariance matrices for machine learning tasks.
method Proximal majorization-minimization method that directly optimizes the Bures-Wasserstein objective.
result ITSPACE achieves lower BW gap solutions faster than other methods.
Paper introduces a method to create robust representations against covariate shifts.
problem Distribution shift between training and testing data in machine learning.
method Introduces a variational objective with two components: discriminative representation and invariant support.
result Optimal representations ensure robustness to covariate shifts, improving performance on DomainBed.
Efficiently solves large portfolio optimization problems by reducing and sparsifying covariance matrices.
problem Large and dense covariance matrices limit efficient portfolio optimization.
method Dimension reduction and increased sparsity based on machine learning predictions.
result Improved portfolio performance and reduced runtime compared to full dense covariance matrices.
E-QRGMM accelerates uncertainty quantification in simulations.
problem Challenges in covariate-dependent uncertainty quantification.
method Integrates cubic Hermite interpolation with gradient estimation.
result Substantially improves computational efficiency and accuracy.
New method improves PCA for high-dimensional data with n < p.
problem PCA struggles in high-dimensional settings with n < p.
method Pairwise differences covariance estimation with four regularized versions.
result Proposed methods outperform existing estimators in high-dimensional data settings.
New estimator improves covariance matrix estimation under distributional uncertainty.
problem Estimating inverse covariance matrix under distributional uncertainty.
method Distributionally robust optimization with Wasserstein ambiguity set.
result Analytical solution as nonlinear shrinkage estimator.
New methods incorporate alpha signals into portfolio construction, improving performance.
problem Signal-blindness in existing portfolio construction methods.
method Introduces three methods: HRP-μ, HRP-Σμ, and CRISP. result CRISP at intermediate γ consistently outperforms other methods. Study addresses covariate mismatch in federated learning, improving model accuracy.
problem Learning from clients with different feature sets in federated learning.
method Developed two approaches for linear prediction under covariate mismatch: plug-in estimator and impute-then-regress strategy.
result Proposed methods provide asymptotic and finite-sample learning rates, improving model accuracy.
A new method for efficient portfolio optimization using graph structures.
problem Optimizing portfolio weights while reducing computational complexity.
method Hierarchical graph structures and Schur complement method.
result Optimal portfolio weights can be computed efficiently by inverting small submatrices.
Deep learning improves covariance matrix estimation for better portfolio risk management.
problem Improving the accuracy of covariance matrix estimation for portfolio risk management.
method Formulated as a learning problem, used deep learning to automatically discover risk factors.
result 1.9% higher explained variance and reduced portfolio risk.
Spectral graph sparsification preserves geometry of GNN embeddings.
problem Maintaining geometric properties of graph neural network embeddings during sparsification.
method Proving spectral sparsification preserves squared pairwise distances, class means, and covariance structure in embedding space.
result Spectral sparsification preserves the geometry of learned embeddings in GNNs.
Efficiently estimates covariance for sub-Weibull vectors with sub-Gaussian rate.
problem Outliers in high-dimensional covariance estimation.
method Cross-Fitted Norm-Truncated Estimator for Sub-Weibull distributions.
result Achieves optimal sub-Gaussian rate with O(Nd2) operations. New method models covariates and responses without parametric assumptions using manifold learning.
problem Losing explanatory power for responses in standard factor models applied to covariates alone.
method Anisotropic diffusion maps for learning low-dimensional embeddings.
result Kalman filtering in diffusion-map coordinates improves joint covariate-response prediction.
This paper presents a geometric-variational approach to continuous and discrete mechanics and field theories. Using multisymplectic geometry, we show that the existence of the fundamental geometric structures as well as their preservation along solutions can be obtained directly from the variational principle. In parti…
Hybrid LLM generates synthetic data preserving causal parameters.
problem Synthetic data fails to accurately estimate causal effects.
method Combines model-based covariate synthesis with separately learned propensity and outcome models.
result Hybrid framework ensures causal structure in synthetic data.
This paper improves STL inference reliability under covariate shift.
problem Ensuring correct STL formulas in real-world settings with distribution shift.
method Proposes a conformalized STL inference framework that addresses covariate shift.
result Significantly improves symbolic learning reliability at deployment time.
A new algorithm speeds up rerandomization for better experiment balance.
problem Achieving optimal covariate balance in randomized experiments.
method Metropolis-Hastings framework with sampling-importance resampling.
result PSRSRR achieves significant speedups while maintaining statistical guarantees.