Study mutual info for community detection with covariate and correlated networks.
problem Community detection with covariate and correlated networks.
method Asymptotic upper bound and MMSE matrix heuristic analysis.
result Explicit characterization of combined information effects.
Steerable E(3) Graph Neural Networks incorporate geometric and physical covariant information.
problem Incorporating covariant information like position, force, velocity, or spin in graph neural networks.
method Steerable E(3) Equivariant Graph Neural Networks (SEGNNs) that use steerable MLPs to incorporate geometric and physical covariant information.
result SEGNNs improve upon classic linear point convolutions and recent equivariant graph networks that send invariant messages.
A new model DKMPP integrates covariates and uses an integration-free method for spatio-temporal point processes.
problem Training intractable deep spatio-temporal point processes with multimodal covariates.
method DKMPP uses a deep kernel to model complex relationships and an integration-free score matching method.
result DKMPP and score-based estimators outperform baseline models in spatio-temporal point processes.
Paper proposes a new method for SP with covariates using PADR and ERM.
problem Stochastic programming with covariate information.
method Empirical risk minimization (ERM) with nonconvex piecewise affine decision rules (PADR).
result The method provides theoretical consistency and computational tractability for nonconvex SP problems.
New method prevents posterior collapse in iVAE models.
problem Posterior collapse in iVAE models where observations and ICs are independent given covariates.
method Developed CI-iVAE by considering a mixture of encoder and posterior distributions in the objective function.
result Prevents posterior collapse, resulting in latent representations with more information of the observations.
Proposes a method to recover sparse tensors with covariate info.
problem Sparse tensor with high missing entries and many zeros.
method Covariate-assisted Sparse Tensor Completion (COSTCO) using latent components.
result 23% accuracy improvement over baseline in advertisement dataset.
Flexible framework integrates machine learning and DRO for uncertain parameter prediction.
problem Limited joint observations of uncertain parameters and covariates.
method Wasserstein, sample robust optimization, and phi-divergence-based ambiguity sets.
result Validation of theoretical and practical benefits in limited data scenarios.
New criterion improves predictive evaluation in weighted inference scenarios.
problem Improving predictive evaluation in scenarios with different likelihoods for estimation and evaluation.
method Developed the posterior covariance information criterion (PCIC) to handle weighted likelihood inference.
result PCIC is asymptotically unbiased for quasi-Bayesian generalization error in weighted inference.
Semi-supervised method boosts two-sample testing with covariate data.
problem Two-sample testing with covariate information.
method Semi-supervised kernel test with asymptotic normality.
result Higher asymptotic power compared to existing methods.
The paper improves ranking by integrating covariates and sparse intrinsic scores.
problem Ranking items with incomplete preference scores explained by covariates.
method Extends BTL model with covariate information and sparse intrinsic scores, using penalized MLE.
result Developed debiased estimator for penalized MLE with distributional properties.
Paper proposes CARE model for ranking with covariates, improving MLE accuracy.
problem Statistical estimation and inference for ranking with covariate information.
method Covariate-Assisted Ranking Estimation (CARE) model, extending Bradley-Terry-Luce (BTL) model.
result Derives optimal rates and asymptotic distributions for MLE of latent scores and covariates.
Contextual information helps identify the best arm more efficiently.
problem Best arm identification with contextual covariate information.
method Proposed a context-aware version of the 'Track-and-Stop' strategy.
result Expected number of arm draws matches lower bound asymptotically.
We describe parallel Markov chain Monte Carlo methods that propagate a collective ensemble of paths, with local covariance information calculated from neighboring replicas. The use of collective dynamics eliminates multiplicative noise and stabilizes the dynamics thus providing a practical approach to difficult anisotr…
One of the most fundamental problems in network study is community detection. The stochastic block model (SBM) is a widely used model, for which various estimation methods have been developed with their community detection consistency results unveiled. However, the SBM is restricted by the strong assumption that all no…
Method uses semi-supervised learning to estimate optimal treatment regimes from medical records.
problem Estimating optimal treatment regimes from electronic medical records.
method Imputation-based semi-supervised method using unlabeled data.
result Proposed method yields more efficient estimators of optimal treatment regimes.
Study integrates machine learning with SAA for optimizing decisions based on uncertain parameters and covariates.
problem Optimizing decisions under uncertain parameters and covariates.
method Data-driven frameworks integrating machine learning prediction models within SAA for scenario generation.
result Consistent and asymptotically optimal solutions under certain conditions, with finite sample guarantees.
Paper uses machine learning in EM framework for better nowcasting.
problem Improving risk estimation with incomplete information.
method Expectation-maximisation framework with machine learning for occurrence and reporting modeling.
result XGBoost-based approach outperforms existing methods in nowcasting.
Chronos-2 forecasts multivariate and covariate data without task-specific training.
problem Limited applicability of existing time series forecasting models to real-world multivariate and covariate data.
method Chronos-2 uses a group attention mechanism for in-context learning across multiple time series.
result Chronos-2 achieves state-of-the-art performance across comprehensive benchmarks.
The paper infers multiple graphs from stationary signals on them.
problem Inferring multiple graphs from signals observed on their nodes.
method Convex optimization method leveraging matrix polynomial commutation.
result High-probability bounds on recovery error provided.
Novel network model estimates mixed-membership structure with covariate information.
problem Estimating latent mixed-membership structure in networks with covariate information.
method Proposes a novel network model that incorporates both community information and node covariate similarities.
result Achieves optimal estimation accuracy for similarity matrix and mixed-membership.
New methods rank players using covariates and comparisons, outperforming existing algorithms.
problem Ranking players based on incomplete and noisy pairwise comparisons.
method Three spectral ranking methods incorporating player covariates.
result Proposed methods outperform existing algorithms in simulations.
Classification is an important tool with many useful applications. Among the many classification methods, Fisher's Linear Discriminant Analysis (LDA) is a traditional model-based approach which makes use of the covariance information. However, in the high-dimensional, low-sample size setting, LDA cannot be directly dep…
We extend multi-way, multivariate ANOVA-type analysis to cases where one covariate is the view, with features of each view coming from different, high-dimensional domains. The different views are assumed to be connected by having paired samples; this is a common setup in recent bioinformatics experiments, of which we a…
We study methods for simultaneous analysis of many noisy experiments in the presence of rich covariate information. The goal of the analyst is to optimally estimate the true effect underlying each experiment. Both the noisy experimental results and the auxiliary covariates are useful for this purpose, but neither data …
Develops model-free methods for event history analysis and efficient covariate adjustment.
problem Estimating treatment effects while accounting for confounding and understanding event history.
method Model-free prediction techniques, Local Covariance Measure (LCM), Debiased Outcome-adapted Propensity Estimator (DOPE), Aalen Covariance Measure (ACM).
result Demonstrates the effectiveness and robustness of the proposed methods in various settings.
The literature provides strong evidence that stock prices can be predicted from past price data. Principal component analysis (PCA) is a widely used mathematical technique for dimensionality reduction and analysis of data by identifying a small number of principal components to explain the variation found in a data set…
We study the problem of treatment effect estimation in randomized experiments with high-dimensional covariate information, and show that essentially any risk-consistent regression adjustment can be used to obtain efficient estimates of the average treatment effect. Our results considerably extend the range of settings …
Variable selection in high dimensional space has challenged many contemporary statistical problems from many frontiers of scientific disciplines. Recent technology advance has made it possible to collect a huge amount of covariate information such as microarray, proteomic and SNP data via bioimaging technology while ob…
In this paper, we propose a new framework to remove parts of the systematic errors affecting popular restoration algorithms, with a special focus for image processing tasks. Generalizing ideas that emerged for ℓ1 regularization, we develop an approach re-fitting the results of standard methods towards the input d…
Optimization algorithms that leverage gradient covariance information, such as variants of natural gradient descent (Amari, 1998), offer the prospect of yielding more effective descent directions. For models with many parameters, the covariance matrix they are based on becomes gigantic, making them inapplicable in thei…
A new method for efficient portfolio optimization using graph structures.
problem Optimizing portfolio weights while reducing computational complexity.
method Hierarchical graph structures and Schur complement method.
result Optimal portfolio weights can be computed efficiently by inverting small submatrices.
Proposes a method to represent high-dimensional covariates for causal inference.
problem Inefficient and unreliable causal inference with high-dimensional covariates.
method Machine-learning-assisted covariate representation approach.
result Statistical reliability and performance guarantees for proposed methods.
AUASE embeds dynamic networks with stability guarantees for node comparison.
problem Stability in dynamic network embeddings for comparing nodes across time.
method Attributed unfolded adjacency spectral embedding (AUASE) for stable unsupervised learning.
result AUASE provides significant improvements in link prediction and node classification.
Learning predictive models from small high-dimensional data sets is a key problem in high-dimensional statistics. Expert knowledge elicitation can help, and a strong line of work focuses on directly eliciting informative prior distributions for parameters. This either requires considerable statistical expertise or is l…
Deep neural networks suffer from over-fitting and catastrophic forgetting when trained with small data. One natural remedy for this problem is data augmentation, which has been recently shown to be effective. However, previous works either assume that intra-class variances can always be generalized to new classes, or e…
Method detects errors in numerical data using regression models.
problem Noise and errors in numerical datasets.
method Introduced veracity scores and a filtering procedure for error detection.
result Method outperforms other approaches in identifying incorrect values.
DeepHazard uses neural networks to predict time-varying survival risks.
problem Traditional survival models assume proportional hazards and do not account for time-varying covariate information.
method DeepHazard is a neural network approach that models time-varying hazards without proportional hazards assumption.
result DeepHazard outperforms existing methods in predicting survival time, as shown by C-index metrics on real datasets.
DeepMPM uses MPM to make deep neural networks more robust against adversarial attacks.
problem Vulnerability of deep neural networks to adversarial attacks.
method Applying MPM to deep neural networks in an end-to-end fashion to minimize an upper bound of misclassification probabilities considering global class information.
result DeepMPM achieves comparable classification performance with CNN but is more robust against adversarial attacks.
The paper develops exact and approximate conformal inference methods for multi-output regression.
problem Uncertainty quantification in multi-output regression predictions.
method Exact derivations and approximations of conformal inference p-values for linear multi-output predictors, and efficient methods for nonlinear predictors.
result Efficient methods for approximating conformal prediction regions for multi-output predictors, both linear and nonlinear.
Proposes L-VAE for longitudinal data analysis.
problem Analyse high-dimensional longitudinal data with missing values.
method Uses a multi-output additive Gaussian process (GP) prior to extend VAE's capability.
result Achieves highly accurate predictive performance.
Develops intrinsic Gaussian process regression for manifold-valued data.
problem Lack of intrinsic Gaussian process methods for manifold-valued response variables.
method Proposes an intrinsic covariance structure and a novel intrinsic Gaussian process regression model.
result Establishes asymptotic properties and shows posterior consistency.
Latent space models are effective tools for statistical modeling and exploration of network data. These models can effectively model real world network characteristics such as degree heterogeneity, transitivity, homophily, etc. Due to their close connection to generalized linear models, it is also natural to incorporat…
A classical problem in causal inference is that of matching, where treatment units need to be matched to control units based on covariate information. In this work, we propose a method that computes high quality almost-exact matches for high-dimensional categorical datasets. This method, called FLAME (Fast Large-scale …
Timer-XL predicts multidimensional time series using a unified Transformer approach.
problem Unified time series forecasting across various tasks and contexts.
method Decoder-only Transformers with a universal TimeAttention mechanism and deft position embedding.
result State-of-the-art performance across multiple forecasting benchmarks.
Method estimates CATE using RCT data to handle hidden confounders.
problem Estimating CATE in the presence of hidden confounders.
method Pseudo-confounder generator and CATE model alignment.
result Method reduces bias in CATE estimation.
New linear algorithms improve wSVMs for multiclass probability estimation.
problem Estimating conditional probabilities for multiclass problems.
method Proposed baseline learning and OVA learning schemes to improve wSVMs.
result Linear algorithms achieve optimal computational efficiency and good estimation accuracy.
This work develops a model to distinguish network and covariate information.
problem Identifying unique network and covariate information.
method Low-rank model with two-step estimation: spectral method followed by refinement.
result The method accurately recovers joint and individual components.
In this paper, we consider the multivariate Bernoulli distribution as a model to estimate the structure of graphs with binary nodes. This distribution is discussed in the framework of the exponential family, and its statistical properties regarding independence of the nodes are demonstrated. Importantly the model can e…