Data augmentation can achieve the same statistical benefits as full augmentation up to an approximation error.
problem Data augmentation in learning problems
method Using Fourier analysis and representation theory of finite groups
result Partial data augmentation achieves the same minimax rates as full augmentation
This work improves fairness in federated learning by using zero-shot data augmentation.
problem Statistical heterogeneity leads to biased and less uniform accuracy across clients in federated learning.
method Proposes a federated learning system with zero-shot data augmentation to mitigate statistical heterogeneity and improve fairness.
result Empirical results show improved test accuracy and fairness across clients.
Develops a statistical framework for self-supervised representation learning using data augmentation.
problem Lack of theoretical understanding of data augmentation in nonlinear settings.
method Augmentation invariant manifold learning framework and stochastic optimization algorithm.
result Improves downstream analysis by exploiting manifold's geometric structure and invariant property of augmented data.
This work proposes a model for geodesic distances and flows on manifolds.
problem Geodesic distances and flows on differentiable manifolds.
method Manifold-augmented Eikonal equation solutions.
result Geodesic flow provides globally length-minimizing curves.
New statistical theory explains contrastive learning effectiveness.
problem Understanding why contrastive learning works well for representation extraction.
method Developed a new theoretical framework based on approximate sufficient statistics.
result Near-sufficient encoders derived from contrastive learning can be adapted for downstream tasks.
Framework uses synthetic data from pretrained models to improve predictive modeling.
problem Limited effectiveness of synthetic data from generative models for improving predictive performance.
method Proposes an end-to-end framework that generates and filters synthetic data through domain-specific statistical methods.
result Consistent improvements in predictive performance across various settings.
Data augmentation, by the introduction of auxiliary variables, has become an ubiquitous technique to improve convergence properties, simplify the implementation or reduce the computational time of inference methods such as Markov chain Monte Carlo ones. Nonetheless, introducing appropriate auxiliary variables while pre…
Data augmented bootstrap unifies various confidence interval construction methods.
problem Constructing confidence intervals from data transformations.
method Data augmented bootstrap (DAB) framework.
result Establishes theoretical coverage results for DAB methods.
Fed-TDA augments federated tabular data to improve performance and privacy.
problem Non-IID data challenges in federated learning.
method Synthesizes tabular data using simple statistics for augmentation.
result Fed-TDA improves test performance and communication efficiency.
This review article surveys data augmentation MCMC algorithms.
problem Sampling from intractable probability distributions.
method Comprehensive study of DA MCMC algorithms, their convergence properties, and acceleration strategies.
result Synthesizes recent developments and provides insights for researchers.
Data augmentation impacts adversarial risk; careful application recommended.
problem Understanding how data augmentation affects adversarial risk in deep learning.
method Empirical analysis using three measures of adversarial risk.
result Data augmentation does not always improve adversarial risk; augmented data influences models more.
Develops Bayesian inference methods for gamma models.
problem Challenges in inference for models with gamma functions.
method Data augmentation scheme using Exponential Reciprocal Gamma distributions.
result Scalable EM and MCMC algorithms developed.
Anomaly detection is the process of finding data points that deviate from a baseline. In a real-life setting, anomalies are usually unknown or extremely rare. Moreover, the detection must be accomplished in a timely manner or the risk of corrupting the system might grow exponentially. In this work, we propose a two lev…
This work analyzes the role of data augmentation in self-supervised learning using RKHS approximation and regression.
problem Limited theoretical understanding of the role of data augmentation in self-supervised learning.
method Geometric characterization of the target function given by augmentation, proving generalization bounds.
result Two generalization bounds are derived, one free of model complexity, the other specific to near-optimal encoders.
The paper analyzes how data augmentation affects the test error in regression models.
problem Understanding the impact of data augmentation on the test error in regression models.
method Characterizes the test error in terms of population quantities and augmentation statistics.
result Provides a tight characterization of the test error in mean squared error.
Enhances machine learning models by preserving data structure, addressing statistical distortions.
problem Statistical distortions in synthetic data generated by Mixup.
method Proposes a generalized mixup method with a flexible weighting scheme to preserve data structure.
result Preserves statistical properties of original data while maintaining model performance.
Synthetic augmentation improves financial machine learning performance in variance-dominant regimes.
problem Data scarcity in financial machine learning.
method Formalized synthetic augmentation, introduced size-matched null augmentation, and developed a non-parametric block permutation test.
result Synthetic augmentation is beneficial only in variance-dominant regimes, such as persistent volatility forecasting.
New offline RL algorithms tackle partial data coverage with optimal performance and practicality.
problem Partial data coverage in offline RL datasets.
method Augmented Lagrangian method applied to MIS formulation for optimal offline RL.
result Statistically optimal offline RL with practical performance, eliminating conservatism.
The study analyzes how data augmentation helps isolate content from style in self-supervised learning.
problem Understanding how data augmentation affects the separation of content and style in self-supervised learning.
method Formulated a latent variable model with content and style components, studied identifiability of latent representation, and introduced a dataset to test the theory.
result Sufficient conditions for identifying the invariant content partition in self-supervised learning.
New statistical models for predicting ranked preferences from partial orders.
problem Statistical models overlook information in list length.
method Composite and augmented ranking models for joint modeling of partial orders and list lengths.
result Augmented ranking models best predict both length and preferences.
Study evaluates synthetic data augmentation for small datasets, highlighting inconsistencies in traditional metrics.
problem Inconsistent validation of synthetic data generated for small sample sizes.
method Proposes a normalized Bottleneck distance metric to evaluate synthetic tabular data.
result Common metrics like propensity scoring and MMD fail for small datasets, showing instability and high variability.
New methods use ML predictions to improve statistical inference.
problem Improving statistical inference using machine learning predictions.
method Prediction-Augmented Trees (PART, PAQ) for reliable statistical analysis.
result PART and PAQ outperform existing methods in various datasets.
We design and study a Contextual Memory Tree (CMT), a learning memory controller that inserts new memories into an experience store of unbounded size. It is designed to efficiently query for memories from that store, supporting logarithmic time insertion and retrieval operations. Hence CMT can be integrated into existi…
A new method uses Hamiltonian Monte Carlo for imputation and augmentation of healthcare data.
problem Missing values in clinical studies lead to biased results and loss of statistical power.
method Folded Hamiltonian Monte Carlo (F-HMC) with Bayesian inference to handle high-dimensional, small sample size datasets.
result The method enriches the quality of data in precision, accuracy, recall, F1 score, and propensity metric.
Paper proposes a data augmentation method for LLM-generated data in market research.
problem Bias in LLM-generated data in market research.
method Statistical data augmentation approach integrating LLM-generated and real data.
result Statistically robust estimators with reduced bias and cost savings.
Efficient synthetic data generation improves model performance on tabular data.
problem Improving model robustness and performance with scarce or low-quality data.
method Hardness characterization to identify high-value training points, generating synthetic data only from these points.
result Synthetic data generated from hardest points outperforms non-targeted methods on tabular datasets.
Synthetic augmentation helps but not always in imbalanced learning.
problem Imbalanced learning causes poor performance on rare classes.
method Developed a statistical framework for synthetic augmentation in imbalanced learning.
result Synthetic augmentation is not always beneficial and depends on the imbalance regime.
Efficiently solves Elastic Net in high dimensions with Newton method.
problem Feature selection in high-dimensional data with non-negligible collinearity.
method Semi-smooth Newton Augmented Lagrangian Method.
result Significantly reduces computational cost compared to competitors.
This chapter introduces quaternion machine learning for 3D rotations.
problem Lack of quaternion machine learning for 3D rotations.
method Augmented statistics, widely linear models, quaternion calculus, mean square estimation.
result Foundation for quaternion machine learning.
Statistical models with constrained probability distributions are abundant in machine learning. Some examples include regression models with norm constraints (e.g., Lasso), probit, many copula models, and latent Dirichlet allocation (LDA). Bayesian inference involving probability distributions confined to constrained d…
Proposes a method to improve CATE estimation by imputing missing potential outcomes.
problem Statistical discrepancy between distinct treatment groups in CATE estimation.
method Contrastive learning approach to reliably impute missing potential outcomes for a subset of individuals.
result Improves the accuracy and robustness of CATE estimation models.
LACD uses unlabeled data to improve conditional diffusion models.
problem Costly and time-consuming acquisition of labeled data.
method Label-augmented conditional diffusion (LACD) with joint denoising score matching.
result LACD converges faster in total variation and Wasserstein-1 distances with sufficient unlabeled data.
Machine Learning (ML) models are applied in a variety of tasks such as network intrusion detection or Malware classification. Yet, these models are vulnerable to a class of malicious inputs known as adversarial examples. These are slightly perturbed inputs that are classified incorrectly by the ML model. The mitigation…
Active inference framework improves U-statistic estimation efficiency.
problem Costly acquisition of labels for U-statistics. method Active inference framework with optimal sampling rule.
result Substantial gains in estimation efficiency over baseline methods.
Self-augmentation improves deep networks for few-shot learning with minimal training data.
problem Improving deep networks' generalization to unseen classes with limited training examples.
method Self-augmentation using self-mix and self-distillation techniques, combined with regional dropout and local representation learning.
result The method outperforms state-of-the-art few-shot learning methods on prevalent benchmarks.
The study examines mixing times of data-augmentation Gibbs samplers for high-dimensional probit regression.
problem Investigating convergence properties of data-augmentation samplers for Bayesian probit regression.
method Using recent results on Gibbs samplers for log-concave targets, the study provides non-asymptotic bounds on mixing times.
result Explicit non-asymptotic bounds on mixing times depend on design matrix and prior precision, holding uniformly over responses.
Variable selection is one of the most important tasks in statistics and machine learning. To incorporate more prior information about the regression coefficients, the constrained Lasso model has been proposed in the literature. In this paper, we present an inexact augmented Lagrangian method to solve the Lasso problem …
NoisyMix boosts model robustness to common corruptions.
problem Improving robustness of neural networks in real-world applications.
method NoisyMix training scheme that uses noisy augmentations in input and feature space.
result NoisyMix produces more robust models with well-calibrated class membership probabilities.
BYOL learns useful representations without batch statistics.
problem BYOL's reliance on batch statistics for representation learning.
method Training BYOL without batch normalization and using alternative normalization schemes.
result BYOL performance comparable to vanilla BYOL without batch normalization.
Proposes a two-stage method for estimating heterogeneous treatment effects using gradient boosting trees.
problem Estimating heterogeneous treatment effects in randomized clinical trials with high-dimensional predictive markers.
method Two-stage statistical learning procedure using gradient boosting trees (XGBoost) to estimate main effects and HTE.
result Improves efficiency in estimating heterogeneous treatment effects through nonparametric function estimation.
Researchers found that avoiding synthetic data generation prevents model collapse in machine learning.
problem Model collapse in machine learning where models degenerate over generations.
method Comparing discard and augment workflows, focusing on Linear Regression.
result Theoretical evidence shows that for Linear Regression, test risk is bounded by π²/6 of original data alone.
Paper proposes distributed optimization for federated learning with theoretical guarantees.
problem Privacy-preserving cross-organizational data collaboration in machine learning.
method Augmented Lagrangian technique for diverse communication topologies, termination criteria, and parameter update mechanisms.
result The proposed framework recovers classical optimization methods and provides strong performance in large-scale federated learning.
This work provides statistical guarantees for GANs that are invariant to certain group symmetries.
problem Learning group-invariant distributions efficiently.
method Study of group-invariant GANs and their performance guarantees.
result Group-invariant GANs require fewer samples and have a reduced discriminator approximation error.
Improves inference from sparse data with hybrid summary statistics.
problem Robust simulation-based inference from limited data.
method Augment traditional summary statistics with neural network outputs to maximize mutual information.
result Improves information extraction and makes inference robust in low-data settings.
Identifying a set of homogeneous clusters in a heterogeneous dataset is one of the most important classes of problems in statistical modeling. In the realm of unsupervised partitional clustering, k-means is a very important algorithm for this. In this technical report, we develop a new k-means variant called Augmented …
Improves AUC for disadvantaged groups by adding features.
problem Reducing cross-group differences in AUC for classification models.
method Feature augmentation to improve AUC for disadvantaged groups.
result Significantly improves AUC for disadvantaged groups.
This article is an overview of supervised machine learning problems for regression and classification. Topics include: kernel methods, training by stochastic gradient descent, deep learning architecture, losses for classification, statistical learning theory, and dimension independent generalization bounds. Implicit re…
Paper proposes MAST to identify stress conditions in forecasting models.
problem Improving reliability and transparency of univariate forecasting models under stress.
method Meta-learning and data augmentation approach to predict stress conditions.
result MAST identifies conditions leading to large errors in forecasting models.