Self-training with noisy student-teacher boosts keyword spotting accuracy.
problem Robust keyword spotting in challenging conditions.
method Aggressive data augmentation and self-training with noisy student-teacher approach.
result Significant accuracy improvement in difficult conditions, up to 60%.
LACD uses unlabeled data to improve conditional diffusion models.
problem Costly and time-consuming acquisition of labeled data.
method Label-augmented conditional diffusion (LACD) with joint denoising score matching.
result LACD converges faster in total variation and Wasserstein-1 distances with sufficient unlabeled data.
LcGAN generates synthetic CT images for hemorrhagic lesion segmentation.
problem Scarce training data for hemorrhagic lesion segmentation.
method Lesion conditional Generative Adversarial Network (LcGAN) for synthetic image generation.
result Segmentation improved by 12.8% with synthetic data augmentation.
AugMask trains diffusion models on incomplete tabular data by augmenting missing values and applying denoising supervision.
problem Training diffusion models on incomplete tabular data with missing values.
method AugMask uses stochastic augmentation and denoising supervision to adapt diffusion models to incomplete data.
result AugMask enables diffusion-based tabular generators to outperform specialized missing-aware baselines across various datasets and missingness regimes.
We study the relationship between Ng's abelian cord ring and SL(2,C) characters of the two-fold branched cover Σ(K). Augmentations, and their corresponding rank, play a central role in the relationship. Our study also leads to a correspondence between trace-free SL(2,C) characters of a knot complement and augmentatio…
Study examines how data augmentation impacts optimization in linear regression.
problem Understanding how data augmentation schedules affect optimization in linear regression.
method Analyzed the effect of augmentation on optimization in linear regression with MSE loss, using classical convex optimization and recent work on implicit bias.
result Proved that under certain joint schedules for learning rate and augmentation scheme, augmented gradient descent converges and characterized the resulting minimum.
A connection between holomorphic and generating family invariants of Legendrian knots is established; namely, that the existence of a ruling (or decomposition) of a Legendrian knot is equivalent to the existence of an augmentation of its contact homology. This result was obtained independently and using different metho…
This work investigates image augmentations for GAN training, improving image quality.
problem Improving the accuracy and robustness of GAN models for image synthesis.
method Systematic study of various image augmentation techniques for GAN training.
result Vanilla GANs can achieve state-of-the-art generation quality with image augmentations.
Improves tabular data augmentation for contrastive learning.
problem Ineffective augmentation techniques for tabular data.
method Class-conditioned and feature-correlation based augmentation.
result Consistently outperforms conventional corruption methods.
Data augmentation improves keyword spotting accuracy in noisy conditions.
problem Maintaining low false reject rates in far-field KWS with playback interference.
method Artificially corrupted training data with mixed music and TV audio.
result 30-45% reduction in false reject rates under audio playback.
SapAugment learns adaptive augmentation policies for better model training.
problem Fixed data augmentation methods often apply the same augmentation to all samples, ignoring sample difficulty.
method SapAugment adapts augmentation parameters based on training loss, learning a sample-adaptive policy.
result SapAugment achieves up to 21% relative reduction in word error rate on LibriSpeech dataset.
Data augmentation impacts adversarial risk; careful application recommended.
problem Understanding how data augmentation affects adversarial risk in deep learning.
method Empirical analysis using three measures of adversarial risk.
result Data augmentation does not always improve adversarial risk; augmented data influences models more.
SpaceGAN enhances geospatial data using deep learning.
problem Challenges in modeling geospatial data with deep learning.
method SpaceGAN, a generative model that learns spatial structures through conditioning on neighbours.
result SpaceGAN produces synthetic samples faithful to real spatial patterns, improving data augmentation and model generalization.
Paper proposes MAST to identify stress conditions in forecasting models.
problem Improving reliability and transparency of univariate forecasting models under stress.
method Meta-learning and data augmentation approach to predict stress conditions.
result MAST identifies conditions leading to large errors in forecasting models.
Proposes T-CGAN for generating time series data with irregular sampling.
problem Generating time series data with irregular sampling and noise.
method Conditional Generative Adversarial Network (CGAN) with deconvolutional and convolutional neural networks, conditioned on timestamps.
result T-CGAN-generated time series perform as well as real data for classification tasks.
Proves accuracy guarantees for self-supervised learning with correlated positive pairs.
problem Lack of theoretical guarantees for self-supervised learning with correlated positive pairs.
method Novel augmentation graph concept and spectral decomposition loss.
result Provably accurate features under linear probe evaluation.
Proposes PG-DA for Bayesian MMNL estimation to handle non-conjugacy.
problem Non-conjugacy in the Bayesian estimation of MMNL models.
method Pólygamma data augmentation technique applied to MMNL estimation.
result Similar posterior estimates for binary choice scenarios, but empirical identification issues for J≥3 alternatives. Enhances RL in target domains with limited data using augmented return.
problem Utilize data from an accessible source domain to improve policy learning in a target domain with scarce data.
method Return Augmented Decision Transformer (REAG) method, which augments the return in the source domain to align with the target domain's optimal trajectory distribution.
result The proposed REAG method achieves the same level of suboptimality as without a dynamics shift, enhancing DT type frameworks' performance in off-dynamics RL.
Autoencoders improve anomaly detection by deriving features and augmenting data.
problem Poor performance of one-class classifiers in high-dimensionality and sparsity.
method Uses autoencoders to derive meaningful latent variables and augment data for OCC training.
result Enhances OCC algorithms' performance and outperforms other methods.
Framework improves target domain prediction using quantile matching.
problem Improving prediction accuracy in data-scarce target domains.
method Conditional quantile matching for distributional alignment.
result Empirical risk minimizer achieves tighter excess risk bound.
CoDA augments data with counterfactuals from local causal structures.
problem Improving sample efficiency in RL with complex dynamic processes.
method Local causal models (LCMs) and Counterfactual Data Augmentation (CoDA).
result CoDA significantly improves RL agent performance in locally factored tasks.
BAGAN balances imbalanced datasets for better image classification.
problem Imbalanced datasets negatively affect deep-learning classifier accuracy.
method Balancing GAN (BAGAN) trains on all available images, balancing majority and minority classes.
result BAGAN generates high-quality images for minority classes, improving classification accuracy.
The paper studies Lipschitz equivalence of self-similar sets and their augmented trees.
problem Lipschitz equivalence of self-similar sets and their boundaries.
method Introducing simple augmented trees and using combinatorial devices to show Lipschitz equivalence.
result Lipschitz equivalence of self-similar sets and their boundaries.
Enhances nighttime vehicle detection using style transfer and augmentation.
problem Nighttime object detection challenges due to lack of lighting and glare.
method Day-to-night style transfer and labeling-free augmentation with CARLA synthetic data.
result Significant improvements in nighttime vehicle detection with YOLO11 model.
YuruGAN generates yuru-chara images using GANs and clustering for small datasets.
problem Generating high-quality yuru-chara images with limited data.
method Class conditional GAN with clustering and data augmentation.
result Improved quality of generated yuru-chara images through clustering and data augmentation.
New approach uses text generation to boost AI agent development.
problem Lack of training data hinders AI agent development.
method Used encoder-decoder generative models, focusing on conditional variational auto-encoders.
result Significantly improved AI agent performance in low-resource cases.
CEB enhances model resilience through simple entropy bottleneck.
problem Improving model robustness against adversarial attacks.
method Conditional Entropy Bottleneck (CEB) combined with data augmentation.
result CEB significantly boosts adversarial robustness on various benchmarks.
The study analyzes how data augmentation helps isolate content from style in self-supervised learning.
problem Understanding how data augmentation affects the separation of content and style in self-supervised learning.
method Formulated a latent variable model with content and style components, studied identifiability of latent representation, and introduced a dataset to test the theory.
result Sufficient conditions for identifying the invariant content partition in self-supervised learning.
New theory explains contrastive learning via overlapping augmented views.
problem Lack of theoretical understanding of contrastive learning.
method Augmentation overlap perspective to improve downstream performance.
result Asymptotically closed bounds for downstream performance under weaker assumptions.
Generative Adversarial Networks improve neural network performance in low-data scenarios.
problem Low data limits neural network performance; standard data augmentation is limited.
method Trains a generative model on existing data to generate new plausible data items.
result Generative Adversarial Networks (DAGAN) significantly enhance neural network accuracy in low-data scenarios.
Improved cross-entropy estimator for likelihood-free inference.
problem Efficient inference for complex models with intractable likelihoods.
method Use neural networks as surrogate models and augment training data with joint likelihood ratio and score.
result New cross-entropy estimator provides improved sample efficiency.
This paper investigates how data augmentation improves linear separation of manifold data.
problem Understanding how data augmentation enhances linear separation of manifold data.
method Investigates the conditions under which self-supervised representations can linearly separate multi-manifold data.
result Self-supervised learning can linearly separate manifolds with a smaller distance than unsupervised learning.
We study a new bordification of the decorated Teichmüller space for a multiply punctured surface F by a space of filtered screens on the surface that arises from a natural elaboration of earlier work of McShane-Penner. We identify necessary and sufficient conditions for paths in this space of filtered screens to yield …
A new method tackles nonconvex optimization with penalties and proximal terms.
problem Nonconvex optimization problems with equality and inequality constraints.
method Inexact proximal augmented Lagrangian method (P-ALM) with adaptive penalty and proximal parameters.
result Effective convergence properties and numerical superiority over traditional methods.
GANs generate new traffic sign images to improve recognition accuracy.
problem Lack of data limits SqueezeNet's performance in traffic sign recognition.
method Applied pix2pix GANs to translate symbolic sign images to real ones for data augmentation.
result Data augmentation with GANs increased classification accuracy for traffic signs.
MicAugment transfers audio style from few seconds of input to match target conditions.
problem Audio model robustness to diverse acquisition conditions.
method Identifies and applies transformations learned from target audio to input audio.
result MicAugment significantly improves model robustness in downstream tasks.
Generative adversarial network improves spectrum sensing accuracy.
problem Lack of sufficient and adaptable training data for spectrum sensing.
method Generative adversarial network (GAN) for data augmentation and adaptation.
result Training data augmentation significantly increases classifier accuracy.
Diffusion models optimize objectives similar to ELBO with Gaussian noise augmentation.
problem Optimizing diffusion models for high perceptual quality.
method Showed diffusion objectives are weighted ELBOs over noise levels, with Gaussian noise augmentation.
result Diffusion objectives equate to ELBO with Gaussian noise augmentation under monotonic weighting.
GeoECG augments ECG data to improve heart disease detection.
problem Insufficient labeled ECG data and vulnerability to adversarial attacks.
method Wasserstein geodesic perturbation for data augmentation.
result Improved accuracy and robustness in ECG-based heart disease detection.
Proposes a method to use causal graph knowledge for better predictive modeling.
problem Lack of effective ways to incorporate causal graph knowledge into predictive models.
method Model-agnostic data augmentation method exploiting CI relations encoded in causal graphs.
result Improves prediction accuracy, especially in small-data scenarios.
Generative data augmentation boosts learning performance in various tasks.
problem Theoretical understanding of generative data augmentation's effect.
method Established a stability bound for non-i.i.d. settings, analyzed Gaussian mixture models and generative adversarial nets.
result Generative data augmentation can improve learning guarantees, especially in small train sets.
Generative adversarial networks improve brain-computer interface performance with limited data.
problem Limited training samples in brain-computer interfaces.
method Conditional Deep Convolutional Generative Adversarial Networks (cDCGAN) for data augmentation.
result Generated artificial EEG data improves classification accuracy in brain-computer interface tasks.
A new framework for systematic graph neural network data augmentation.
problem Diversity and difficulty in choosing graph neural network data augmentation techniques.
method Comprehensive framework capturing all previous RDAs, formal universality proof, automatic training method.
result Improved state of the art through new RDAs and impartial comparison.
Proposes a method to improve CATE estimation by imputing missing potential outcomes.
problem Statistical discrepancy between distinct treatment groups in CATE estimation.
method Contrastive learning approach to reliably impute missing potential outcomes for a subset of individuals.
result Improves the accuracy and robustness of CATE estimation models.
We propose an efficient algorithm for sparse signal reconstruction problems. The proposed algorithm is an augmented Lagrangian method based on the dual sparse reconstruction problem. It is efficient when the number of unknown variables is much larger than the number of observations because of the dual formulation. More…
Unified method for multi-defect microscopy image restoration with limited training data.
problem Challenges in applying deep learning methods due to limited training data for multi-defect microscopy images.
method Two-stage approach: data augmentation with GAN and conditional GAN training.
result Proposed method gives comparable or superior results to existing methods in image quality restoration.
We analyze the convergence behaviour of a recently proposed algorithm for regularized estimation called Dual Augmented Lagrangian (DAL). Our analysis is based on a new interpretation of DAL as a proximal minimization algorithm. We theoretically show under some conditions that DAL converges super-linearly in a non-asymp…
Augmented bridge matching preserves coupling information between distributions.
problem Preserving the original empirical pairing in flow and bridge matching processes.
method Augmenting the velocity field with initial sample point information.
result Simple modification recovers coupling information without losing Markovian property.