Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

3266519771,302 · Jun 202019922001200920182026
48 results for training improvement

Paper improves adversarial training using a learned optimizer.

problem Improving robustness of deep learning models against adversarial attacks.
method Empirically identified PGD attack's limitations and used a learning-to-learn framework to train an adaptive inner optimizer.
result The proposed framework consistently improves model robustness over traditional adversarial training methods.

Self-training outperforms pre-training on COCO object detection and segmentation datasets.

problem The effectiveness of pre-training in improving object detection and segmentation models is limited.
method Investigated self-training as an alternative method to utilize additional data.
result Self-training consistently improves model performance across various dataset sizes and data augmentation levels.

FIRE PBT improves neural network training by focusing on long-term performance.

problem Greedy decision mechanisms in PBT lead to poor long-term performance.
method FIRE PBT uses a fitness metric to encourage long-term performance over short-term improvements.
result FIRE PBT outperforms PBT on ImageNet and matches hand-tuned learning rates.

Adaptive networks improve model robustness through conditional normalization.

problem Limited robustness of adversarial-trained networks due to network capacity and training samples.
method Proposes a conditional normalization module to adapt networks during adversarial training.
result Adaptive networks outperform both clean validation accuracy and robustness compared to non-adaptive counterparts.

We improve private training accuracy with learning rate schedules and matrix factorizations.

problem Private training with learning rate schedules and correlated noise.
method General upper and lower bounds for learning rate schedules, memory-efficient constructions, and schedule-aware factorizations.
result Schedule-aware factorizations improve accuracy in private training.

Training on mixed distributions improves test performance even when components are unrelated.

problem Improving test performance with mismatched training and test distributions.
method Analyzing mixture distributions with different training and test proportions.
result Distribution shift can be beneficial, improving test performance even when components are unrelated.

Improved neural network robustness with instance-specific perturbation margins.

problem Adversarial training fails to generalize well to unperturbed test set.
method Instance adaptive adversarial training with sample-specific perturbation margins.
result Test accuracy improves with a marginal drop in robustness.

Pruning improves model generalization in over-parameterized models, contradicting traditional theories.

problem Pruning's effect on generalization in over-parameterized models.
method Empirical study on standard pruning algorithms and additional regularization effects.
result Pruning leads to better training and regularization, improving generalization.

Generative models improve adversarial robustness by adding synthetic data.

problem Improving robustness in machine learning models trained on limited data.
method Using synthetic data generated from a large dataset to augment the original training set.
result Generative models can significantly reduce the robust-accuracy gap compared to models trained with additional real data.

SPAT improves adversarial robustness by preserving semantics in adversarial training.

problem Adversarial examples often have different semantics than original data, introducing unintended biases.
method Semantics-preserving adversarial training (SPAT) that encourages pixel perturbation shared among all classes.
result SPAT improves adversarial robustness and achieves state-of-the-art results in CIFAR-10 and CIFAR-100.

Improves information cascade models using contrastive training and DSTs.

problem Improving models of information cascades using limited labeled data.
method Proposes a contrastive training procedure for models of information cascades as directed spanning trees (DSTs).
result Unsupervised training with additional content features achieves significantly better results, reaching half the accuracy of a fully supervised model.

Rob-GAN combines generator, discriminator, and adversarial attack for improved robustness and quality.

problem Improving robustness and quality of GAN-generated images under adversarial attacks.
method Rob-GAN framework that jointly optimizes generator and discriminator in the presence of adversarial attacks.
result Rob-GAN improves convergence speed, image quality, and robustness of discriminators under strong adversarial attacks.

AGMMNs improve learning of copula models by adaptively selecting kernels.

problem Learning dependence structures in copula models.
method Adaptive bandwidth selection for MMD in GMMNs, increasing kernels based on validation loss.
result AGMMNs significantly improve training performance over GMMNs and parametric models.

Meta-learning improves DNN generalization on standard supervised learning.

problem Improving deep neural networks' generalization without adding more parameters.
method MLTP simulates meta-training by considering a batch of samples as a task, optimizing for both current and new tasks.
result MLTP consistently improves DNN generalization across various sizes and datasets.

Improves sample quality of generative models using energy-based methods.

problem Low sample quality in generative models.
method Constructs an energy function on latent space, trains an energy-based model, and generates improved samples.
result Significant improvement in sample quality with minimal computational overhead.

We simplify diffusion models by defining a design space and improving FID scores.

problem Complexity in diffusion-based generative models.
method Defined a design space, improved sampling and training processes, and pre-conditioned score networks.
result Improved FID scores of 1.79 for CIFAR-10 and 1.97 for unconditional settings.

Cost-effective method improves and re-purposes pre-trained GANs by fine-tuning class-embeddings.

problem Fine-tuning BigGANs from scratch is impractical due to instability and high computational cost.
method Fine-tuning only the class-embedding layer of pre-trained GANs.
result Significantly improved realism and diversity of samples, re-purposed for new tasks, and de-biased or improved diversity.

Joint training improves model accuracy by selectively using privileged information.

problem Two-stage training can lead to model failure with noisy privileged information.
method Joint training of two models to use privileged information selectively.
result Joint training outperforms two-stage baselines on synthetic and real-world tasks.

Improved disentanglement of data factors using recursive training.

problem Current unsupervised disentanglement methods are inconsistent and fail to achieve levels of disentanglement seen in supervised approaches.
method Introduced PBT for VAEs, used UDR for heuristic scoring, and developed recursive rPU-VAE approach.
result Recursive training leads to robust disentanglement of data factors across multiple datasets.

Noise improves model quality in non-linear neural networks during decentralized training.

problem Improving generalization of locally trained neural networks.
method Injecting noise into the weights of neural networks during decentralized training.
result Noise injection improves model quality for non-linear neural networks, but not for linear models.