Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

2785558331,110 · Jun 202019922001200920182026
48 results for Training Framework

Solve-training trains neural nets to map physical solutions efficiently.

problem Representing complex physical solutions with neural networks.
method Variational training using loss functions from physical models.
result Effective neural network representation of solution maps without expensive labels.

Framework for few-shot relation classification with minimal training data.

problem Few-shot relation classification with limited training data.
method Meta-learning framework that combines instance and support knowledge.
result Framework outperforms state-of-the-art results and achieves competitive performance with large training data.

Unified framework for supervised classification with diverse training data.

problem Handling different types of training data for supervised classification.
method Generalized robust risk minimization (GRRM) with probabilistic transformations.
result GRRM can handle various training data types and new supervision schemes.

Develops a robust training framework to detect backdoor attacks in DNNs.

problem Vulnerability of DNNs to backdoor attacks by poisoned training data.
method Collider framework selects prominent samples based on geometric structures and coreset selection objective.
result Significantly reduces backdoor success rate in various poisoned datasets.

PaRoT simplifies robust training for deep neural networks.

problem Training deep neural networks to be robust to small input changes.
method Developed a practical framework on TensorFlow for robust training without code modifications.
result PaRoT's performance is comparable to existing methods and is easy to use on real-world models.

A framework for training deep networks in Apache Spark.

problem Expensive and time-consuming training of deep networks with large data and model parameters.
method Data and model-parallel, distributed training over Apache Spark clusters.
result Significant speedup and scalability for deep network training.

Paper develops a statistical framework for quantized training of deep neural networks.

problem Lack of theoretical understanding of gradient quantization in FQT.
method Presented a statistical framework for analyzing FQT algorithms, viewing quantized gradient as a stochastic estimator of QAT gradient.
result Developed two novel gradient quantizers with smaller variance than existing per-tensor quantizer.

Paper tackles model vulnerabilities by reconstructing training data.

problem Reconstructing training data from model parameters poses a security risk.
method Developed a mathematical framework and score matching method for both Bayesian and non-Bayesian models.
result First score matching framework for reconstructing data in Bayesian models.

This paper proposes a curriculum learning framework for NMT to reduce training time and improve performance.

problem Slow training and need for heuristics in NMT systems.
method A curriculum learning framework that decides training samples based on estimated difficulty and model competence.
result Up to 70% decrease in training time and up to 2.2 BLEU accuracy improvements.

Framework infers conservation laws from trained neural networks.

problem Building reduced models of complex systems from physical data.
method Derives conservation laws from symmetries of dynamics in trained DNNs using Noether's theorem.
result Consistent results with previous studies for metastable collective motion systems.

ZipML framework trains models at low precision with provable guarantees and significant speedups.

problem Training machine learning models at low precision to achieve speedups and maintain accuracy.
method ZipML framework, double sampling, variance-optimal stochastic quantization, approximation to non-linear models.
result Training at low precision with ZipML framework achieves up to 6.5x speedups and maintains accuracy.

We propose a new framework for estimating generative models via an adversarial process, in which we simultaneously train two models: a generative model G that captures the data distribution, and a discriminative model D that estimates the probability that a sample came from the training data rather than G. The training…

2014-06-10abs ↗pdf ↗

New framework minimizes model complexity for improved few-shot learning.

problem Empirical benefits of pre-training scale with data size but lack theoretical explanation.
method Complexity Minimization framework for meta-representation learning.
result Theoretical analysis shows error rate improves with more meta-training data.

Chicle tackles elastic machine learning training by avoiding micro-tasks.

problem Elasticity and load balancing in distributed machine learning training.
method Chicle is a new elastic distributed training framework that exploits machine learning algorithms to implement elasticity and load balancing without micro-tasks.
result Chicle achieves performance competitive with state-of-the-art rigid frameworks while enabling elastic execution and dynamic load balancing.

ENSURE framework trains deep image recon algorithms without clean data.

problem Lack of clean, fully sampled ground-truth data for deep learning image reconstruction.
method Introduces ENSURE framework, a generalization of SURE and GSURE to random sampling patterns.
result ENSURE loss function is an unbiased estimate for true mean-square error.

New framework improves LLM performance by avoiding forgetting during sequential training stages.

problem Forgetting during sequential training stages of LLMs.
method Proposes a joint post-training framework with theoretical convergence guarantees.
result Empirically outperforms sequential post-training framework by up to 23%.

A new framework trains RBMs deterministically for unsupervised learning.

problem Training and evaluation of RBMs with weak interactions.
method TAP mean-field approximation for generalized latent-variable models.
result Effective deterministic training and interesting unsupervised learning features demonstrated.

This work proposes a new adversarial training method based on L2L framework.

problem Training robust neural networks against adversarial attacks.
method Generic learning-to-learn (L2L) framework to learn an optimizer and a robust classifier.
result L2L outperforms existing adversarial training methods in classification accuracy and computational efficiency.

POLAR framework interprets word embeddings using polar opposites.

problem Lack of interpretability in pre-trained word embeddings.
method Adopt semantic differentials and polar opposites to transform embeddings.
result Interpretable word embeddings maintain performance comparable to original embeddings.

We present a novel view that unifies two frameworks that aim to solve sequential prediction problems: learning to search (L2S) and recurrent neural networks (RNN). We point out equivalences between elements of the two frameworks. By complementing what is missing from one framework comparing to the other, we introduce a…

2016-07-18abs ↗pdf ↗

Enhances robustness of AT frameworks to multiple perturbations without increasing training complexity.

problem Defending against the union of multiple perturbations in adversarial training.
method SNAP technique that augments a network with shaped noise to enhance robustness.
result 14%-to-20% improvement in adversarial accuracy for ResNet-18 on CIFAR-10.

Framework improves gradient estimation for faster training convergence.

problem Efficiently estimating noisy gradients in stochastic optimization.
method Dynamic adaptive importance sampling combining multiple distributions.
result Adaptively weighted multiple importance sampling yields superior gradient estimates.

Develops a framework for parallel and distributed neural network training.

problem Training neural networks in a distributed environment with sparse connectivity.
method Customizes a non-convex optimization framework over networks, including dynamic consensus and parallel optimization.
result Guarantees convergence to a stationary solution under mild assumptions.

A Bayesian framework models adversarial uncertainty for robust machine learning.

problem Vulnerability of machine learning models to adversarial attacks.
method Formal Bayesian framework that models adversarial uncertainty through a stochastic channel, articulating probabilistic assumptions.
result Explicitly modeling adversarial uncertainty leads to improved robustification strategies.

Paper develops a theory explaining contrastive pre-training for multimodal AI.

problem Limited theoretical understanding of contrastive pre-training for multi-modal AI.
method Introduces approximate sufficient statistics and Joint Generative Hierarchical Model.
result Near-minimizers of contrastive loss are approximately sufficient, enabling diverse downstream tasks.

Unsupervised pre-training improves model generalization, but lacks theoretical understanding.

problem Lack of theoretical understanding of unsupervised pre-training's impact on model generalization.
method Introduces a novel theoretical framework to analyze and enhance generalization.
result Enhances understanding of unsupervised pre-training and fine-tuning, proposing a new regularization method.