A new method computes gradients without backpropagation.
problem Optimization of machine learning models.
method Forward mode automatic differentiation to compute gradients.
result Forward gradient is an unbiased estimate of the gradient, eliminating the need for backpropagation.
Paper introduces FoMoH for optimization without backpropagation.
problem Optimizing machine learning models without backpropagation.
method Second-order hyperplane search, forward-mode stochastic gradient method, hyper-dual numbers, FoMoH.
result Developed a novel optimization algorithm that avoids backpropagation.
We study two procedures (reverse-mode and forward-mode) for computing the gradient of the validation error with respect to the hyperparameters of any iterative learning algorithm such as stochastic gradient descent. These procedures mirror two methods of computing gradients for recurrent neural networks and have differ…
Backpropagation-free trunk training improves model performance on various benchmarks.
problem Memory inefficiency and noisy gradient estimates in deep network training.
method Split Forward Gradient (Split-FG) method that splits network into trunk and head, estimating only trunk gradient.
result Split-FG achieves better performance than pure forward-gradient training and backpropagation on various benchmarks.
Proposes a new learning method for RBMs that combines strengths of forward and reverse KLD.
problem Underfitting and mode-collapse issues in RBM learning.
method Ratio divergence learning using target energy.
result Significantly outperforms other learning methods in energy function fitting, mode-covering, and stability.
FDS tackles long horizon hyperparameter optimization issues.
problem Memory scaling and gradient degradation in long horizon tasks.
method Forward-mode differentiation with sharing (FDS).
result Significantly outperforms greedy gradient-based alternatives.
CT compares two distributions using Bayes' theorem and chain rule.
problem Measuring the difference between two probability distributions.
method Conditional transport (CT) using chain rule and Bayes' theorem.
result CT strikes a good balance between mode-covering and mode-seeking behaviors.
We quantify forgetting in post-training models, distinguishing mass and drift.
problem Understanding and preventing forgetting in post-training generative models.
method Developed theoretical results under a two-mode mixture abstraction, formalizing mass and drift forgetting.
result Forgetting can be precisely quantified based on divergence direction, geometric overlap, and training regime.
Improved KL divergence estimators for normalizing flows lead to faster convergence and better approximations.
problem Estimating KL divergences for normalizing flows efficiently and accurately.
method Path-gradient estimators for reverse and forward KL divergences.
result Path-gradient estimators lead to faster convergence and better approximation results.
Geometric AD framework simplifies derivative computation in JAX.
problem Efficient and accurate automatic differentiation.
method Jet functors and Weil algebras for geometric analysis.
result Unified view of derivative propagation with algebraic exactness.
Forward gradients improve neural network training without backpropagation issues.
problem Training neural networks without backpropagation's locking and memorization problems.
method Using directional derivatives in forward differentiation mode, with biased guesses based on feedback from small auxiliary networks.
result Using gradients from a local loss as a candidate direction improves Forward Gradient methods.
The paper proposes a method to compute higher infinitesimals in numerical and symbolic analysis.
problem Computing higher-order derivatives with higher infinitesimals.
method Automatic differentiation in terms of C-infinity rings and Weil algebras.
result A unifying theoretical framework for multivariate higher-order derivatives.
This paper improves non-asymptotic bounds for denoising diffusions, focusing on the Ornstein-Uhlenbeck process.
problem Improving non-asymptotic bounds for denoising diffusions, especially for the Ornstein-Uhlenbeck process.
method Explicit non-asymptotic bounds on forward diffusion error in total variation, considering multi-modal data distributions.
result The Ornstein-Uhlenbeck process cannot be significantly improved in terms of reducing terminal time T for multi-modal data distributions. Bayesian neural networks benefit from fully marginalizing over all modes to improve generalization.
problem Bayesian neural networks suffer from multimodal posterior distributions that can lead to suboptimal generalization.
method Use appropriate Bayesian sampling tools to fully marginalize over all posterior modes.
result Training with full marginalization improves the ability of the network to reason between multiple candidate solutions.
Transformers can emulate various algorithms by prompting, proving universality.
problem How to emulate algorithms using fixed-weight Transformers.
method Two modes of in-context algorithm emulation: task-specific and prompt-programmable. Constructing prompts that encode algorithm parameters into token representations.
result Fixed-weight Transformers can emulate a broad class of algorithms via prompts.
Push-forward models struggle to fit multimodal distributions due to high Lipschitz constants.
problem Expressivity of push-forward generative models in fitting multimodal distributions.
method Analyzing the Lipschitz constant and its relation to the total variation distance and Kullback-Leibler divergence.
result Push-forward models require high Lipschitz constants to approximate multimodal distributions, leading to a trade-off between expressivity and stability.
BDMBC clusters data with varying densities using a new PLLS measure.
problem Finding clusters with varying densities in data.
method Bagged k-distance with PLLS for mode estimation. result BDMBC achieves optimal convergence rates for mode and level set estimation.
This research examines the geometry of latent spaces in push-forward generative models.
problem Tendency of deep generative models to output samples outside target distribution support.
method Geometric measure theory and truncation method to enforce simplicial cluster structure.
result Proves sufficient condition for optimality in latent space geometry.
Neural nets improve plasma equilibrium modeling for NSTX-U.
problem Modeling plasma equilibrium and shape control for NSTX-U.
method Developed two neural networks: Eqnet and Pertnet.
result NNs offer faster and more flexible prediction of plasma scenarios.
Dropout neural networks can approximate any function with high probability.
problem Approximating functions with dropout neural networks.
method Two universal approximation theorems for dropout neural networks in random and deterministic modes.
result Dropout neural networks can approximate any function in probability and in Lq. This paper is concerned with the numerical solution of model-based, Bayesian inverse problems. We are particularly interested in cases where the cost of each likelihood evaluation (forward-model call) is expensive and the number of un- known (latent) variables is high. This is the setting in many problems in com- putat…
ACA method improves gradient estimation for neural ODEs, reducing error and training time.
problem Inaccurate gradient estimation methods hinder the performance of neural ODEs on benchmark tasks.
method Adaptive Checkpoint Adjoint (ACA) method that applies trajectory checkpointing, deletes redundant components, and supports adaptive solvers.
result ACA reduces error rate by half and training time by half compared to adjoint and naive methods on image classification tasks.
New method uses neural networks to forecast spatial-temporal data.
problem Probabilistic forecasting of spatio-temporal data with causal structure.
method MMAF-guided learning with ensemble of stochastic feed-forward neural networks.
result Forecasting remains calibrated across multiple time horizons.
The need to efficiently calculate first- and higher-order derivatives of increasingly complex models expressed in Python has stressed or exceeded the capabilities of available tools. In this work, we explore techniques from the field of automatic differentiation (AD) that can give researchers expressive power, performa…
New method uses MMAF-guided learning for spatio-temporal probabilistic forecasts.
problem Probabilistic forecasting of spatio-temporal data with causal structure.
method Generalized Bayesian methodology, MMAF-guided learning, ensemble of stochastic feed-forward neural networks.
result Forecast performance comparable to, and sometimes better than, deep learning architectures.
TREK uses distillation to help students solve hard problems.
problem Stalled progress on hard prompts when current policy lacks useful reasoning trajectories.
method TREK combines distillation and reinforcement learning to expand student support.
result TREK significantly improves student performance on mathematical reasoning and agentic tasks.
Periodic surfaces have a limited number of bending modes, equal to their membrane modes.
problem Understanding the limitations of bending modes in periodic surfaces.
method Analyzing deformation modes of periodic, piecewise smooth, simply connected surfaces.
result Effective membrane modes and bending modes are orthogonal, limiting the total number of modes to 3.
Sparse-mode DMD disambiguates local and global modes in spatiotemporal data.
problem Disambiguating local and global modes in spatiotemporal data.
method Sparse-mode DMD with sparsity-promoting regularization.
result Explicitly constructs discrete and continuous spectra.
A new training method speeds up ResNet training by 3x with minimal accuracy loss.
problem Training ResNets is slow due to dependencies between modules.
method Serial-parallel hybrid training strategy with data augmentation and downsampling.
result Significant speedup over traditional methods with comparable accuracy.
Continuous Hidden Markov Models for Equity Returns
problem Generating synthetic equity returns that match real return characteristics
method Continuous Hidden Markov Models
result Recovered volatility clustering and narrowed kurtosis gap
Motivated by advantages of current-mode design, this brief contribution explores the implementation of weight matrices in neuromemristive systems via current-mode memristor crossbar circuits. After deriving theoretical results for the range and distribution of weights in the current-mode design, it is shown that any we…
We propose a fast second-order method that can be used as a drop-in replacement for current deep learning solvers. Compared to stochastic gradient descent (SGD), it only requires two additional forward-mode automatic differentiation operations per iteration, which has a computational cost comparable to two standard for…
In recent years, deep learning researchers have focused on how to find the interpretability behind deep learning models. However, today cognitive competence of human has not completely covered the deep learning model. In other words, there is a gap between the deep learning model and the cognitive mode. How to evaluate…
For dynamical systems that can be modelled as asymptotically stable linear systems forced by Gaussian noise, this paper develops methods to infer or estimate their modes from observations in real time. The modes can be real or complex. For a real mode, we wish to infer its damping rate and mode shape. For a complex mod…
Characterizes neutral deformation modes of minimal surfaces.
problem Understanding the energy content of deformation modes of minimal surfaces.
method Analyzes the energy content of stretching, drilling, and bending modes of minimal surfaces.
result All isometries of a minimal surface are globally neutral and give rise to soft elasticity.
Mathematical analysis shows annealing prevents mode collapse in Gaussian mixtures.
problem Mode collapse in variational inference for multimodal distributions.
method Analyzed annealing strategies for Gaussian mixtures, derived formulas, and tested on neural networks.
result Appropriately chosen annealing schemes can robustly prevent mode collapse.
Parsimonious Dynamic Mode Decomposition selects sparse modes robustly.
problem Manual tuning of sparsity parameters in traditional DMD.
method Time-delay embedding and Orthogonal Matching Pursuit.
result Autonomously determines optimally sparse subset of modes.
Geodesics connect model modes in neural network loss landscapes.
problem Connecting modes in neural network loss landscapes.
method Reframed mode connectivity in Information Geometry, hypothesized geodesics as mode-connecting paths, proposed algorithm to approximate geodesics.
result Geodesics achieve mode connectivity in neural networks.
A new, efficient k-modes algorithm improves clustering of categorical data.
problem Clustering categorical data using existing methods like k-means is inefficient. method Developed a novel k-modes algorithm called OTQT, which improves on existing methods. result OTQT finds more accurate clusters per iteration and is faster overall.
Multimodal clustering is an unsupervised technique for mining interesting patterns in n-adic binary relations or n-mode networks. Among different types of such generalized patterns one can find biclusters and formal concepts (maximal bicliques) for 2-mode case, triclusters and triconcepts for 3-mode case, closed $n…
A geometric account explains why 'The Dress' is ambiguous, predicting observable signatures in image processing.
problem Understanding and predicting ambiguity in image processing, particularly in intrinsic image decomposition.
method Geometric analysis of intrinsic image decomposition, focusing on the discontinuous switch in prior-mode sections.
result Predicted signatures in albedo Jacobian and Fernet curvature can be observed in various models and datasets.
EDLP samples flat modes in discrete spaces using entropy.
problem Sampling flat modes in discrete spaces is challenging.
method EDLP uses a continuous auxiliary variable and local entropy to guide sampling.
result EDLP consistently outperforms traditional methods in various tasks.
Proposes a Gaussian process for Koopman mode decomposition.
problem Estimating Koopman mode decomposition quantities and latent variables.
method Unsupervised Gaussian process for simultaneous estimation.
result Efficient parameter estimation through low-rank approximations.
Deep learning helps remove secondary B-mode polarization to detect primordial gravitational waves.
problem Removing secondary B-mode polarization from CMB data to detect primordial gravitational waves. method Applied deep learning (ResUNet-CMB) to estimate and remove multiple sources of secondary B-mode polarization. result Deep learning can produce nearly optimal, unbiased estimates of the amplitude of primordial gravitational waves.
A new method for continual learning in GANs learns new modes with limited data.
problem Learning new target modes with limited samples while preserving previously learned ones.
method Mode-affinity score for generative modeling, generator replay, and weighted label generation.
result Gains over state-of-the-art methods, even with fewer training samples.
Empirical study shows GANs overfit and drop modes when training is deterministic.
problem Understanding overfitting and mode drop in GAN training.
method Empirical analysis of GAN training with and without stochasticity.
result GANs overfit and drop modes when training is deterministic.
The paper finds shape modes for vortices in a specific sigma model.
problem Existence of internal modes in CP1 vortices. method Developed a geometric formalism based on the Bogomol'nyi decomposition of the energy functional.
result Proved the existence of at least one shape mode for a general CP1 vortex solution. Dynamic Mode Decomposition (DMD) yields a linear, approximate model of a system's dynamics that is built from data. We seek to reduce the order of this model by identifying a reduced set of modes that best fit the output. We adopt a model selection algorithm from statistics and machine learning known as Least Angle Reg…