Backpropagation-free trunk training improves model performance on various benchmarks.
problem Memory inefficiency and noisy gradient estimates in deep network training.
method Split Forward Gradient (Split-FG) method that splits network into trunk and head, estimating only trunk gradient.
result Split-FG achieves better performance than pure forward-gradient training and backpropagation on various benchmarks.
We study the knot invariant called trunk, as defined by Ozawa, and the relation of the trunk of a satellite knot with the trunk of its companion knot. Our first result is trunk(K)≥n⋅trunk(J) where trunk(⋅) denotes the trunk of a knot, K is a satellite knot with companion J, and …
Compared to in-clinic balance training, in-home training is not as effective. This is, in part, due to the lack of feedback from physical therapists (PTs). Here, we analyze the feasibility of using trunk sway data and machine learning (ML) techniques to automatically evaluate balance, providing accurate assessments out…
Satellite knots have a higher trunk number than their base knots.
problem Understanding the relationship between satellite knots and their base knots.
method Using the Thurston norm and properties of satellite patterns.
result The trunk number of satellite knots is strictly greater than the product of the Thurston norm and the trunk number of their base knots.
In this paper, we investigate three geometrical invariants of knots, the height, the trunk and the representativity. First, we give a conterexample for the conjecture which states that the height is additive under connected sum of knots. We also define the minimal height of a knot and give a potential example which has…
Random sampling improves DeepONet training efficiency without sacrificing accuracy.
problem Training DeepONet models with high computational and memory costs.
method Random sampling of inputs in the trunk network of DeepONet.
result Significant reduction in training time with comparable accuracy.
The trunk of a knot in S3, defined by Makoto Ozawa, is a measure of geometric complexity similar to the bridge number or width of a knot. We prove that for any two knots K1 and K2, we have tr(K1#K2)=max{tr(K1),tr(K2)}, confirming a conjecture of Ozawa. Another conjecture of Ozawa asserts that any…
We introduce two numerical invariants, the waist and the trunk of knots. The waist of a closed incompressible surface in the complement of a knot is defined as the minimal intersection number of all compressing disks for the surface in the 3-sphere and the knot. Then the waist of a knot is defined as the maximal waist …
A new training method improves stability and generalization of DeepONets.
problem Training deep operator networks (DeepONets) is challenging due to nonconvex and nonlinear nature.
method Two-step training method: first train trunk network, then branch network. Introduced Gram-Schmidt orthonormalization.
result Generalization error estimate and numerical examples demonstrating effectiveness.
Improved DeepONet variants using Transformer cross-conditioning enhance PDE solution efficiency.
problem Solving partial differential equations efficiently and accurately.
method Transformer-inspired DeepONet variants with bidirectional cross-conditioning.
result Improved efficiency and accuracy compared to modified DeepONet, with variant effectiveness tied to PDE characteristics.
We construct a new invariant-the trunkenness-for volume-perserving vector fields on S^3 up to volume-preserving diffeomorphism. We prove that the trunkenness is independent from the helicity and that it is the limit of a knot invariant (called the trunk) computed on long pieces of orbits.
AMORE uses neural operators to efficiently predict multiple thermochemical states in stiff chemical kinetics.
problem Efficiently integrating stiff chemical kinetics systems to reduce computational cost.
method Developed AMORE, a framework of adaptive multi-output operator network with two adaptive loss functions.
result Demonstrated improved accuracy and efficiency in predicting thermochemical states from initial conditions.
DeepONet learns operators for PDEs with varying parameters and initial conditions.
problem Learning operators for partial differential equations with different parameters or initial conditions.
method DeepONet uses a Branch net and Trunk net to minimize error between evaluated and expected outputs, incorporating a scalar auxiliary variable approach for energy dissipation.
result DeepONet can accurately approximate operators for PDEs with varying parameters or initial conditions.
We define combinatorial analogues of stable and unstable minimal surfaces in the setting of weighted pseudomanifolds. We prove that, under mild conditions, such combinatorial minimal surfaces always exist. We use a technique, adapted from work of Johnson and Thompson, called thin position. Thin position is defined usin…
Flattenings of knotted surfaces help define new invariants.
problem Understanding and quantifying knotted surfaces in 4-sphere.
method Using hyperbolic decompositions and projections onto 2-sphere.
result Introduced layering, trunk, and partition number invariants.
Proposes a new method to learn operators for stochastic problems using DeepONet with autoencoder.
problem Efficiently solve forward and inverse stochastic problems with limited data.
method MultiAuto-DeepONet, a multi-resolution autoencoder DeepONet model.
result The model effectively handles high-dimensional stochastic inputs and reduces the number of trainable parameters.
The main result of this paper is a new classification theorem for links (smooth embeddings in codimension 2). The classifying space is the rack space (defined in [Trunks and classifying spaces, Applied Categorical Structures, 3 (1995) 321--356]) and the classifying bundle is the first James bundle (defined in "James bu…
New invariants defined for volume-preserving flows on 3-manifolds.
problem Defining invariants for volume-preserving flows.
method Extending wrapping number and trunk to define invariants of links and flows.
result Wrappingness and trunkenness are not functions of helicity.
Study approximates operators on labelled conditional distributions for non-exchangeable systems.
problem Approximating operators on constrained probability measures for non-exchangeable systems.
method Combines cylindrical approximations and DeepONet-type neural architecture for finite-dimensional representations.
result Establishes a universal approximation theorem for continuous operators on Mλ. DeepONets improve surrogate modeling for engineering systems.
problem Accurately modeling complex PDEs for engineering systems.
method DeepONets specialize in approximating mathematical operators for PDEs.
result DeepONets achieve high prediction accuracy and zero-shot capability.
Innovative rack theory applied to Legendrian links.
problem Classifying and distinguishing Legendrian links.
method Purely rack-theoretic approach, Legendrian Reidemeister moves, cusps, homogeneous representations, modules.
result Invariant distinguishes infinitely many Legendrian unknots and trefoils.
While it is widely known that neural networks are universal approximators of continuous functions, a less known and perhaps more powerful result is that a neural network with a single hidden layer can approximate accurately any nonlinear continuous operator. This universal approximation theorem is suggestive of the pot…
Deep learning framework predicts surface texture parameters and their uncertainties.
problem Predicting surface texture parameters and their uncertainties from multi-instrument datasets.
method Reproducible deep learning framework using multi-instrument dataset, quantile and heteroscedastic heads for uncertainty modeling, and post-hoc conformal calibration.
result High fidelity predictions (R2: Ra 0.9824, Rz 0.9847, RONt 0.9918) and well-modelled uncertainty targets (Ra_uncert 0.9899, Rz_uncert 0.9955).
Freezing of gait (FoG) is a common gait disability in Parkinson's disease, that usually appears in its advanced stage. Freeze episodes are associated with falls, injuries, and psychological consequences, negatively affecting the patients' quality of life. For detecting FoG episodes automatically, a highly accurate dete…
DeepONets enhance spatial-temporal surrogates for structural dynamics.
problem Creating full spatial-temporal surrogates for dynamical systems under uncertainty.
method Proposed Full-Field Extended DeepONet (FExD) to learn full solution operator across multiple degrees of freedom.
result FExD achieves superior accuracy and computational efficiency compared to other models.
Enhanced DeepONet framework with uncertainty quantification for complex operators.
problem Learning complex operators with uncertainty quantification.
method Generalised variational inference (GVI) using Rényi's α-divergence.
result Superior predictive accuracy and uncertainty quantification.
Study examines how body segments respond to random vibrations.
problem Understanding human body responses to random vibrations.
method 35 participants were tested with random noise signals. Multiple linear regression models were created to determine influential predictors of peak translational gains.
result Multiple predictors, including motion direction and body segment, significantly influence peak translational gains.
The paper analyzes geometric densities and compression radii for knot types.
problem Optimizing geometric quantities associated with knot types.
method Develops a factorization framework for scale-covariant size functionals.
result Different minimizing sequences for density, compression, packing, and ropelength problems.
Taylorized training improves neural network training at finite width.
problem Understanding and improving neural network training at finite width.
method Training the k-th order Taylor expansion of the neural network at initialization.
result Taylorized training agrees with full neural network training better as k increases and can significantly close the performance gap.
Self-training outperforms pre-training on COCO object detection and segmentation datasets.
problem The effectiveness of pre-training in improving object detection and segmentation models is limited.
method Investigated self-training as an alternative method to utilize additional data.
result Self-training consistently improves model performance across various dataset sizes and data augmentation levels.
Un-trained neural networks outperform trained methods in MRI reconstruction.
problem Accelerated MRI reconstruction with minimal training data.
method Variation of Deep Decoder without training data.
result Un-trained approach significantly outperforms other methods in reconstruction accuracy.
Meta-learning performance depends on train-validation split type.
problem Understanding the importance of train-validation split in meta-learning.
method Theoretical and experimental study comparing train-val and train-train methods.
result Train-train method can achieve strictly better excess loss in realizable cases.
New method reveals how training data influence diffusion model outputs.
problem Difficulty in assessing training data impact on diffusion model outputs.
method Use of ensembles trained on carefully engineered splits of training data to identify influential training examples.
result Demonstrated the viability of ensembles as generative models and validity of assessing influence.
Paper shows adversarial training can be fooled by new type of noise.
problem Adversarial training can be fooled by new types of noise.
method Designing ADVIN, a new type of inducing noise.
result ADVIN can degrade adversarial training robustness by 99.9%.
New MIP methods improve training of integer-valued neural networks.
problem Training integer-valued neural networks with limited data and resources.
method Formulated new MIP models to optimize training efficiency and handle more data.
result Significantly outperforms previous state-of-the-art methods in accuracy, training time, and data usage.
Free adversarial training reduces the generalization gap compared to vanilla method.
problem Improving generalization in adversarial training.
method Analysis of algorithmic stability in free adversarial training.
result Free adversarial training shows a lower generalization gap.
Making neural networks robust against adversarial inputs has resulted in an arms race between new defenses and attacks. The most promising defenses, adversarially robust training and verifiably robust training, have limitations that restrict their practical applications. The adversarially robust training only makes the…
learn2mix trains neural nets faster by adjusting class proportions dynamically.
problem Training neural nets efficiently with limited resources and imbalanced classes.
method Adaptive class proportion adjustment during training.
result Neural nets trained with learn2mix converge faster than static methods.
Loss-guided training accelerates node embedding methods on graphs.
problem Training efficiency in graph learning methods with implicit positive examples.
method Dynamic adjustment of training distribution based on loss values.
result Significant acceleration in training and computation over static methods.
Paper explores fast adversarial training to improve robustness with less computation.
problem Efficiently defending against adversarial examples.
method Integrates simple self-attacks for faster training, focusing on overfitting recovery.
result Shows superior robust accuracy with reduced training time compared to strong adversarial training.
Stochastic gradient decent~(SGD) and its variants, including some accelerated variants, have become popular for training in machine learning. However, in all existing SGD and its variants, the sample size in each iteration~(epoch) of training is the same as the size of the full training set. In this paper, we propose a…
Crowdsourced training of large neural networks with decentralized Mixture-of-Experts.
problem Expensive training of large neural networks limits research contributions.
method Learning@home: decentralized Mixture-of-Experts for large, poorly connected participants.
result Performance and reliability of Learning@home surpass conventional distributed training.
Pre-training is crucial for learning deep neural networks. Most of existing pre-training methods train simple models (e.g., restricted Boltzmann machines) and then stack them layer by layer to form the deep structure. This layer-wise pre-training has found strong theoretical foundation and broad empirical support. Howe…
New method reconstructs significant parts of training data from neural networks.
problem Understanding and reconstructing training data from neural networks.
method Proposes a novel reconstruction scheme based on recent theoretical results about neural network training.
result Shows that a significant fraction of training data can be reconstructed from neural network parameters.
The paper studies the asymptotic behavior of adversarial training under ℓ∞-perturbation.
problem Theoretical guarantees for sparsity-recovery in adversarial training.
method Investigation of the asymptotic distribution of the adversarial training estimator in generalized linear models.
result The asymptotic distribution of the adversarial training estimator under ℓ∞-perturbation could have a positive probability mass at 0 when the true parameter is 0. This paper introduces 'General Cyclical Training' for neural networks.
problem Improving training efficiency and performance of neural networks.
method Cyclical training phases with varying hyperparameters, batch sizes, loss functions, and data augmentation.
result Cyclical weight decay, softmax temperature, and gradient clipping enhance model accuracy.
Speeds up training and inference by pruning entire channels before training.
problem Training and inference speed in deep neural networks.
method Structured pruning applied before training, focusing on removing entire channels and hidden units.
result 2x speedup in training and 3x speedup in inference.
The paper predicts loss scaling across different datasets and compute scales.
problem Predicting loss scaling across different datasets and compute scales.
method Derive shifted power law relationships between train and test losses.
result Shifted power law relationships hold for various datasets and tasks, improving prediction accuracy.