This paper tackles discontinuous neural networks for better approximation of piecewise continuous functions.
problem Limitation of neural networks in approximating piecewise continuous functions due to discontinuities.
method Proposes a decoupled two-step procedure to train a discontinuous deep neural network model.
result Provides approximation guarantees for the proposed model in piecewise continuous function spaces.
Kolmogorov neural networks can represent various types of functions.
problem Representing different types of functions with neural networks.
method Continuous, discontinuous bounded or unbounded activation functions in a two hidden layer model.
result Kolmogorov neural networks can represent continuous, discontinuous bounded and all unbounded multivariate functions.
New proof shows neural networks can represent all multivariate functions.
problem Representing all multivariate functions with neural networks.
method Proved that three-layer neural networks can represent both continuous and discontinuous functions.
result Three-layer neural networks can represent all multivariate functions, including discontinuous ones.
Cut-DeepONet handles discontinuities and sharp transitions in neural operators.
problem Neural operators struggle with discontinuities and sharp transitions in PDEs.
method Two-stage training framework that explicitly models discontinuities via a lifting strategy and input-dependent discontinuity prediction.
result Cut-DeepONet outperforms state-of-the-art methods on benchmark PDEs with low-resolution datasets.
Paper analyzes SGHMC for non-convex optimization with discontinuous gradients.
problem Training neural networks with ReLU activation.
method Non-asymptotic convergence analysis of SGHMC with discontinuous gradients.
result Explicit upper bounds for expected excess risk in non-convex optimization.
New neural network designs learn contact dynamics efficiently.
problem Learning contact dynamics in robotics from noisy data.
method Physically structured neural networks.
result Data-efficient learning of discontinuous contact events.
In neural networks, it is often desirable to work with various representations of the same space. For example, 3D rotations can be represented with quaternions or Euler angles. In this paper, we advance a definition of a continuous representation, which can be helpful for training deep neural networks. We relate this t…
Framework uses deep learning and statistical models to solve PDEs with discontinuous coefficients.
problem Solving PDEs with discontinuous coefficients.
method Two-stage physics-informed deep learning and statistical mixture models.
result Framework achieves adaptability and accurate parameter identification.
In this paper, we present a novel and principled approach to learn the optimal transport between two distributions, from samples. Guided by the optimal transport theory, we learn the optimal Kantorovich potential which induces the optimal transport map. This involves learning two convex functions, by solving a novel mi…
This paper analyzes the training dynamics of binary neural networks using information bottleneck.
problem Training binary neural networks is challenging due to discontinuity in activation functions.
method The approach uses the Information Bottleneck principle to analyze BNN training dynamics.
result Training dynamics of BNNs are different from DNNs, with both phases occurring simultaneously.
DGNet solves complex dynamical systems with neural networks and constraints.
problem Real-time accurate solutions for large-scale complex systems.
method Model-constrained discontinuous Galerkin Network (DGNet) for compressible Euler equations.
result DGNet achieves out-of-distribution generalization and improved stability.
New algorithm tackles optimization problems with discontinuous gradients in finance and insurance.
problem Optimization problems with discontinuous stochastic gradients in finance and insurance.
method Langevin dynamics based algorithm e-THεO POULA. result Non-asymptotic error bounds and expected excess risk estimates for e-THεO POULA. We propose a simple change to existing neural network structures for better defending against gradient-based adversarial attacks. Instead of using popular activation functions (such as ReLU), we advocate the use of k-Winners-Take-All (k-WTA) activation, a C0 discontinuous function that purposely invalidates the neural …
New method finds optimal learning rates for neural nets.
problem Finding optimal learning rates in stochastic neural networks.
method Gradient-only line searches using Non-negative Associative Gradient Projection Points (NN-GPPs).
result Learning rates can be reliably resolved as step sizes along search directions.
Mini-batch sub-sampling in neural network training is unavoidable, due to growing data demands, memory-limited computational resources such as graphical processing units (GPUs), and the dynamics of on-line learning. In this study we specifically distinguish between static mini-batch sub-sampled loss functions, where mi…
We develop a novel multi-fidelity framework that goes far beyond the classical AR(1) Co-kriging scheme of Kennedy and O'Hagan (2000). Our method can handle general discontinuous cross-correlations among systems with different levels of fidelity. A combination of multi-fidelity Gaussian Processes (AR(1) Co-kriging) and …
Deep neural networks improve surrogate models for non-smooth quantities in uncertain geometries.
problem Building accurate surrogates for non-smooth quantities in uncertain geometries.
method Deep neural networks for point evaluation of solutions to interface problems with geometric uncertainties.
result Neural networks provide good surrogates without suffering from the curse of dimensionality.
A method for identifying NPWARX models with arbitrary domains using probabilistic mixture models.
problem Identifying hybrid system models with discontinuous maps.
method Probabilistic mixture model with a neural network for nonlinear partitioning and Expectation Maximization for parameter estimation.
result Demonstrated on a nonlinear piece-wise problem with discontinuous maps.
For node level graph encoding, a recent important state-of-art method is the graph convolutional networks (GCN), which nicely integrate local vertex features and graph topology in the spectral domain. However, current studies suffer from several drawbacks: (1) graph CNNs relies on Chebyshev polynomial approximation whi…
In this work, we propose a generalized likelihood ratio method capable of training the artificial neural networks with some biological brain-like mechanisms,.e.g., (a) learning by the loss value, (b) learning via neurons with discontinuous activation and loss functions. The traditional back propagation method cannot tr…
New algorithm trains deep neural networks without global optimization.
problem Training deep neural networks efficiently and without global optimization.
method Uses random complex exponential activation functions and Markov Chain Monte Carlo sampling.
result Consistently attains theoretical approximation rate for residual networks.
Novel active learning framework using sparse approximation for efficient model training.
problem Efficient model training with limited labeled data.
method Formulates batch active learning as sparsity-constrained discontinuous optimization problems, using greedy or proximal iterative hard thresholding algorithms.
result Achieves competitive performance with lower computational complexity across different settings.
In recent years, an increasing number of neural network models have included derivatives with respect to inputs in their loss functions, resulting in so-called double backpropagation for first-order optimization. However, so far no general description of the involved derivatives exists. Here, we cover a wide array of s…
New method reduces errors in pricing and sensitivities for discontinuous payoffs.
problem Errors in pricing and sensitivities for discontinuous payoffs in digital and barrier options.
method Alternative methods for estimating sensitivities, including likelihood ratio and hybrid methods.
result New methods substantially reduce test errors in prices and sensitivities.
Deep learning solves barrier options with stochastic volatility.
problem Solving barrier options with stochastic volatility.
method Unsupervised deep learning neural networks trained to satisfy PDE and boundary conditions.
result Neural networks accurately price barrier options in a single framework.
Low-variance gradient estimation is crucial for learning directed graphical models parameterized by neural networks, where the reparameterization trick is widely used for those with continuous variables. While this technique gives low-variance gradient estimates, it has not been directly applicable to discrete variable…
Active learning improves GP regression on complex, high-dimensional data.
problem Improving Gaussian Process regression in high-dimensional spaces with discontinuous functions.
method Combines manifold learning with active learning to optimize data selection and reduce dimensionality.
result Superior performance over random learning in synthetic data experiments.
Neural Laplace models diverse DEs in the Laplace domain for better dynamics.
problem Inadequate ODEs for long-range dependencies and discontinuities.
method Unified framework in Laplace domain, using stereographic map for smoothness.
result Superior performance in diverse DEs, including complex history dependency and abrupt changes.
A new algorithm learns model regimes and parameters efficiently.
problem Learning high-dimensional parameters with discontinuous jumps.
method Differentiable interacting multiple model particle filter with gradient descent.
result Superior numerical performance compared to previous methods.
We study the emergence of sparse representations in neural networks. We show that in unsupervised models with regularization, the emergence of sparsity is the result of the input data samples being distributed along highly non-linear or discontinuous manifold. We also derive a similar argument for discriminatively trai…
Wiatowski and Bölcskei, 2015, proved that deformation stability and vertical translation invariance of deep convolutional neural network-based feature extractors are guaranteed by the network structure per se rather than the specific convolution kernels and non-linearities. While the translation invariance result appli…
Neural networks solve the Dirichlet problem for Monge-Ampère equations.
problem Solving the Dirichlet problem for the Monge-Ampère equation.
method Using deep input convex neural networks to find the unique convex solution.
result Deep input convex neural networks can solve the Monge-Ampère Dirichlet problem.
A method for efficient approximate inference on discrete distributions.
problem Applying SVGD to discrete distributions.
method Transforming discrete distributions to piecewise continuous distributions for SVGD application.
result Outperforms traditional algorithms and ensemble methods on discrete graphical models.
Cubic spline smoothing improves interpolation between irregularly sampled data.
problem Interpolation discontinuity in recurrent neural networks for irregularly sampled sequences.
method Cubic spline smoothing compensation module trained end-to-end with ODE-RNN.
result Improves interpolation between irregularly sampled data points.
Bayesian estimators for causal inference using hierarchical Gaussian Processes.
problem Estimating causal effects in sharp and fuzzy RD/RK designs.
method Hierarchical Gaussian Process models for regression and classification.
result Hierarchical GP models improve precision and coverage of RD/RK estimations.
NSR enables neural networks to reason with continuous numbers and extrapolate.
problem Quantitative reasoning and extrapolation in neural networks.
method Proposes Neural Status Register (NSR) for continuous number reasoning.
result NSR achieves extrapolation to numbers many orders of magnitude larger than training data.
Surrogate strategies are used widely for uncertainty quantification of groundwater models in order to improve computational efficiency. However, their application to dynamic multiphase flow problems is hindered by the curse of dimensionality, the saturation discontinuity due to capillarity effects, and the time-depende…
Proposes VCNet for estimating ADRFs of continuous treatments.
problem Estimating ADRFs of continuous treatments from observational data.
method VCNet for improved model expressiveness and continuity; targeted regularization for finite sample performance.
result Improves model expressiveness and continuity of ADRFs.
Neural networks can approximate high-dimensional classifiers with ReLU networks under margin conditions.
problem Approximating high-dimensional discontinuous classifiers with neural networks.
method Using ReLU neural networks with three hidden layers, approximating a classifier with a Barron-regular decision boundary.
result High-dimensional discontinuous classifiers can be approximated with a rate of n−1 under strong margin conditions. Study shows overparametrization can shift and bend loss landscapes, affecting signal recovery.
problem Understanding how overparametrization affects loss landscapes in neural networks.
method Field theory analysis of Hessian spectrum at initialization.
result Overparametrization can shift the BBP transition point, potentially reaching weak-recovery threshold.
Paper introduces a differentiable regularizer for condition number to improve neural network stability.
problem Maintaining numerical stability in neural networks to ensure reliable and performant models.
method Introduces a novel differentiable regularizer for the condition number of weight matrices.
result Derives a differentiable formula for the gradient of the regularizer, promoting matrices with low condition numbers.
Unified framework identifies nonlinear systems using characteristic curves and neural networks.
problem Balancing interpretability and flexibility in nonlinear system identification.
method Combines differential equation structure with neural networks, using characteristic curves as modular components.
result NN-CC approach outperforms other methods in complex nonlinear systems.
Deep networks adapt to function regularity and data distribution.
problem Understanding deep learning's adaptability to function regularity and data distribution.
method Developed nonparametric approximation and estimation theories for a broad class of functions using deep ReLU networks.
result Deep neural networks are adaptive to different regularity of functions and nonuniform data distributions.
TUSLA algorithm solves non-convex optimization problems with ReLU activations.
problem Non-convex stochastic optimization with super-linearly growing and discontinuous gradients.
method Non-asymptotic analysis of TUSLA algorithm for non-convex learning.
result TUSLA provides non-asymptotic error bounds in Wasserstein distances for non-convex learning.
We describe an approach to understand the peculiar and counterintuitive generalization properties of deep neural networks. The approach involves going beyond worst-case theoretical capacity control frameworks that have been popular in machine learning in recent years to revisit old ideas in the statistical mechanics of…
SGD learns sparse parities near computational limits with discontinuous phase transitions.
problem Learning sparse parities in deep learning.
method Empirical and theoretical analysis of SGD on sparse parities.
result SGD makes progress on sparse parities via Fourier gap, not Langevin-like mechanism.
We study layered neural networks of rectified linear units (ReLU) in a modelling framework for stochastic training processes. The comparison with sigmoidal activation functions is in the center of interest. We compute typical learning curves for shallow networks with K hidden units in matching student teacher scenarios…
Unified reinforcement learning and stochastic processes with action-driven processes.
problem Combining reinforcement learning and stochastic processes for efficient control.
method Action-driven processes, leveraging control-as-inference, and minimizing Kullback-Leibler divergence.
result Action-driven processes unify reinforcement learning and stochastic processes, equivalent to maximum entropy reinforcement learning.