A new method to rescale ReLU neural networks based on path-lifting.
problem Lack of principled ways to leverage rescaling symmetries in ReLU neural networks.
method Introduces a geometrically motivated criterion to rescale neural network parameters, aligning a kernel in the path-lifting space with a chosen reference.
result Proposed method can speed up training and aligns a kernel in the path-lifting space with a chosen reference.
Paper introduces non-adversarial training for Neural SDEs using signature kernel scores.
problem Stability and mode collapse issues in adversarial training of Neural SDEs.
method Uses signature kernel scores as objective function for non-adversarial training.
result Non-adversarial training leads to better performance and more stable models.
The paper analyzes the role of ReLU gates in deep learning networks.
problem Understanding the role of gates in deep learning networks.
method Developed neural path features (NPF) and neural path values (NPV) to characterize the active sub-networks during training.
result The neural path kernel associated with NPFs is a fundamental quantity that characterizes the information stored in the gates of a DNN.
Signature kernel scoring rule improves weather forecasting by capturing temporal and spatial dependencies.
problem Lack of suitable scoring rules for probabilistic weather forecasting.
method Reframe weather variables as continuous paths using iterated integrals (signature kernels) to capture temporal and spatial dependencies.
result Signature kernel scoring rule outperforms conventional methods in weather forecasting, especially for long-term forecasts.
The behavior of the gradient descent (GD) algorithm is analyzed for a deep neural network model with skip-connections. It is proved that in the over-parametrized regime, for a suitable initialization, with high probability GD can find a global minimum exponentially fast. Generalization error estimates along the GD path…
Framework for training stochastic spiking neural networks with rough signals.
problem Training stochastic spiking neural networks with noisy spike timing and dynamics.
method Rough path theory and signature kernels for gradient computation.
result Pathwise gradients of SSNNs' trajectories and event times exist and satisfy a recursive relation.
Scalable machine learning with path signatures for time series and graphs.
problem Challenges in real-world time series and graph data.
method Combines rough path theory with probabilistic, deep, and kernel methods.
result Scalable models for time series and graph data.
Blending multiple convolutional kernels is proved advantageous in neural architecture design. However, current two-stage neural architecture search methods are mainly limited to single-path search spaces. How to efficiently search models of multi-path structures remains a difficult problem. In this paper, we are motiva…
Develops a new solver for path-dependent PDEs using signature kernels.
problem Solving path-dependent PDEs (PPDEs) efficiently and accurately.
method Uses signature kernels to solve PPDEs by approximating the solution with minimal norm in a reproducing kernel Hilbert space.
result Proves the consistency of the numerical scheme, ensuring convergence to PPDE solutions as the number of collocation points increases.
Tree++ graph kernel captures similarities at multiple granularities.
problem Lack of scale-adaptivity in existing graph kernels.
method Tree++ uses truncated BFS trees and super paths to represent graphs at different granularities.
result Tree++ achieves best classification accuracy on real-world graphs.
Kernel for Lévy rough paths derived from PDE system.
problem Computing similarity measures for Lévy rough paths.
method Developed a PDE system for the expected signature of inhomogeneous Lévy processes.
result Gaussian martingales' expected signature kernel satisfies a Goursat PDE.
Characterizes neural kernel and NNGP for various activations.
problem Understanding neural kernels and NNGP for non-RELU activations.
method Characterization of RKHS for various activation functions.
result Broad class of non-infinitely smooth activations generate equivalent RKHSs at different depths.
New method uses path signatures for efficient likelihood estimation in time-series data.
problem Intractable likelihood functions in complex dynamic models.
method Kernel classifier based on path signatures for sequential data.
result Path signatures yield highly performant classifiers, even with low sample numbers.
Functional input neural networks approximate continuous functions on weighted spaces.
problem Approximating continuous functions on infinite-dimensional weighted spaces.
method Additive family mapping, non-linear activation, linear readouts, Stone-Weierstrass theorem.
result Global universal approximation of continuous functions on weighted spaces.
Poor approximators found in neural networks and random feature models.
problem Understanding why certain neural networks and models perform poorly in approximating functions.
method Established a scale separation of Kolmogorov width type and applied it to neural networks and random feature models.
result Reproducing kernel Hilbert spaces and two-layer neural networks are poor L2-approximators for certain functions. In a rigorous construction of the path integral for supersymmetric quantum mechanics on a Riemann manifold, based on Bär and Pfäffle's use of piecewise geodesic paths, the kernel of the time evolution operator is the heat kernel for the Laplacian on forms. The path integral is approximated by the integral of a form on …
In the framework of path integral the evolution operator kernel for the Merton-Garman Hamiltonian is constructed. Based on this kernel option formula is obtained, which generalizes the well-known Black-Scholes result. Possible approximation numerical schemes for path integral calculations are proposed.
The study uses response theory to understand RNNs processing input signals.
problem Understanding how RNNs process sequential data.
method Deriving a Volterra series representation for SRNNs output using response theory from nonequilibrium statistical mechanics.
result SRNNs can be viewed as kernel machines operating on a reproducing kernel Hilbert space associated with the response feature.
Study shows how feature weighting affects neural network regularization.
problem Understanding how feature weighting influences neural network regularization.
method Derived equivalence paths connecting different weighting matrices and ridge regularization levels.
result Ridge estimators trained on weighted features are asymptotically equivalent when evaluated against test vectors.
We study a simplification of GAN training: the problem of transporting particles from a source to a target distribution. Starting from the Sobolev GAN critic, part of the gradient regularized GAN family, we show a strong relation with Optimal Transport (OT). Specifically with the less popular dynamic formulation of OT …
Volterra signature provides a clear, interpretable feature for history-dependent systems.
problem Learning from non-Markovian time series with implicit memory mechanisms.
method Develops Volterra signature as a tensor algebra representation weighted by a temporal kernel, proving injectivity and universal approximation.
result Volterra signature leads to linear functionals and universal approximation, improving dynamic learning tasks.
Regularization effect found in neural feature alignment.
problem Implicit regularization in deep learning models.
method Geometrical viewpoint and analysis of Rademacher complexity.
result Neural features align along task-relevant directions, leading to regularization.
Estimates path-valued data using signature metrics and local kernels.
problem Nonparametric regression and classification for path-valued data.
method Combines signature transform and local kernel regression.
result Establishes convergence bounds and demonstrates competitive accuracy.
Paper establishes a generalization bound for gradient flow using a data-dependent kernel.
problem Understanding the generalization properties of gradient-based optimization methods.
method Establishes a generalization bound for gradient flow through a data-dependent kernel called the loss path kernel (LPK).
result The LPK captures the entire training trajectory and leads to tighter generalization guarantees.
The paper defines conditions for Gaussian process sample path regularity.
problem Lack of understanding of Gaussian process sample path regularity.
method Analyzes covariance kernels to determine sample path regularity.
result Necessary and sufficient conditions for Hölder regularity are provided.
Conditional diffusion models can approximate target distributions well with Gaussian-mixture reverse kernels.
problem Approximating target distributions in conditional diffusion models.
method Using finite Gaussian mixtures with ReLU-network logits as reverse kernels, reducing the problem to static conditional density approximation.
result The resulting neural reverse-kernel class is dense in conditional KL divergence under exact terminal matching.
In this survey article, we review the relation between heat kernels and path integrals. In particular, we review recent results on the approximation of the Wiener measure on compact manifold by measures on (finite-dimensional) spaces of piece-wise geodesics.
Selecting important features in non-linear or kernel spaces is a difficult challenge in both classification and regression problems. When many of the features are irrelevant, kernel methods such as the support vector machine and kernel ridge regression can sometimes perform poorly. We propose weighting the features wit…
The paper studies local heat kernel properties on smooth manifolds.
problem Understanding heat kernel properties in open convex sets of smooth Riemannian manifolds.
method Utilizes path integral formulation to investigate properties like uniqueness, symmetry, and asymptotics.
result Uniqueness and symmetry of Seeley-DeWitt coefficients are established.
PathBoost boosts graph-level predictions using path-based features.
problem Graph-level classification and regression challenges.
method Gradient tree boosting method for graph-level prediction.
result PathBoost outperforms graph neural networks and graph kernel approaches in many cases.
This work extends Gaussian process priors to neural operators for function space mappings.
problem Improving uncertainty quantification in deep neural networks.
method Extending Gaussian process priors to neural operators with conditions for convergence and computation of covariance functions.
result Arbitrary-depth neural operators with Gaussian kernels converge to function-valued GPs, enabling posterior computation in regression scenarios.
Noether's theorem clarifies how symmetries in neural networks influence learning.
problem Understanding how symmetries in neural networks affect learning.
method Systematic study of symmetry interactions with learning algorithms using Noether's theorem.
result Symmetries impose restrictions on the optimization path, leading to conserved quantities.
EntroPath learns manifold geometry from diffusion paths.
problem Learning geodesic geometry from data graphs with spurious shortcuts.
method Maximum Entropy Path Ensemble Embedding (MERW) with k-step diffusion paths.
result EntroPath converges to squared geodesic distance in the short-time limit.
The study reveals how attention paths in Transformers influence learning outcomes.
problem Understanding the theoretical basis of Transformers' performance.
method Developed a statistical mechanics theory for a simplified attention network.
result The predictor statistics are influenced by the combination of attention paths.
This work derives closed-form expressions computing the expectation of co-presence and of number of co-occurrences of nodes on paths sampled from a network according to general path weights (a bag of paths). The underlying idea is that two nodes are considered as similar when they often appear together on (preferably s…
A new algorithm for high-dimensional hedging problems.
problem High-dimensional, path-dependent hedging problems.
method Signature-based algorithm using operator-valued kernels and geometric rough paths.
result Theoretical guarantees on existence and uniqueness of a global minimum.
A new graph kernel uses LCS and Wasserstein distance for better graph comparisons.
problem Graph learning methods can be limited by information from distant vertices and path length constraints.
method Proposes a Graph Kernel based on LCS similarity and Wasserstein distance in a novel metric space.
result The new kernel emphasizes comparisons between similar paths and reduces information loss.
We investigate the short-time expansion of the heat kernel of a Laplace type operator on a compact Riemannian manifold and show that the lowest order term of this expansion is given by the Fredholm determinant of the Hessian of the energy functional on a space of finite energy paths. This is the asymptotic behavior to …
sig-MMD tests compare path distributions using kernel methods.
problem Comparing path distributions in stochastic processes.
method Signature kernel for path space valued distributions.
result sig-MMD can lead to Type 2 errors in limited data settings.
Can we automatically design a Convolutional Network (ConvNet) with the highest image classification accuracy under the latency constraint of a mobile device? Neural Architecture Search (NAS) for ConvNet design is a challenging problem due to the combinatorially large design space and search time (at least 200 GPU-hours…
Efficiently computes sparse signature coefficients using kernels.
problem Lack of efficient methods for sparse signature coefficients.
method Signature kernels and PDE-based methods.
result Sparse groups of signature coefficients can be isolated effectively.
Following Feynman's prescription for constructing a path integral representation of the propagator of a quantum theory, a short-time approximation to the propagator for imaginary time, N=1 supersymmetric quantum mechanics on a compact, even-dimensional Riemannian manifold is constructed. The path integral is interprete…
Real time application of deep learning algorithms is often hindered by high computational complexity and frequent memory accesses. Network pruning is a promising technique to solve this problem. However, pruning usually results in irregular network connections that not only demand extra representation efforts but also …
This paper studies neural networks with bounded norms to avoid the curse of dimensionality.
problem The curse of dimensionality in approximating functions by neural networks.
method Investigates over-parameterized two-layer neural networks with norm constraints in RKHS.
result Improved sample complexity and generalization bounds for neural networks with bounded norms.
ICON-OCnet solves optimal execution problems with neural networks and few examples.
problem Optimal order execution in markets with unknown price impact.
method Transformer-based neural network architecture (ICON-OCnet) that learns price impact from few examples and applies it to optimal execution strategies.
result ICON-OCnet accurately infers price impact models and retrieves optimal execution strategies for various propagator kernels.
Can we automatically design a Convolutional Network (ConvNet) with the highest image classification accuracy under the runtime constraint of a mobile device? Neural architecture search (NAS) has revolutionized the design of hardware-efficient ConvNets by automating this process. However, the NAS problem remains challen…
Study proves existence, uniqueness, and positivity of solutions to a complex volatility model.
problem Modeling equity index and spot volatility with path-dependent features and general kernels.
method Proved existence and uniqueness of a continuous solution to a Stochastic Volterra Equation (SVE) with non-convolutional, non-bounded kernels and non-Lipschitz coefficients.
result Positivity of the volatility process under certain conditions on the kernels.
Adapts flow matching for MCMC to improve sampling efficiency.
problem Improving sampling efficiency in MCMC for complex distributions.
method Combines Markov chain and CNFs to learn a path between distributions.
result Achieves similar performance to state-of-the-art methods but with lower computational cost.