Sparse butterfly network replaces dense layers in neural networks, improving expressibility and performance.
problem Improving expressibility and performance of neural networks with dense layers.
method Replacing dense layers with a butterfly network architecture.
result The proposed architecture significantly reduces the number of weights from quadratic to nearly linear, with comparable or better performance.
Deep networks, especially convolutional neural networks (CNNs), have been successfully applied in various areas of machine learning as well as to challenging problems in other scientific and engineering fields. This paper introduces Butterfly-Net, a low-complexity CNN with structured and sparse cross-channel connection…
WideBNet learns inverse scattering from wide-band data efficiently and stably.
problem Learning the inverse scattering map from wide-band scattering data.
method Combines butterfly factorization, FFT, and deep learning.
result WideBNet requires fewer training points and has stable training dynamics.
In unsupervised domain adaptation (UDA), classifiers for the target domain (TD) are trained with clean labeled data from the source domain (SD) and unlabeled data from TD. However, in the wild, it is difficult to acquire a large amount of perfectly clean labeled data in SD given limited budget. Hence, we consider a new…
We give an explicit handy (and cocycle-free) description of the groupoid of weak maps between two crossed-modules in terms of certain digrams of groups which we we call a {\em butterflies}. We define composition of butterflies and this way find a bicategory that is naturally biequivalent to the 2-category of pointed ho…
ButterflyFlow uses butterfly matrices for efficient invertible layers in normalizing flows.
problem Building efficient invertible layers for complex probability distributions.
method Proposes butterfly layers for invertible linear layers, leveraging their ability to capture complex structures.
result ButterflyFlow achieves strong density estimation and significantly better log-likelihoods on various datasets.
Characterizes no Butterfly arbitrage in SVI model parameters.
problem No Butterfly arbitrage in SVI implied total variance formula.
method Characterization using intermediary condition from Fukasawa (2012) and rescaling of SVI parameters.
result Simple range conditions on SVI parameters ensure no Butterfly arbitrage.
Simplified Butterfly-Net2 improves CNN efficiency in solving PDEs and signal processing tasks.
problem Improving CNN efficiency in solving PDEs and signal processing tasks.
method Introducing BNet2, a simplified Butterfly-Net, and Fourier transform initialization.
result BNet2 achieves similar accuracy as CNN but with fewer parameters and improves accuracy over randomly initialized CNN.
We show how to integrate a weak morphism of Lie algebra crossed-modules to a weak morphism of Lie 2-groups. To do so we develop a theory of butterflies for 2-term L_infty algebras. In particular, we obtain a new description of the bicategory of 2-term L_infty algebras. We use butterflies to give a functorial constructi…
Traditional anatomical analyses captured only a fraction of real phenomic information. Here, we apply deep learning to quantify total phenotypic similarity across 2468 butterfly photographs, covering 38 subspecies from the polymorphic mimicry complex of Heliconius erato and Heliconius melpomene. E…
Study on 2-bridge knots, proving equivariant concordance order is infinite.
problem Equivariant concordance of 2-bridge knots.
method Formula for butterfly polynomial, two proofs of non-equivariant sliceness, new invariant for strongly invertible knots.
result Equivariant concordance order of 2-bridge knots is infinite.
Unified framework for inference in complex nonlinear processes.
problem Challenges in inferring nonlinear continuous stochastic processes with sparse observations and complex topologies.
method Neural Backward Filtering Forward Guiding (NBFFG) framework that constructs a variational posterior using a proxy linear-Gaussian process.
result Empirical results show NBFFG outperforms baselines on synthetic benchmarks and high-dimensional phylogenetic analysis tasks.
Unified 3D R-matrices from quantum cluster algebra.
problem Constructing new solutions to the tetrahedron equation.
method Symmetric butterfly quiver, quantum cluster algebra, quantum dilogarithms, q-Weyl algebra.
result Unified 3D R-matrices from various sources.
Support Vector Regression (SVR) has achieved high performance on forecasting future behavior of random systems. However, the performance of SVR models highly depends upon the appropriate choice of SVR parameters. In this study, a novel BOA-SVR model based on Butterfly Optimization Algorithm (BOA) is presented. The perf…
In this paper we test for the sensitive dependence on initial conditions (the so called "butterfly effect") of energy futures time series (heating oil, natural gas), and thus the determinism of those series. This paper is distinguished from previous studies in the following points: first, we reread existent works in th…
Efficient trainable front-end for neural speech enhancement.
problem Inefficient STFT front-ends in neural speech enhancement models.
method Butterfly mechanism for Fast Fourier Transform, trainable STFT window.
result Accuracy and efficiency improvements for low-compute systems.
We study surfaces of constant positive Gauss curvature in Euclidean 3-space via the harmonicity of the Gauss map. Using the loop group representation, we solve the regular and the singular geometric Cauchy problems for these surfaces, and use these solutions to compute several new examples. We give the criteria on the …
The study classifies points on ruled surfaces in 4-space based on geometric properties.
problem Characterizing points on smooth ruled surfaces in 4-space.
method Contact with transverse planes, binary differential equations, and projective transformations.
result Parabolic points on ruled surfaces in 4-space can be classified as butterfly hyperbolic, parabolic, or elliptic based on the discriminant of a binary differential equation.
Characterizes smiles in delta satisfying specific conditions.
problem Characterizing no butterfly arbitrage smiles in delta.
method Using parametrization of the smile in delta, we characterize the set of smiles.
result Obtained a parametrization of the set via one real number and three positive functions.
This paper is devoted to the application of an l1 -minimisation technique to construct an arbitrage-free call-option surface. We propose a nononparametric approach to obtaining model-free call option surfaces that are perfectly consistent with market quotes and free of static arbitrage. The approach is inspired from…
Paper studies singularities of timelike minimal surfaces in Minkowski 3-space.
problem Exploring singularities of timelike minimal surfaces in Minkowski 3-space.
method Existence and non-existence theorems, criteria for specific singularities.
result Various singularities unique to timelike minimal surfaces, including cuspidal butterfly and (2,5)-cuspidal edge. Unified market making controls risk, arbitrage, and volatility surfaces.
problem Market making risk, arbitrage, and volatility surface consistency.
method Constrained RL and stochastic control for risk-sensitive execution and hedging.
result Agent achieves positive P&L with zero calendar and butterfly violations.
Fast linear transforms are ubiquitous in machine learning, including the discrete Fourier transform, discrete cosine transform, and other structured transformations such as convolutions. All of these transforms can be represented by dense matrix-vector multiplication, yet each has a specialized and highly efficient (su…
We investigate singularities of all parallel surfaces to a given regular surface. In generic context, the types of singularities of parallel surfaces are cuspidal edge, swallowtail, cuspidal lips, cuspidal beaks, cuspidal butterfly and 3-dimensional D4± singularities. We give criteria for these singularities type…
This paper proposes a new algorithm for controlling classification results by generating a small additive perturbation without changing the classifier network. Our work is inspired by existing works generating adversarial perturbation that worsens classification performance. In contrast to the existing methods, our wor…
Graev's nerve implies invariant Einstein metrics on homogeneous spaces.
problem Existence of invariant Einstein metrics on homogeneous spaces.
method Lie-theoretic definition of Graev's nerve and curvature estimates.
result Detailed description of Graev's work and curvature estimates.
Study geometric singular solutions of generalized Monge-Ampère equations.
problem Solving generalized Monge-Ampère equations on a plane.
method Using exterior differential systems and Cauchy characteristics.
result Criteria for geometric singular solutions to be equivalent to specific types.
This paper explores loss landscapes of sparse neural networks, finding unique characteristics compared to dense networks.
problem Understanding the loss landscape of sparse neural networks, especially one-hidden-layer networks.
method Analyzes sparse networks with dense and sparse final layers, focusing on linear and non-linear models.
result Sparse networks can have no spurious valleys under certain conditions, but spurious valleys and minima can exist for wide sparse networks.
This work introduces a method to compare sparse neural network topologies using graph theory.
problem Comparing and understanding sparse neural network topologies, especially during training.
method Introducing Neural Network Sparse Topology Distance (NNSTD) to measure distances between different sparse neural networks.
result Sparse neural networks can outperform over-parameterized models without further structure optimization.
We simplify SVI volatility smile constraints for three sub-SVIs without numerical methods.
problem No arbitrage constraints for SVI volatility smiles.
method Explicit domain derivation for sub-SVIs without numerical procedures.
result Explicit no arbitrage domains for Symmetric SVI, Vanishing Upward/Downward SVI, and SSVI.
Dynamic Sparse Training finds efficient sparse networks from scratch.
problem Finding efficient sparse neural networks.
method Jointly optimizes network parameters and sparsity with trainable thresholds.
result Achieves state-of-the-art performance with minimal performance loss.
Deep learning solves wave-based inverse problems, including super-resolution imaging.
problem Solving inverse wave scattering problems across all length scales.
method Wide-band butterfly network coupled with dynamic noise injection.
result Framework successfully solves super-resolution imaging problems.
New method finds sparse networks without labels, improving performance.
problem Sparse connectivity in neural networks to reduce memory and energy demands.
method Neural Tangent Transfer method to find sparse networks without labels.
result Sparse networks achieve higher classification performance and faster convergence.
Guarantees sparse recovery for neural networks with iterative hard thresholding.
problem Recovering sparse network weights in neural networks.
method Structural properties of sparse network weights and iterative hard thresholding algorithm.
result Simple iterative hard thresholding algorithm recovers sparse network weights exactly using linear memory.
We describe a robust calibration algorithm of a set of SSVI slices (i.e. a set of 3 SSVI parameters θ,ρ,φ attached to each option maturity available on the market), which grants that these slices are free of Butterfly and Calendar-Spread arbitrage. Given such a set of consistent SSVI parameters, we show that …
There is vast empirical evidence that given a set of assumptions on the real-world dynamics of an asset, the European options on this asset are not efficiently priced in options markets, giving rise to arbitrage opportunities. We study these opportunities in a generic stochastic volatility model and exhibit the strateg…
The paper improves lower bounds on deep neural network trajectories.
problem Understanding the expressivity of deep neural networks.
method Generalized method for lower bounding trajectory growth in random sparse deep ReLU networks.
result Trajectory growth can remain exponential in depth with sparse variants of random nets.
Sparse deep neural networks(DNNs) are efficient in both memory and compute when compared to dense DNNs. But due to irregularity in computation of sparse DNNs, their efficiencies are much lower than that of dense DNNs on regular parallel hardware such as TPU. This inefficiency leads to poor/no performance benefits for s…
Sparse neural networks can match dense models on Lipschitz functions.
problem Sparse networks are more efficient but lack theoretical guarantees.
method Formal model of sparse networks, LSH-based routing function, Lipschitz function approximation.
result Sparse networks can approximate dense networks on Lipschitz functions.
Proposes a method to construct risk-neutral marginals from arbitrage-free option prices.
problem Lack of risk-neutral marginals that are free of arbitrage and easy to use.
method Explicit construction of risk-neutral marginals from discrete arbitrage-free option prices.
result Explicit construction guarantees risk-neutral marginals free of butterfly and calendar arbitrage.
Law derived for neural networks with sparse connections.
problem Understanding the behavior of neural networks with sparse connections.
method Law of large numbers for empirical distribution of parameters derived.
result Law for neural networks with sparse connections derived.
Method trains sparse neural networks without sacrificing accuracy.
problem Training sparse neural networks limits model size.
method Updates sparse network topology during training.
result Requires fewer FLOPs to achieve accuracy.
Sparse linear models improve neural network debuggability.
problem Improving neural network interpretability and debugging.
method Using sparse linear models over learned deep feature representations.
result The approach leads to more debuggable and accurate neural networks.
SPICE estimates sparse linear dynamic networks without hyperparameters.
problem Estimating topology and dynamics of sparse linear dynamic networks.
method SPICE (Sparse Iterative Covariance Estimation) method in an iterative framework.
result Directly reveals the underlying topology of the network.
Paper optimizes Laplacian regularization for sparse network clustering.
problem Improving spectral clustering in sparse networks.
method Formally determines optimal Laplacian regularization.
result Proper regularization is closely tied to state-of-the-art techniques.
USN improves neural networks with uniform sparse connectivity.
problem Overfitting and limited scalability in classical neural networks.
method Uniform sparse network (USN) with even and sparse connectivity.
result USN outperforms state-of-the-art sparse network models in accuracy, speed, and robustness.
SnAp approximates RTRL for online training of sparse recurrent networks.
problem Training large sparse recurrent networks online is computationally expensive.
method Sparse n-step Approximation (SnAp) of the RTRL influence matrix.
result SnAp with n=2 remains tractable for highly sparse networks and outperforms backpropagation through time.
We propose a method for solving statistical mechanics problems defined on sparse graphs. It extracts a small Feedback Vertex Set (FVS) from the sparse graph, converting the sparse system to a much smaller system with many-body and dense interactions with an effective energy on every configuration of the FVS, then learn…