Initialization of parameters in deep neural networks has been shown to have a big impact on the performance of the networks (Mishkin & Matas, 2015). The initialization scheme devised by He et al, allowed convolution activations to carry a constrained mean which allowed deep networks to be trained effectively (He et al.…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A hybrid model combines BPH and HE distributions for better heavy-tailed distribution approximation.
The Chirikov standard map and the 2D Froeschlé map are investigated. A few thousand values of the Hurst exponent (HE) and the maximal Lyapunov exponent (mLE) are plotted in a mixed space of the nonlinear parameter versus the initial condition. Both characteristic exponents reveal remarkably similar structures in this s…
Single layer Feedforward Neural Network(FNN) is used many a time as a last layer in models such as seq2seq or could be a simple RNN network. The importance of such layer is to transform the output to our required dimensions. When it comes to weights and biases initialization, there is no such specific technique that co…
It has been noted in existing literature that over-parameterization in ReLU networks generally improves performance. While there could be several factors involved behind this, we prove some desirable theoretical properties at initialization which may be enjoyed by ReLU networks. Specifically, it is known that He initia…
Proves positive mass theorem for hyperbolic manifolds with ends.
Deep neural networks (DNNs) form the backbone of almost every state-of-the-art technique in the fields such as computer vision, speech processing, and text analysis. The recent advances in computational technology have made the use of DNNs more practical. Despite the overwhelming performances by DNN and the advances in…
In his 1954 paper about the initial value problem for 2D hyperbolic nonlinear PDEs, P. Lax declared that he had "a strong reason to believe" that there must exist a well-defined class of "not genuinely nonlinear" nonlinear PDEs. In 1978 G. Boillat coined the term "completely exceptional" to denote it. In the case of $2…
The Residual Network (ResNet), proposed in He et al. (2015), utilized shortcut connections to significantly reduce the difficulty of training, which resulted in great performance boosts in terms of both training and generalization error. It was empirically observed in He et al. (2015) that stacking more layers of resid…
This paper analyzes the Lipschitz constants of deep neural networks with random weights.
The paper establishes principles for initializing and designing GNNs with ReLU activations to avoid oversmoothing and correlation collapse.
Consider an agent who enters a financial market on day t = 0 with an initial capital amount x. He invests this amount on stocks and the money market, and by day t = T, has generated a wealth W . He is given a convex class of probability measures (called scenarios) and a real-valued function (or floors) corresponding to…
New method stabilizes deep neural networks by setting Lyapunov exponent to zero.
New CNN initialization scheme derived from modern architectures.
Our study analyzes how neural network initialization affects privacy and utility in overparameterized models.
We prove that two-layer (Leaky)ReLU networks initialized by e.g. the widely used method proposed by He et al. (2015) and trained using gradient descent on a least-squares loss are not universally consistent. Specifically, we describe a large class of one-dimensional data-generating distributions for which, with high pr…
Nicolas-Auguste Tissot (1824--1897) was a French mathematician and cartographer. He introduced a tool which became known among geographers under the name ``Tissot indicatrix'', and which was widely used during the first half of the twentieth century in cartography. This is a graphical representation of a field of ellip…
Bitcoin volatility analysis shows decreasing HE with longer sampling periods.
New method for initializing RBM weights without datasets.
Unified learning-rate scale for CNNs and ResNets, avoiding depth imbalance.
Cryptotree enables accurate predictions on encrypted data using Random Forests.
New method for initializing low-rank neural networks improves performance.
Optimizes deep neural network initialization variance for better performance.
Mini-Hes improves LFA model performance on HDI tasks with missing data.
We prove that on Fano manifolds, the Kähler-Ricci flow produces a "most destabilising" degeneration, with respect to a new stability notion related to the H-functional. This answers questions of Chen-Sun-Wang and He. We give two applications of this result. Firstly, we give a purely algebro-geometric formula for the su…
Consider an American option that pays G(X^*_t) when exercised at time t, where G is a positive increasing function, X^*_t := \sup_{s\le t}X_s, and X_s is the price of the underlying security at time s. Assuming zero interest rates, we show that the seller of this option can hedge his position by trading in the underlyi…
We describe a general family of curved-crease folding tessellations consisting of a repeating "lens" motif formed by two convex curved arcs. The third author invented the first such design in 1992, when he made both a sketch of the crease pattern and a vinyl model (pictured below). Curve fitting suggests that this init…
A Bayesian agent learns about the structure of a stationary process from ob- serving past outcomes. We prove that his predictions about the near future become ap- proximately those he would have made if he knew the long run empirical frequencies of the process.
Adversarial training (AT) is one of the most effective defenses against adversarial attacks for deep learning models. In this work, we advocate incorporating the hypersphere embedding (HE) mechanism into the AT procedure by regularizing the features onto compact manifolds, which constitutes a lightweight yet effective …
This paper studies bounds for the Lipschitz constant of random neural networks.
Study solves sub-Laplacian equivalence on a specific Heisenberg group.
Proves initial data on big bang singularities for Einstein-nonlinear scalar field equations lead to unique solutions.
We review some ideas of Grothendieck and others on actions of the absolute Galois group Γ Q of Q (the automorphism group of the tower of finite extensions of Q), related to the geometry and topology of surfaces (mapping class groups, Teichm{ü}ller spaces and moduli spaces of Riemann surfaces). Grothendieck's motivation…
Paper finds how Steklov eigenvalues change on graphs and trees.
We solve the equivalence problem for the orthogonally separable webs on the three-sphere under the action of the isometry group. This continues a classical project initiated by Olevsky in which he solved the corresponding canonical forms problem. The solution to the equivalence problem together with the results by Olev…
In this paper we consider a modification of the classical Merton portfolio optimization problem. Namely, an investor can trade in financial asset and consume his capital. He is additionally endowed with a one unit of an indivisible asset which he can sell at any time. We give a numerical example of calculating the opti…
The paper discusses Nirenberg's work on geometric problems and his personality.
Study on the complexity of 1D ReLU neural networks, proving growth in linear regions.
Theory explains deep nonlinear networks' plateaus and transitions.
In a recent work, Galloway [9] proved a local foliation theorem by MOTSs for a 3-dimensional initial data set with mean curvature in a 4-dimensional spacetime when (under suitable assumptions) has a stable spherical MOTS which achieves an upper bound for the area. H…
Georg de Buquoy, Lord de Vaux, lived in Nove Hrady, Prague and Cerveny Hradek for most of his productive life. From his extensive scientific contributions, both theoretical and experimental, we expand here the discussion of his contributions to mathematical economy. He is mainly celebrated as the first persons to defin…
Herbert Gr{ö}tzsch is the main founder of the theory of quasicon-formal mappings. We review five of his papers, written between 1928 and 1932, that show the progress of his work from conformal to quasiconformal geometry. This will give an idea of his motivation for introducing quasicon-formal mappings, of the problems …
The paper proves rigidity and uniformization theorems for infinite circle patterns and convex polyhedra in hyperbolic 3-space.
We study pricing and (super)hedging for American options in an imperfect market model with default, where the imperfections are taken into account via the nonlinearity of the wealth dynamics. The payoff is given by an RCLL adapted process . We define the {\em seller's superhedging price} of the American option a…
T. Mochizuki constructs a theory of variations of wild Hodge structure for which the underlying flat connection can have irregular singularities at infinity. He extends in this way the correspondence of Corlette and Simpson between irreducible flat bundles and stables Higgs bundles, taking into account objects with irr…
In this paper, the author considers the numerical computation of CVA for large systems by Mote Carlo methods. He introduces two types of stochastic mesh methods for the computations of CVA. In the first method, stochastic mesh method is used to obtain the future value of the derivative contracts. In the second method, …
In 2003, S.-s. Chern began a study of almost-complex structures on the 6-sphere, with the idea of exploiting the special properties of its well-known almost-complex structure invariant under the exceptional group . While he did not solve the (currently still open) problem of determining whether there exists an int…
Plane Delaunay triangulations are rigid under Luo's discrete conformal change.