Method measures weight similarity in neural networks using normalization and statistical inference.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
GAN normalizes CT scans for consistent radiomic feature values.
Traditionally, multi-layer neural networks use dot product between the output vector of previous layer and the incoming weight vector as the input to activation function. The result of dot product is unbounded, thus increases the risk of large variance. Large variance of neuron makes the model sensitive to the change o…
Proves a theorem similar to Moser's using a normalization method.
I classify spacelike self-similar shrinking solutions of the mean curvature flow in pseudo-euclidean space in arbitrary codimension, if the mean curvature vector is not a null vector and the principal normal vector is parallel in the normal bundle. Moreover, I exclude the existence of such self-shrinkers in several cas…
We present a new similarity measure based on information theoretic measures which is superior than Normalized Compression Distance for clustering problems and inherits the useful properties of conditional Kolmogorov complexity. We show that Normalized Compression Dictionary Size and Normalized Compression Dictionary En…
A major breakthrough in the theory of topological algorithms occurred in 1992 when Hyam Rubinstein introduced the idea of an almost normal surface. We explain how almost normal surfaces emerged naturally from the study of geodesics and minimal surfaces. Patterns of stable and unstable geodesics can be used to character…
A virtual link diagram is called normal if the associated abstract link diagram is checkerboard colorable, and a virtual link is normal if it has a normal diagram as a representative. Normal virtual links have some properties similar to classical links.In this paper, we introduce a method of converting a virtual link d…
A new method speeds up SoftMax normalization for embedding learning.
Distance, normals, and double normals for real plane curves with singularities
Topological normal generation proved for mapping class groups of certain surfaces.
We propose a theoretical framework for thinking about score normalization, which confirms that normalization is not needed under (admittedly fragile) ideal conditions. If, however, these conditions are not met, e.g. under data-set shift between training and runtime, our theory reveals dependencies between scores that c…
The paper studies estimating the normalizing constant using queries to a black-box function in RKHS.
Feature normalization prevents collapse in non-contrastive learning dynamics.
This paper reviews normalization techniques for DNNs.
Most network-based machine learning methods assume that the labels of two adjacent samples in the network are likely to be the same. However, assuming the pairwise relationship between samples is not complete. The information a group of samples that shows very similar pattern and tends to have similar labels is missed.…
From a sequence of similarity networks, with edges representing certain similarity measures between nodes, we are interested in detecting a change-point which changes the statistical property of the networks. After the change, a subset of anomalous nodes which compares dissimilarly with the normal nodes. We study a sim…
A 1-bridge torus knot in a 3-manifold of genus is a knot drawn on a Heegaard torus with one bridge. We give two types of normal forms to parameterize the family of 1-bridge torus knots that are similar to the Schubert's normal form and the Conway's normal form for 2-bridge knots. For a given Schubert's normal f…
We propose a family of near-metrics based on local graph diffusion to capture similarity for a wide class of data sets. These quasi-metametrics, as their names suggest, dispense with one or two standard axioms of metric spaces, specifically distinguishability and symmetry, so that similarity between data points of arbi…
We introduce the anti-profile Support Vector Machine (apSVM) as a novel algorithm to address the anomaly classification problem, an extension of anomaly detection where the goal is to distinguish data samples from a number of anomalous and heterogeneous classes based on their pattern of deviation from a normal stable c…
Spectral clustering is a technique that clusters elements using the top few eigenvectors of their (possibly normalized) similarity matrix. The quality of spectral clustering is closely tied to the convergence properties of these principal eigenvectors. This rate of convergence has been shown to be identical for both th…
Recent seminal work at the intersection of deep neural networks practice and random matrix theory has linked the convergence speed and robustness of these networks with the combination of random weight initialization and nonlinear activation function in use. Building on those principles, we introduce a process to trans…
In this paper we introduce the fourth fundamental form for the hypersurfaces in and the space-like hypersurfaces in and discuss the conformality of the normal Gauss maps of the hypersurfaces in and . Particularly, we discuss the surfaces with conformal normal Gauss maps in…
We investigate three-dimensional surfaces where the normal vector forms a constant angle with the radius vector. These surfaces naturally extend equiangular (logarithmic) spirals in the plane.
We present a new approximation to the normal distribution quantile function. It has a similar form to the approximation of Beasley and Springer [3], providing a maximum absolute error of less than . This is less accurate than [3], but still sufficient for many applications. However it is faster than …
The method of "random Fourier features (RFF)" has become a popular tool for approximating the "radial basis function (RBF)" kernel. The variance of RFF is actually large. Interestingly, the variance can be substantially reduced by a simple normalization step as we theoretically demonstrate. We name the improved scheme …
The paper argues that normalized mutual information is biased in clustering and community detection.
This paper uses linear rational splines for invertible modeling, offering a simpler inverse and similar costs.
Using geodesic currents, we provide a theoretical justification for some of the experimental results regarding the behavior of Whitehead's algorithm on non-minimal inputs, that were obtained by Haralick, Miasnikov and Myasnikov via pattern recognition methods. In particular we prove that the images of "random" elements…
Laplace kernel and Neural Tangent Kernels are shown to be nearly identical for normalized data.
For option pricing models and heavy-tailed distributions, this study proposes a continuous-time stochastic volatility model based on an arithmetic Brownian motion: a one-parameter extension of the normal stochastic alpha-beta-rho (SABR) model. Using two generalized Bougerol's identities in the literature, the study sho…
The paper defines -normality for contact and paracontact manifolds and explores their properties.
Unified toolkit for comparing neural representations using SRTD and NTS.
Stochastic approximation proves asymptotic normality for non-smooth problems.
We study the dynamics of the normal implied volatility in a local volatility model, using a small-time expansion in powers of maturity T. At leading order in this expansion, the asymptotics of the normal implied volatility is similar, up to a different definition of the moneyness, to that of the log-normal volatility. …
Unified understanding of neural representation similarity measures.
Laboratory test results are an important and generally high dimensional component of a patient's Electronic Health Record (EHR). We train embedding representations (via Word2Vec and GloVe) for LOINC codes of laboratory tests from the EHRs of about 80,000 patients at a cancer center. To include information about lab tes…
Fractal Lipschitz-Killing curvature measures C^f_k(F,.), k = 0, ..., d, are determined for a large class of self-similar sets F in R^d. They arise as weak limits of the appropriately rescaled classical Lipschitz-Killing curvature measures C_k(F_r,.) from geometric measure theory of parallel sets F_r for small distances…
Self Normalizing Flows improve normalizing flows by reducing computational complexity.
In recent years, deep metric learning has achieved promising results in learning high dimensional semantic feature embeddings where the spatial relationships of the feature vectors match the visual similarities of the images. Similarity search for images is performed by determining the vectors with the smallest distanc…
Normalized nonnegative models assign probability distributions to users and random variables to items; see [Stark, 2015]. Rating an item is regarded as sampling the random variable assigned to the item with respect to the distribution assigned to the user who rates the item. Models of that kind are highly expressive. F…
Firm size data usually do not show the normality that is often assumed in statistical analysis such as regression analysis. In this study we focus on two firm size data: the number of employees and sale. Those data deviate considerably from a normal distribution. To improve the normality of those data we transform them…
Study on self-similar surfaces and their mapping class groups generated by involutions.
New variational flows improve Monte Carlo and normalization tasks.
Most network-based protein (or gene) function prediction methods are based on the assumption that the labels of two adjacent proteins in the network are likely to be the same. However, assuming the pairwise relationship between proteins or genes is not complete, the information a group of genes that show very similar p…
Log-Normal Multiplicative Dynamics improves low-precision training of neural networks.
Doubly-stochastic normalization improves robustness to heteroskedastic noise.
We investigate how the final parameters found by stochastic gradient descent are influenced by over-parameterization. We generate families of models by increasing the number of channels in a base network, and then perform a large hyper-parameter search to study how the test error depends on learning rate, batch size, a…