The paper explains the richness scale of wide neural networks.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study on how initialization scale affects neural network training regimes.
Simplifies deep learning scaling analysis without sacrificing accuracy.
The rich-get-richer mechanism (agents increase their ``wealth'' randomly at a rate proportional to their holdings) is often invoked to explain the Pareto power-law distribution observed in many physical situations, such as the degree distribution of growing scale free nets. We use two different analytical approaches, a…
A recent line of work studies overparametrized neural networks in the "kernel regime," i.e. when the network behaves during training as a kernelized linear predictor, and thus training with gradient descent has the effect of finding the minimum RKHS norm solution. This stands in contrast to other studies which demonstr…
Ridge regression reveals surprising high-dimensional behaviors via random matrix theory.
A fundamental challenge in developing high-impact machine learning technologies is balancing the need to model rich, structured domains with the ability to scale to big data. Many important problem areas are both richly structured and large scale, from social and biological networks, to knowledge graphs and the Web, to…
Let M be a compact manifold. We show the identity component of the group of self-homeomorphisms of M has a well-defined quasi-isometry type, and study its large scale geometry. Through examples, we relate this large scale geometry to both the topology of M and the dynamics of group actions on M. T…
Motivated by the widespread adoption of large-scale A/B testing in industry, we propose a new experimentation framework for the setting where potential experiments are abundant (i.e., many hypotheses are available to test), and observations are costly; we refer to this as the experiment-rich regime. Such scenarios requ…
Study shows optimal model performance at critical level of feature learning.
Transformers capture combinatorial tasks with bounded error and logarithmic sample dependence.
The Empirical Mode Decomposition (EMD) provides a tool to characterize time series in terms of its implicit components oscillating at different time-scales. We apply this decomposition to intraday time series of the following three financial indices: the S\&P 500 (USA), the IPC (Mexico) and the VIX (volatility index US…
Tools of the theory of critical phenomena, namely the scaling analysis and universality, are argued to be applicable to large complex web-like network structures. Using a detailed analysis of the real data of the International Trade Network we argue that the scaled link weight distribution has an approximate log-normal…
Multi-Entity Dependence Learning (MEDL) explores conditional correlations among multiple entities. The availability of rich contextual information requires a nimble learning scheme that tightly integrates with deep neural networks and has the ability to capture correlation structures among exponentially many outcomes. …
Study reveals how initialization scale affects training accuracy in linear networks.
We provide an exact solution to the ideal-gas-like models studied in econophysics to understand the microscopic origin of Pareto-law. In these class of models the key ingredient necessary for having a self-organized scale-free steady-state distribution is the trading or collision rule where agents or particles save a d…
In recent years, a rich variety of shrinkage priors have been proposed that have great promise in addressing massive regression problems. In general, these new priors can be expressed as scale mixtures of normals, but have more complex forms and better properties than traditional Cauchy and double exponential priors. W…
Optimization of neural networks scales with γ, revealing unique loss curves and optimal learning rates.
We present Blitzkriging, a new approach to fast inference for Gaussian processes, applicable to regression, optimisation and classification. State-of-the-art (stochastic) inference for Gaussian processes on very large datasets scales cubically in the number of 'inducing inputs', variables introduced to factorise the mo…
Paper introduces scalable clustering for large datasets with outliers.
We present an algorithm, HOMER, for exploration and reinforcement learning in rich observation environments that are summarizable by an unknown latent state space. The algorithm interleaves representation learning to identify a new notion of kinematic state abstraction with strategic exploration to reach new states usi…
We present Distributed Equivalent Substitution (DES) training, a novel distributed training framework for large-scale recommender systems with dynamic sparse features. DES introduces fully synchronous training to large-scale recommendation system for the first time by reducing communication, thus making the training of…
Grokking occurs when neural networks transition from lazy to rich training dynamics, fitting initial features before generalizing.
New metric measures dynamical richness without relying on accuracy.
Topological data analysis offers a rich source of valuable information to study vision problems. Yet, so far we lack a theoretically sound connection to popular kernel-based learning techniques, such as kernel SVMs or kernel PCA. In this work, we establish such a connection by designing a multi-scale kernel for persist…
This study examines cores within superclusters, highlighting their transitional nature and dynamical state.
Motivated by the rich geometry of conformal Riemannian manifolds and by the recent development of geometries modeled on homogeneous spaces with semisimple and parabolic, Weyl structures and preferred connections are introduced in this general framework. In particular, we extend the notions of scales, clos…
New shape representation for airfoils improves design and manufacturing.
Improves speaker verification for variable-duration utterances using a feature pyramid module.
Wide CNNs outperform infinite width networks, revealing scaling laws.
In the past few decades considerable effort has been expended in characterizing and modeling financial time series. A number of stylized facts have been identified, and volatility clustering or the tendency toward persistence has emerged as the central feature. In this paper we propose an appropriately defined conditio…
Probabilistic methods for classifying text form a rich tradition in machine learning and natural language processing. For many important problems, however, class prediction is uninteresting because the class is known, and instead the focus shifts to estimating latent quantities related to the text, such as affect or id…
Datasets are growing not just in size but in complexity, creating a demand for rich models and quantification of uncertainty. Bayesian methods are an excellent fit for this demand, but scaling Bayesian inference is a challenge. In response to this challenge, there has been considerable recent work based on varying assu…
Using the United Nations Commodity Trade Statistics Database [http://comtrade.un.org/db/] we construct the Google matrix of the world trade network and analyze its properties for various trade commodities for all countries and all available years from 1962 to 2009. The trade flows on this network are classified with th…
Continuous latent time series models are prevalent in Bayesian modeling; examples include the Kalman filter, dynamic collaborative filtering, or dynamic topic models. These models often benefit from structured, non mean field variational approximations that capture correlations between time steps. Black box variational…
DeformRS certifies deep networks against various input deformations.
MCFNet recovers spatial detail and fuses it with semantic information for real-time segmentation.
Using a model of wealth distribution where traders are characterized by quenched random saving propensities and trade among themselves by bipartite transactions, we mimic the enhanced rates of trading of the rich by introducing the preferential selection rule using a pair of continuously tunable parameters. The biparti…
The paper proposes an efficient method to scale Bayesian inference for mixed multinomial logit models to very large datasets.
Measuring entity relatedness is a fundamental task for many natural language processing and information retrieval applications. Prior work often studies entity relatedness in static settings and an unsupervised manner. However, entities in real-world are often involved in many different relationships, consequently enti…
SEMASIA provides a large dataset of latent representations for model comparison.
CHEER boosts poor models using rich model knowledge.
Graph embedding methods transform high-dimensional and complex graph contents into low-dimensional representations. They are useful for a wide range of graph analysis tasks including link prediction, node classification, recommendation and visualization. Most existing approaches represent graph nodes as point vectors i…
Consider a Riemannian metric on two-torus. We prove that the question of existence of polynomial first integrals leads naturally to a remarkable system of quasi-linear equations which turns out to be a Rich system of conservation laws. This reduces the question of integrability to the question of existence of smooth (q…
A dynamical model of capital exchange is introduced in which a specified amount of capital is exchanged between two individuals when they meet. The resulting time dependent wealth distributions are determined for a variety of exchange rules. For ``greedy'' exchange, an interaction between a rich and a poor individual r…
Kernel learning methods are among the most effective learning methods and have been vigorously studied in the past decades. However, when tackling with complicated tasks, classical kernel methods are not flexible or "rich" enough to describe the data and hence could not yield satisfactory performance. In this paper, vi…
Study explores properties of bipartite knots.
Generative adversarial networks (GAN) are a powerful subclass of generative models. Despite a very rich research activity leading to numerous interesting GAN algorithms, it is still very hard to assess which algorithm(s) perform better than others. We conduct a neutral, multi-faceted large-scale empirical study on stat…