Study excess capacity in neural networks using Rademacher complexity.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Let be a complete non-compact Riemannian manifold together with a function , which weights the Hausdorff measures associated to the Riemannian metric. In this work we assume lower or upper radial bounds on some weighted or unweighted curvatures of to deduce comparisons for the weighted isoperimetric qu…
This paper presents a general framework for norm-based capacity control for weight normalized deep neural networks. We establish the upper bound on the Rademacher complexities of this family. With an normalization where , and , we discuss properties of a width-independent ca…
Study on neural networks' storage capacity and solution space structure.
Unweighted matrix factorization can match or outperform weighted methods in recommender systems.
DDN dynamically combines weights for domain-specific models, improving performance.
Dropout controls model capacity in deep learning and matrix completion.
This paper proposes a method to improve few-shot learning by generating multi-level weight-centric features.
Capacity-Constrained Online Convex Optimization with Delayed Feedback
Classical results on the statistical complexity of linear models have commonly identified the norm of the weights as a fundamental capacity measure. Generalizations of this measure to the setting of deep networks have been varied, though a frequently identified quantity is the product of weight norms of each la…
In this paper, we propose the nonlinearity generation method to speed up and stabilize the training of deep convolutional neural networks. The proposed method modifies a family of activation functions as nonlinearity generators (NGs). NGs make the activation functions linear symmetric for their inputs to lower model ca…
We study the computational capacity of a model neuron, the Tempotron, which classifies sequences of spikes by linear-threshold operations. We use statistical mechanics and extreme value theory to derive the capacity of the system in random classification tasks. In contrast to its static analog, the Perceptron, the Temp…
Extends DAMs to Gaussian distributions for efficient pattern storage and retrieval.
A long standing open problem in the theory of neural networks is the development of quantitative methods to estimate and compare the capabilities of different architectures. Here we define the capacity of an architecture by the binary logarithm of the number of functions it can compute, as the synaptic weights are vari…
We improve deep threshold networks' memorization capacity exponentially.
Nonlinear RNNs' memory capacity varies widely, making it impractical.
We show that the non pluripolar product of positive currents is a bimeromorphic invariant. Under some natural assumptions, we show that the (weighted) energy associated to big cohomology classes are also bimeromorphic invariants. We compare the weighted energy functionals of currents with respect to different cohomolog…
gLSTM improves graph neural networks by increasing storage capacity to prevent over-squashing.
AON improves neural network generalization by making weights approximately orthogonal.
A new model clusters network nodes based on relative edge weights.
Study on CNNs' learning rates and approximation capacities.
Rectified Linear Units (ReLU) have become the main model for the neural units in current deep learning systems. This choice has been originally suggested as a way to compensate for the so called vanishing gradient problem which can undercut stochastic gradient descent (SGD) learning in networks composed of multiple lay…
Theory of learning with weight-distribution constraints.
Many state-of-the-art results obtained with deep networks are achieved with the largest models that could be trained, and if more computation power was available, we might be able to exploit much larger datasets in order to improve generalization ability. Whereas in learning algorithms such as decision trees the ratio …
Investment decisions shift earlier as patience decreases, with implications for pasting conditions.
Generative approach speeds hyperparameter tuning for machine learning models.
Understanding how neural networks learn remains one of the central challenges in machine learning research. From random at the start of training, the weights of a neural network evolve in such a way as to be able to perform a variety of tasks, like classifying images. Here we study the emergence of structure in the wei…
Sparse codes improve optimal control tasks with correlated inputs.
Following the recent work on capacity allocation, we formulate the conjecture that the shattering problem in deep neural networks can only be avoided if the capacity propagation through layers has a non-degenerate continuous limit when the number of layers tends to infinity. This allows us to study a number of commonly…
An implicit goal in works on deep generative models is that such models should be able to generate novel examples that were not previously seen in the training data. In this paper, we investigate to what extent this property holds for widely employed variational autoencoder (VAE) architectures. VAEs maximize a lower bo…
Study bounds Rademacher complexity of Fourier neural operators.
Let be a compact Kähler manifold and $\om$ a smooth closed form of bidegree which is nonnegative and big. We study the classes ${\mathcal E}_χ(X,\om)$ of $\om$-plurisubharmonic functions of finite weighted Monge-Ampère energy. When the weight has fast growth at infinity, the corresponding functions are …
Corrects distribution shift in target shift scenarios using importance weighting.
Spectral algorithms improve under covariate shift with novel weighted techniques.
In this article, we propose the notion of the general -affine capacity and prove some basic properties for the general -affine capacity, such as affine invariance and monotonicity. The newly proposed general -affine capacity is compared with several classical geometric quantities, e.g., the volume, the -var…
We study various capacities on compact Kähler manifolds which generalize the Bedford-Taylor Monge-Ampère capacity. We then use these capacities to study the existence and the regularity of solutions of complex Monge-Ampère equations.
Solves a discrete logarithmic Minkowski problem for electrostatic p-capacity.
MeliusNet improves binary neural networks to match MobileNet-v1 accuracy.
Exploiting different representations, or views, of the same object for better clustering has become very popular these days, which is conventionally called multi-view clustering. Generally, it is essential to measure the importance of each individual view, due to some noises, or inherent capacities in description. Many…
New regularizer improves neural network robustness and generalization.
We continue our study of the Complex Monge-Ampère Operator on the Weighted Pluricomplex energy classes. We give more characterizations of the range of the classes by the Complex Monge-Ampère Operator. In particular, we prove that a non-negative Borel measure is the Monge-Ampère of a unique function …
A new neural network model using weighted Lehmer means and multiplets.
The variational autoencoder (VAE; Kingma, Welling (2014)) is a recently proposed generative model pairing a top-down generative network with a bottom-up recognition network which approximates posterior inference. It typically makes strong assumptions about posterior inference, for instance that the posterior distributi…
Given two or more Deep Neural Networks (DNNs) with the same or similar architectures, and trained on the same dataset, but trained with different solvers, parameters, hyper-parameters, regularization, etc., can we predict which DNN will have the best test accuracy, and can we do so without peeking at the test data? In …
We introduce the concept of pseudo symplectic capacities which is a mild generalization of that of symplectic capacities. As a generalization of the Hofer-Zehnder capacity we construct a Hofer-Zehnder type pseudo symplectic capacity and estimate it in terms of Gromov-Witten invariants. The (pseudo) symplectic capacitie…
A new memory system handles non-stationary environments by self-sizing and retaining memories.
Generalizes memory and forecasting capacities for nonlinear recurrent networks with dependent inputs.
New cyclicity measures defined in weighted Besov spaces, with stability and geometric analysis.