New bound for neural networks with full-rank weights, independent of network width.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Completes the space of vector-valued one-forms on manifolds.
Study Betti and Hodge numbers of solvmanifolds from integer polynomials.
New metrics defined for full-rank correlation matrices, ensuring unique operations.
We consider the teacher-student setting of learning shallow neural networks with quadratic activations and planted weight matrix , where is the width of the hidden layer and is the data dimension. We study the optimization landscape associated with the empirical and the popula…
New method for initializing low-rank neural networks improves performance.
Theoretical analysis of the error landscape of deep neural networks has garnered significant interest in recent years. In this work, we theoretically study the importance of noise in the trajectories of gradient descent towards optimal solutions in multi-layer neural networks. We show that adding noise (in different wa…
This paper describes a versatile method that accelerates multichannel source separation methods based on full-rank spatial modeling. A popular approach to multichannel source separation is to integrate a spatial model with a source model for estimating the spatial covariance matrices (SCMs) and power spectral densities…
Study of Eisenstein series linked to hyperbolic cusps.
We consider the problem of statistical inference for ranking data, specifically rank aggregation, under the assumption that samples are incomplete in the sense of not comprising all choice alternatives. In contrast to most existing methods, we explicitly model the process of turning a full ranking into an incomplete on…
The paper introduces structured variational families to improve scalability in black-box variational inference.
Optimizes wide low-rank neural networks for reduced parameters and cost.
An isomorphism of symplectically tame smooth pseudocomplex structures on the complex projective plane which is a homeomorphism and differentiable of full rank at two points is smooth.
Paper solves Christoffel-Minkowski and Weingarten curvature problems in hyperbolic space.
Method quantifies uncertainty in full ranking with new CP approach.
This article addresses the modeling of reverberant recording environments in the context of under-determined convolutive blind source separation. We model the contribution of each source to all mixture channels in the time-frequency domain as a zero-mean Gaussian random variable whose covariance encodes the spatial cha…
The paper analyzes how low-rank layers in neural networks improve generalization.
PLD distills knowledge using choice-theoretic Plackett-Luce model.
In this paper, we develop a relative error bound for nuclear norm regularized matrix completion, with the focus on the completion of full-rank matrices. Under the assumption that the top eigenspaces of the target matrix are incoherent, we derive a relative upper bound for recovering the best low-rank approximation of t…
Slow feature analysis (SFA) is a method for extracting slowly varying features from a quickly varying multidimensional signal. An open source Matlab-implementation sfa-tk makes SFA easily useable. We show here that under certain circumstances, namely when the covariance matrix of the nonlinearly expanded data does not …
Simple algorithms identify best items or full rankings from choice-based feedback.
This paper solves the Christoffel problem in hyperbolic space and its equivalent on spheres.
New methods provide stable ranking without assumptions on data distributions.
The expressive power of a Gaussian process (GP) model comes at a cost of poor scalability in the data size. To improve its scalability, this paper presents a low-rank-cum-Markov approximation (LMA) of the GP model that is novel in leveraging the dual computational advantages stemming from complementing a low-rank appro…
Recent advances in matrix completion enable data imputation in full-rank matrices by exploiting low dimensional (nonlinear) latent structure. In this paper, we develop a new model for high rank matrix completion (HRMC), together with batch and online methods to fit the model and out-of-sample extension to complete new …
PLUMAGE improves large model training efficiency and stability.
Geodesics found in deep linear networks.
Low-rank MPPCA improves importance sampling in high dimensions.
In this paper, we present some theoretical work to explain why simple gradient descent methods are so successful in solving non-convex optimization problems in learning large-scale neural networks (NN). After introducing a mathematical tool called canonical space, we have proved that the objective functions in learning…
Determinantal point processes (DPPs) have garnered attention as an elegant probabilistic model of set diversity. They are useful for a number of subset selection tasks, including product recommendation. DPPs are parametrized by a positive semi-definite kernel matrix. In this work we present a new method for learning th…
New insights into attention mechanisms reveal dramatic trade-offs between rank and heads.
New algorithm ranks players from partial comparisons with optimal rate.
Identifying components and estimating mixing weights in unlabeled finite mixtures under marginal independence.
Analyzes Hessian spectrum for neural networks near optimal learning.
Simple perturbation of Vafa-Witten equations leads to transversality.
New metrics improve landing algorithms for orthogonality constraints.
High-dimensional linear classifiers, such as the support vector machine (SVM) and distance weighted discrimination (DWD), are commonly used in biomedical research to distinguish groups of subjects based on a large number of features. However, their use is limited to applications where a single vector of features is mea…
We consider the problem of learning a one-hidden-layer neural network: we assume the input is from Gaussian distribution and the label , where is a nonnegative vector in with , is a full-rank weight matrix, and is a n…
The paper analyzes convergence properties of NGA and PAMe for -norm PCA.
The paper addresses calibration in label ranking, a structured prediction task.
A necessary condition for a connection in a vector bundle to be locally metric is for its curvature matrix, which consists of forms, to be skew symmetric with respect to some local frame. In this paper we give a simple algorithm that can be used to decide when a matrix of forms is equivalent to a skew symmetric…
In this letter, we propose a new identification criterion that guarantees the recovery of the low-rank latent factors in the nonnegative matrix factorization (NMF) model, under mild conditions. Specifically, using the proposed criterion, it suffices to identify the latent factors if the rows of one factor are \emph{suf…
Determinantal point processes (DPPs) are an elegant model for encoding probabilities over subsets, such as shopping baskets, of a ground set, such as an item catalog. They are useful for a number of machine learning tasks, including product recommendation. DPPs are parametrized by a positive semi-definite kernel matrix…
SVD training reduces DNN rank and computation load without SVD per step.
Compressing DNNs is important for the real-world applications operating on resource-constrained devices. However, we typically observe drastic performance deterioration when changing model size after training is completed. Therefore, retraining is required to resume the performance of the compressed models suitable for…
Algorithm learns two-layer residual units using ReLU activations from samples.
Optimal smooth subspaces approximate large data sets efficiently.
Sparse codes improve optimal control tasks with correlated inputs.