Unified framework for nonconvex matrix completion with linearly parameterized factors.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
PrecGD restores linear convergence in over-parameterized nonconvex matrix factorization.
This paper improves reinforcement learning efficiency for large-scale MDPs.
We study the linear contextual bandit problem with finite action sets. When the problem dimension is , the time horizon is , and there are candidate actions per time period, we (1) show that the minimax expected regret is for every algorithm, and (2) introduce a V…
Policy gradient converges linearly with Hadamard parameterization in tabular settings.
Small initialization improves tensor recovery from noisy data.
We show that there is a family of pseudo-Anosov braids independently parameterized by the braid index and the (canonical) length whose smallest conjugacy invariant sets grow exponentially in the braid index and linearly in the length and conclude that the conjugacy problem remains exponential in the braid index under t…
This work improves the lottery ticket hypothesis by reducing over-parameterization requirement.
In this paper, we theoretically prove that gradient descent can find a global minimum of non-convex optimization of all layers for nonlinear deep neural networks of sizes commonly encountered in practice. The theory developed in this paper only requires the practical degrees of over-parameterization unlike previous the…
The paper axiomatizes strong emergence in parameterized field theories and proves existence theorems.
DMSTF models spatio-temporal data with deep Markov priors.
New methods solve tensor-on-tensor regression with unknown rank, revealing benefits of over-parameterization.
Gradient descent converges linearly for neural networks with specific conditions.
Variational Bayesian Inference is a popular methodology for approximating posterior distributions over Bayesian neural network weights. Recent work developing this class of methods has explored ever richer parameterizations of the approximate posterior in the hope of improving performance. In contrast, here we share a …
A parsimonious model reduces over-parameterization in skewed matrix variate mixtures.
We study the generalization properties of stochastic gradient methods for learning with convex loss functions and linearly parameterized functions. We show that, in the absence of penalizations or constraints, the stability and approximation properties of the algorithm can be controlled by tuning either the step-size o…
We reveal a model rank that predicts successful recovery of target functions at overparameterization.
Global results are proved about the way in which Boyland's forcing partial order organizes a set of braid types: those of periodic orbits of Smale's horseshoe map for which the associated train track is a star. This is a special case of a conjecture introduced in a previous paper, which claims that forcing organizes al…
New method tackles over-parameterized matrix sensing with FGD, improving statistical and computational complexity.
Variational inference offers scalable and flexible tools to tackle intractable Bayesian inference of modern statistical models like Bayesian neural networks and Gaussian processes. For largely over-parameterized models, however, the over-regularization property of the variational objective makes the application of vari…
We consider the problem of clustering a set of high-dimensional data points into sets of low-dimensional linear subspaces. The number of subspaces, their dimensions, and their orientations are unknown. We propose a simple and low-complexity clustering algorithm based on thresholding the correlations between the data po…
Binary data matrices can represent many types of data such as social networks, votes, or gene expression. In some cases, the analysis of binary matrices can be tackled with nonnegative matrix factorization (NMF), where the observed data matrix is approximated by the product of two smaller nonnegative matrices. In this …
FedAvg converges linearly to global minimum in federated learning with partial participation.
In this paper, we introduce a parameterized discrete curvature (-curvature) for piecewise linear metrics on polyhedral surfaces, which is a generalization of the classical discrete curvature. A discrete uniformization theorem is established for the parameterized discrete curvature, which generalizes the discrete uni…
Gradient descent with small initialization solves matrix completion without regularization.
Pruning improves model generalization in over-parameterized models, contradicting traditional theories.
Motivated by models of human decision making proposed to explain commonly observed deviations from conventional expected value preferences, we formulate two stochastic multi-armed bandit problems with distorted probabilities on the reward distributions: the classic -armed bandit and the linearly parameterized bandit…
We consider reinforcement learning in parameterized Markov Decision Processes (MDPs), where the parameterization may induce correlation across transition probabilities or rewards. Consequently, observing a particular state transition might yield useful information about other, unobserved, parts of the MDP. We present a…
We consider whether algorithmic choices in over-parameterized linear matrix factorization introduce implicit regularization. We focus on noiseless matrix sensing over rank- positive semi-definite (PSD) matrices in , with a sensing mechanism that satisfies restricted isometry properties (RIP)…
We propose randomized least-squares value iteration (RLSVI) -- a new reinforcement learning algorithm designed to explore and generalize efficiently via linearly parameterized value functions. We explain why versions of least-squares value iteration that use Boltzmann or epsilon-greedy exploration can be highly ineffic…
Multiresolution Matrix Factorization (MMF) was recently introduced as a method for finding multiscale structure and defining wavelets on graphs/matrices. In this paper we derive pMMF, a parallel algorithm for computing the MMF factorization. Empirically, the running time of pMMF scales linearly in the dimension for spa…
Boolean matrix factorization and Boolean matrix completion from noisy observations are desirable unsupervised data-analysis methods due to their interpretability, but hard to perform due to their NP-hardness. We treat these problems as maximum a posteriori inference problems in a graphical model and present a message p…
Geometrically classifies maps from R^0|2 to any manifold, unifying theories.
Entropy-regularized NPG methods converge linearly in discounted MDPs.
We show that strongly contracting geodesics in Outer space project to parameterized quasigeodesics in the free factor complex. This result provides a converse to a theorem of Bestvina--Feighn, and is used to give conditions for when a subgroup of has a quasi-isometric orbit map into the free …
Graded Transformers embed algebraic structure in neural networks through graded transformations.
We address the rectangular matrix completion problem by lifting the unknown matrix to a positive semidefinite matrix in higher dimension, and optimizing a nonconvex objective over the semidefinite factor using a simple gradient descent scheme. With random observations of a $n_1 \times n…
The paper analyzes how over-parameterization affects GD convergence in matrix sensing problems.
New algorithms improve binary neural network configurations.
Consider a Markov decision process (MDP) that admits a set of state-action features, which can linearly express the process's probabilistic transition model. We propose a parametric Q-learning algorithm that finds an approximate-optimal policy using a sample size proportional to the feature dimension and invariant …
Over-parameterized models, such as DeepNets and ConvNets, form a class of models that are routinely adopted in a wide variety of applications, and for which Bayesian inference is desirable but extremely challenging. Variational inference offers the tools to tackle this challenge in a scalable way and with some degree o…
We show that the gradient descent algorithm provides an implicit regularization effect in the learning of over-parameterized matrix factorization models and one-hidden-layer neural networks with quadratic activations. Concretely, we show that given random linear measurements of a rank positive s…
Improves conditions for mode connectivity in deep neural networks.
We seek to improve the data efficiency of neural networks and present novel implementations of parameterized piece-wise polynomial activation functions. The parameters are the y-coordinates of n+1 Chebyshev nodes per hidden unit and Lagrangian interpolation between the nodes produces the polynomial on [-1, 1]. We show …
APGD algorithm efficiently recovers over-parameterized matrices from noisy measurements.
BPNNs learn to solve combinatorial problems faster and more accurately.
We consider probabilistic PCA and related factor models from a Bayesian perspective. These models are in general not identifiable as the likelihood has a rotational symmetry. This gives rise to complicated posterior distributions with continuous subspaces of equal density and thus hinders efficiency of inference as wel…
The optimization of multilayer neural networks typically leads to a solution with zero training error, yet the landscape can exhibit spurious local minima and the minima can be disconnected. In this paper, we shed light on this phenomenon: we show that the combination of stochastic gradient descent (SGD) and over-param…