New bound for neural networks with full-rank weights, independent of network width.
problem Understanding generalization of neural networks with full-rank weight matrices.
method Using Koopman operators to derive a tighter generalization bound for full-rank weight matrices.
result The bound is tighter than existing norm-based bounds when condition numbers are small.
Completes the space of vector-valued one-forms on manifolds.
problem Metric incompleteness of the space of full-ranked one-forms.
method Distance equality and quotient structures.
result Concrete description of the metric completion of the space of full-ranked one-forms.
Study Betti and Hodge numbers of solvmanifolds from integer polynomials.
problem Computing Betti and Hodge numbers of solvmanifolds constructed from integer polynomials.
method Analyzing de Rham and Dolbeault cohomology of solvmanifolds under algebraic conditions.
result Explicit generating polynomials for Hodge numbers in quasi full rank case.
We analyze incomplete ranking data, modeling coarsening and studying rank aggregation methods.
problem Statistical inference for incomplete ranking data, especially under rank-dependent coarsening.
method Modeling rank-dependent coarsening, studying Plackett-Luce distribution, and analyzing rank aggregation methods.
result The ability to recover a target ranking from incomplete observations, despite coarsening bias, is theoretically addressed.
New metrics defined for full-rank correlation matrices, ensuring unique operations.
problem No suitable problem statement as the abstract does not describe a problem to be solved.
method New Riemannian metrics defined on full-rank correlation matrices, providing unique operations.
result Unique Riemannian logarithm and Fréchet mean defined for full-rank correlation matrices.
New methods provide stable ranking without assumptions on data distributions.
problem Stability issues in ranking problems with noisy data.
method Developed a stability framework and two ranking operators.
result Guaranteed stability without assumptions on data distributions.
Slow feature analysis (SFA) is a method for extracting slowly varying features from a quickly varying multidimensional signal. An open source Matlab-implementation sfa-tk makes SFA easily useable. We show here that under certain circumstances, namely when the covariance matrix of the nonlinearly expanded data does not …
This article addresses the modeling of reverberant recording environments in the context of under-determined convolutive blind source separation. We model the contribution of each source to all mixture channels in the time-frequency domain as a zero-mean Gaussian random variable whose covariance encodes the spatial cha…
Method quantifies uncertainty in full ranking with new CP approach.
problem Uncertainty quantification in full ranking with unknown ground truth.
method Transductive Conformal Prediction (CP) method to construct distribution-free bounds.
result Valid prediction sets and false coverage proportion control for full ranking.
Study of Eisenstein series linked to hyperbolic cusps.
problem Understanding Eisenstein series associated with hyperbolic cusps.
method Analyzing cohomology classes and intertwining operators.
result Different cusps correspond to linearly independent cohomology classes.
New model for high rank matrix completion with online and batch methods.
problem Matrix completion for high rank matrices with latent structure.
method Kernel trick to map data into a high dimensional feature space, explicit parametrization of low dimensional subspace, online fitting procedure.
result Online method can handle streaming data and adapt to non-stationary latent structure.
The expressive power of a Gaussian process (GP) model comes at a cost of poor scalability in the data size. To improve its scalability, this paper presents a low-rank-cum-Markov approximation (LMA) of the GP model that is novel in leveraging the dual computational advantages stemming from complementing a low-rank appro…
A fast method for multichannel source separation using jointly diagonalizable SCMs.
problem Computational inefficiency and poor performance in multichannel source separation.
method Restricts SCMs to jointly-diagonalizable but full-rank matrices, proposing efficient algorithms.
result Significant speedup and improved performance compared to original methods.
The paper introduces structured variational families to improve scalability in black-box variational inference.
problem Scalability issues in black-box variational inference, especially for large datasets and hierarchical models.
method Developed structured variational families that achieve better iteration complexity of O(N) compared to full-rank families.
result Structured variational families can achieve better scaling with respect to dataset size N, improving iteration complexity from O(N^2) to O(N).
Simple algorithms identify best items or full rankings from choice-based feedback.
problem Learning to identify the best item or full ranking from choice-based feedback.
method Nested Elimination (NE) and Nested Partition (NP) algorithms.
result NE is worst-case asymptotically optimal, NP is optimal up to a constant factor.
New algorithm clusters data from full or incomplete datasets.
problem Clustering incomplete data with standard methods.
method Fusion penalties for subspace clustering.
result Approach performs comparably to state-of-the-art with complete data, and better with missing data.
Paper solves Christoffel-Minkowski and Weingarten curvature problems in hyperbolic space.
problem Christoffel-Minkowski and Weingarten curvature problems in hyperbolic space.
method Proved existence of solutions using a new full rank theorem.
result Existence of smooth, origin-symmetric, strictly horospherically convex solutions.
An isomorphism of symplectically tame smooth pseudocomplex structures on the complex projective plane which is a homeomorphism and differentiable of full rank at two points is smooth.
Determinantal point processes (DPPs) have garnered attention as an elegant probabilistic model of set diversity. They are useful for a number of subset selection tasks, including product recommendation. DPPs are parametrized by a positive semi-definite kernel matrix. In this work we present a new method for learning th…
New algorithm ranks players from partial comparisons with optimal rate.
problem Ranking players from partial pairwise comparisons.
method Divide-and-conquer approach, local MLE within groups.
result Optimal ranking algorithm with minimax rate.
In this paper, we develop a relative error bound for nuclear norm regularized matrix completion, with the focus on the completion of full-rank matrices. Under the assumption that the top eigenspaces of the target matrix are incoherent, we derive a relative upper bound for recovering the best low-rank approximation of t…
This paper solves the Christoffel problem in hyperbolic space and its equivalent on spheres.
problem Prescribing curvatures for convex hypersurfaces in hyperbolic space.
method Proving a full rank theorem to establish the existence of solutions.
result Existence of solutions to the Christoffel problem and its equivalent Nirenberg-Kazdan-Warner problem on spheres.
Deep linear networks avoid spurious local minima under certain conditions.
problem Existence of spurious local minima in deep linear networks.
method Reduction to two-layer case, quadratic loss analysis, and perturbation argument to show full rank property.
result Deep linear networks have no spurious local minima under specific conditions.
Optimal smooth subspaces approximate large data sets efficiently.
problem Approximating large data sets with invariant subspaces.
method Smooth functions under lattice translations or crystallographic groups, with optimal selection of Paley-Wiener space.
result Optimal lattice selection enhances approximation efficiency.
Geodesics found in deep linear networks.
problem Finding shortest paths in deep neural networks.
method Derived ODEs and explicit solutions for geodesics.
result Horizontal straight lines are geodesics in invariant manifold.
New insights into attention mechanisms reveal dramatic trade-offs between rank and heads.
problem Dramatic trade-offs between rank and number of heads in attention mechanisms.
method Presented a simple target function and proved theoretical limits.
result Full-rank attention is necessary for long contexts, while low-rank is sufficient for short ones.
Determinantal point processes (DPPs) are an elegant model for encoding probabilities over subsets, such as shopping baskets, of a ground set, such as an item catalog. They are useful for a number of machine learning tasks, including product recommendation. DPPs are parametrized by a positive semi-definite kernel matrix…
Simple perturbation of Vafa-Witten equations leads to transversality.
problem Transversality of Vafa-Witten moduli space.
method Simple perturbation of Vafa-Witten equations, proving transversality for SU(2) or SO(3) structure groups. result For generic perturbation parameter, the full rank part of the moduli space satisfies transversality.
New metrics improve landing algorithms for orthogonality constraints.
problem Optimizing landing algorithms with orthogonality constraints.
method Proposed a family of metrics over full-rank matrices to enhance landing algorithms.
result Natural extension of β-metric improves landing performance.
The paper analyzes convergence properties of NGA and PAMe for L1-norm PCA.
problem Finite-step convergence of L1-norm PCA algorithms. method Conditional subgradient and alternating maximization interpretations of NGA, and PAMe with extrapolation.
result Iterative points of modified NGA and PAMe remain constant after finitely many steps under certain conditions.
The paper addresses calibration in label ranking, a structured prediction task.
problem Calibration in label ranking is not well understood and often poorly calibrated.
method Formalized calibration for label ranking, developed a hierarchy of notions, and empirically evaluated models.
result Popular label ranking models are often poorly calibrated, with differences between sub-ranking and top-k metrics.
Gradient descent works well for large NNs due to convexity in a transformed space.
problem Why gradient descent works well in non-convex NN optimization.
method Introduced canonical space and disparity matrix to prove convexity.
result Gradient descent converges to global minimum in large NNs.
Algorithm learns two-layer residual units using ReLU activations from samples.
problem Learning two-layer residual units from samples.
method Design layer-wise objectives as functionals, formulate ERM as QP, solve using LP, prove statistical consistency.
result Strong statistical consistency and robustness of the algorithm.
New method clusters incomplete data by fusing subspaces.
problem Learning low-dimensional structures from highly incomplete data.
method Assign each datum to its own subspace, then fuse subspaces of the same cluster.
result Our method performs comparably to state-of-the-art with complete data and better with missing data.
A necessary condition for a connection in a vector bundle to be locally metric is for its curvature matrix, which consists of 2 forms, to be skew symmetric with respect to some local frame. In this paper we give a simple algorithm that can be used to decide when a matrix of 2 forms is equivalent to a skew symmetric…
Deep ReLU networks with extra parameters have mostly good loss landscapes.
problem Finding good local minima in the loss landscape of deep neural networks.
method Analyzing shallow and deep ReLU networks with extra parameters on a generic dataset.
result Most activation patterns correspond to regions with no bad local minima.
Gradient descent recovers planted weights in shallow neural networks with quadratic activations.
problem Learning shallow neural networks with quadratic activations and planted weights.
method Analysis of optimization landscape, gradient descent, semicircle law for Wishart ensemble.
result Gradient descent can recover planted weights if initialized below an energy barrier.
Gradient flow in parameters equals linear interpolation in outputs.
problem Understanding and optimizing training algorithms in deep learning.
method Proving equivalence between gradient flow in parameter space and linear interpolation in output space, and deriving formulas for global minima.
result Gradient flow in parameters can be transformed into linear interpolation in outputs, leading to global minima.
New method for initializing low-rank neural networks improves performance.
problem Training low-rank neural networks efficiently and accurately.
method Inspired by function approximation, proposes a novel low-rank initialization framework.
result Demonstrates significant gap between spectral and low-rank initialization approaches.
Compressing data helps learn Mahalanobis metrics effectively.
problem Learning Mahalanobis metrics in high-dimensional spaces.
method Randomly compress data to train a full-rank metric in a reduced feature space.
result Theoretical guarantees on error for Mahalanobis metric learning, independent of ambient dimension.
The paper recovers missing data entries of high-rank matrices using polynomial polynomials.
problem Recovering missing entries of high-rank matrices with low intrinsic dimension.
method Developed a new polynomial matrix completion method using the kernel trick and relaxation of rank objective.
result Identified complete matrix of minimum intrinsic dimension by minimizing rank in high-dimensional feature space.
MRTL learns interpretable spatial patterns efficiently.
problem Efficient and interpretable spatial analysis in various fields.
method Multiresolution Tensor Learning (MRTL) algorithm.
result 4~5x speedup with accurate and interpretable latent factors.
The paper tackles sparse graph learning under Laplacian-related constraints, improving upon existing methods.
problem Learning a sparse undirected graph from multivariate data under Laplacian-related constraints.
method Modifications to penalized log-likelihood approaches to enforce total positivity and lasso/adaptive lasso penalties using ADMM.
result The proposed constrained adaptive lasso approach significantly outperforms existing Laplacian-based approaches.
Matrices of (approximate) low rank are pervasive in data science, appearing in recommender systems, movie preferences, topic models, medical records, and genomics. While there is a vast literature on how to exploit low rank structure in these datasets, there is less attention on explaining why the low rank structure ap…
The paper analyzes privacy leakage in federated learning using linear algebra and optimization theory.
problem Privacy leakage in federated learning despite its promise for data privacy.
method Theoretical analysis from linear algebra and optimization theory perspectives.
result Derives sufficient conditions to prevent data reconstruction attacks and establishes an upper bound on privacy leakage.
New criterion ensures recovery of latent factors in NMF with mild conditions.
problem Identifying latent factors in nonnegative matrix factorization (NMF) under mild conditions.
method Proposed a new identification criterion based on the scatteredness of one factor's rows in the nonnegative orthant.
result Latent factors can be provably identified from the NMF model with minimal structural assumptions.
This paper explores approximations for fully Bayesian Gaussian Process Regression.
problem Learning in Gaussian Process models through hyperparameter adaptation.
method Two approximation schemes: Hamiltonian Monte Carlo and Variational Inference.
result Predictive performance analysis on various benchmark datasets.
Study on the limits of learning HMM parameters under various conditions.
problem Understanding the conditions under which hidden Markov model parameters can be learned.
method Nonasymptotic minimax upper and lower bounds, thresholds analysis.
result Nonasymptotic minimax bounds match up to constants, showing learnable thresholds.