Paper proposes ABDR for convex subspace clustering with adaptive block diagonal representation.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper solves matrix blind joint block diagonalization with noise.
Characterizes Anosov reducible representations in terms of eigenvalues.
Subspace clustering is a useful technique for many computer vision applications in which the intrinsic dimension of high-dimensional data is often smaller than the ambient dimension. Spectral clustering, as one of the main approaches to subspace clustering, often takes on a sparse representation or a low-rank represent…
Randomized block-diagonal preconditioning improves parallel learning convergence.
New theory allows simultaneous block-diagonalization of commuting operator fields.
Adaptive gradient approaches that automatically adjust the learning rate on a per-feature basis have been very popular for training deep networks. This rich class of algorithms includes Adagrad, RMSprop, Adam, and recent extensions. All these algorithms have adopted diagonal matrix adaptation, due to the prohibitive co…
We propose a modular extension of backpropagation for the computation of block-diagonal approximations to various curvature matrices of the training objective (in particular, the Hessian, generalized Gauss-Newton, and positive-curvature Hessian). The approach reduces the otherwise tedious manual derivation of these mat…
New insights into Hessian structure of neural networks reveal two forces.
Efficiently approximates Sparse PCA with significant speedups and minor error.
Deep latent-variable models learn representations of high-dimensional data in an unsupervised manner. A number of recent efforts have focused on learning representations that disentangle statistically independent axes of variation by introducing modifications to the standard objective function. These approaches general…
BEGIN network models binary data without parametric assumptions.
Localized sketching improves matrix multiplication and ridge regression complexity.
Second-order methods for neural network optimization have several advantages over methods based on first-order gradient descent, including better scaling to large mini-batch sizes and fewer updates needed for convergence. But they are rarely applied to deep learning in practice because of high computational cost and th…
Gaussian graphical models are widely utilized to infer and visualize networks of dependencies between continuous variables. However, inferring the graph is difficult when the sample size is small compared to the number of variables. To reduce the number of parameters to estimate in the model, we propose a non-asymptoti…
A Semi-Hidden Markov Model (SHMM) for bursty error channels is defined by a state transition probability matrix , a prior probability vector , and the state dependent output symbol error probability matrix . Several processes are utilized for estimating , and from a given empirically obtained or sim…
Improved sparse Gaussian processes using structured scaling matrices and Power-EP framework.
New method for estimating financial covariance matrices efficiently.
Homogeneous links were introduced by Peter Cromwell, who proved that the projection surface of these links, that given by the Seifert algorithm, has minimal genus. Here we provide a different proof, with a geometric rather than combinatorial flavor. To do this, we first show a direct relation between the Seifert matrix…
DKLM learns adaptive kernels for robust nonlinear subspace clustering.
A deep neural network is a hierarchical nonlinear model transforming input signals to output signals. Its input-output relation is considered to be stochastic, being described for a given input by a parameterized conditional probability distribution of outputs. The space of parameters consisting of weights and biases i…
New model handles complex non-linear relationships with hidden graph structures.
This paper tackles model selection for MoE models in high-dimensional data.
New diagonal knots found with non-torus structure.
In real-world applications, not all instances in multi-view data are fully represented. To deal with incomplete data, Incomplete Multi-view Learning (IML) rises. In this paper, we propose the Joint Embedding Learning and Low-Rank Approximation (JELLA) framework for IML. The JELLA framework approximates the incomplete d…
We introduce the notion of Haantjes algebra: It consists of an assignment of a family of operator fields on a differentiable manifold, each of them with vanishing Haantjes torsion. They are also required to satisfy suitable compatibility conditions. Haantjes algebras naturally generalize several known interesting geome…
Develops efficient quasi-Newton methods for training deep neural networks.
We define a second-order neural network stochastic gradient training algorithm whose block-diagonal structure effectively amounts to normalizing the unit activations. Investigating why this algorithm lacks in robustness then reveals two interesting insights. The first insight suggests a new way to scale the stepsizes, …
Sharp pseudospectral bounds prevent transient amplification in coupled gradient descent.
Study spectral flow on a warped cylinder with special boundary conditions.
It is shown that, in four dimensions, it is possible to introduce coordinates so that an analytic metric locally takes block diagonal form. i.e. one can find coordinates such that for where . We call a coordinate system in which the metric takes this for…
This paper develops dimension-agnostic inference methods for high-dimensional data.
Mini-Hes improves LFA model performance on HDI tasks with missing data.
Study local structure of knot group representations into SL(n,C).
New estimators reduce computation for Kendall's tau and conditional Kendall's tau matrices under structural assumptions.
We propose a fast general projection-free metric learning framework, where the minimization objective is a convex differentiable function of the metric matrix , and resides in the set of generalized graph Laplacian matrices for con…
A new metric learning framework for signed graphs using Gershgorin disc alignment.
The paper identifies redundant columns in matrices for feature selection and clustering.
We simplify matrix computations for block matrices, especially useful for covariance and correlation matrices.
Develops large-sample theory for non-stationary source separation.
Advanced optimization algorithms such as Newton method and AdaGrad benefit from second order derivative or second order statistics to achieve better descent directions and faster convergence rates. At their heart, such algorithms need to compute the inverse or inverse square root of a matrix whose size is quadratic of …
We show local rigidity of hyperbolic triangle groups generated by reflections in pairs of -dimensional subspaces of obtained by composition of the geometric representation in with the diagonal embeddings into and .
Algorithm solves robust linear regression with block Lewis weights.
Constructs six-dimensional braid group representations for knot detection.
We study the existence of invariant Einstein metrics on real flag manifolds associated to simple and non-compact split real forms of complex classical Lie algebras whose isotropy representation decomposes into two or three irreducible sub-representations. In this situation, one can have equivalent sub-modules, leading …
It is challenging to develop stochastic gradient based scalable inference for deep discrete latent variable models (LVMs), due to the difficulties in not only computing the gradients, but also adapting the step sizes to different latent factors and hidden layers. For the Poisson gamma belief network (PGBN), a recently …
Motivated by large-scale Collaborative-Filtering applications, we present a Non-Commuting Latent Factor (NCLF) tensor-completion approach for modeling three-way arrays, which is diagonal like the standard PARAFAC, but wherein different terms distinguish different kinds of three-way relations of co-clusters, as determin…
We present a method based on the orthogonal symmetric non-negative matrix tri-factorization of the normalized Laplacian matrix for community detection in complex networks. While the exact factorization of a given order may not exist and is NP hard to compute, we obtain an approximate factorization by solving an optimiz…