This paper solves matrix blind joint block diagonalization with noise.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper proposes ABDR for convex subspace clustering with adaptive block diagonal representation.
New insights into Hessian structure of neural networks reveal two forces.
New theory allows simultaneous block-diagonalization of commuting operator fields.
Randomized block-diagonal preconditioning improves parallel learning convergence.
Adaptive gradient approaches that automatically adjust the learning rate on a per-feature basis have been very popular for training deep networks. This rich class of algorithms includes Adagrad, RMSprop, Adam, and recent extensions. All these algorithms have adopted diagonal matrix adaptation, due to the prohibitive co…
Improved sparse Gaussian processes using structured scaling matrices and Power-EP framework.
Gaussian graphical models are widely utilized to infer and visualize networks of dependencies between continuous variables. However, inferring the graph is difficult when the sample size is small compared to the number of variables. To reduce the number of parameters to estimate in the model, we propose a non-asymptoti…
New model handles complex non-linear relationships with hidden graph structures.
Localized sketching improves matrix multiplication and ridge regression complexity.
We propose a modular extension of backpropagation for the computation of block-diagonal approximations to various curvature matrices of the training objective (in particular, the Hessian, generalized Gauss-Newton, and positive-curvature Hessian). The approach reduces the otherwise tedious manual derivation of these mat…
Homogeneous links were introduced by Peter Cromwell, who proved that the projection surface of these links, that given by the Seifert algorithm, has minimal genus. Here we provide a different proof, with a geometric rather than combinatorial flavor. To do this, we first show a direct relation between the Seifert matrix…
A Semi-Hidden Markov Model (SHMM) for bursty error channels is defined by a state transition probability matrix , a prior probability vector , and the state dependent output symbol error probability matrix . Several processes are utilized for estimating , and from a given empirically obtained or sim…
This paper tackles model selection for MoE models in high-dimensional data.
New method for estimating financial covariance matrices efficiently.
New estimators reduce computation for Kendall's tau and conditional Kendall's tau matrices under structural assumptions.
We define a second-order neural network stochastic gradient training algorithm whose block-diagonal structure effectively amounts to normalizing the unit activations. Investigating why this algorithm lacks in robustness then reveals two interesting insights. The first insight suggests a new way to scale the stepsizes, …
We introduce the notion of Haantjes algebra: It consists of an assignment of a family of operator fields on a differentiable manifold, each of them with vanishing Haantjes torsion. They are also required to satisfy suitable compatibility conditions. Haantjes algebras naturally generalize several known interesting geome…
Efficiently approximates Sparse PCA with significant speedups and minor error.
Develops efficient quasi-Newton methods for training deep neural networks.
Second-order methods for neural network optimization have several advantages over methods based on first-order gradient descent, including better scaling to large mini-batch sizes and fewer updates needed for convergence. But they are rarely applied to deep learning in practice because of high computational cost and th…
DKLM learns adaptive kernels for robust nonlinear subspace clustering.
A deep neural network is a hierarchical nonlinear model transforming input signals to output signals. Its input-output relation is considered to be stochastic, being described for a given input by a parameterized conditional probability distribution of outputs. The space of parameters consisting of weights and biases i…
Paper proposes efficient methods for clustering and signal recovery in high-dimensional data with block structures.
A fast metric learning framework using Gershgorin disc alignment.
In real-world applications, not all instances in multi-view data are fully represented. To deal with incomplete data, Incomplete Multi-view Learning (IML) rises. In this paper, we propose the Joint Embedding Learning and Low-Rank Approximation (JELLA) framework for IML. The JELLA framework approximates the incomplete d…
Sharp pseudospectral bounds prevent transient amplification in coupled gradient descent.
Subspace clustering is a useful technique for many computer vision applications in which the intrinsic dimension of high-dimensional data is often smaller than the ambient dimension. Spectral clustering, as one of the main approaches to subspace clustering, often takes on a sparse representation or a low-rank represent…
It is shown that, in four dimensions, it is possible to introduce coordinates so that an analytic metric locally takes block diagonal form. i.e. one can find coordinates such that for where . We call a coordinate system in which the metric takes this for…
Improved method for unbiased causal discovery in presence of unobserved confounding.
Using a Bayesian approach, we consider the problem of recovering sparse signals under additive sparse and dense noise. Typically, sparse noise models outliers, impulse bursts or data loss. To handle sparse noise, existing methods simultaneously estimate the sparse signal of interest and the sparse noise of no interest.…
Characterizes Anosov reducible representations in terms of eigenvalues.
A new metric learning framework for signed graphs using Gershgorin disc alignment.
The paper identifies redundant columns in matrices for feature selection and clustering.
BEGIN network models binary data without parametric assumptions.
Proposes TFCL to mitigate negative transfer in MTL by collaborating across features and tasks.
Deep latent-variable models learn representations of high-dimensional data in an unsupervised manner. A number of recent efforts have focused on learning representations that disentangle statistically independent axes of variation by introducing modifications to the standard objective function. These approaches general…
We introduce the Kronecker factored online Laplace approximation for overcoming catastrophic forgetting in neural networks. The method is grounded in a Bayesian online learning framework, where we recursively approximate the posterior after every task with a Gaussian, leading to a quadratic penalty on changes to the we…
Develops large-sample theory for non-stationary source separation.
Advanced optimization algorithms such as Newton method and AdaGrad benefit from second order derivative or second order statistics to achieve better descent directions and faster convergence rates. At their heart, such algorithms need to compute the inverse or inverse square root of a matrix whose size is quadratic of …
We obtain the natural diagonal almost product and locally product structures on the total space of the cotangent bundle of a Riemannian manifold. We find the Riemannian almost product (locally product) and the (almost) para-Hermitian cotangent bundles of natural diagonal lift type. We prove the characterization theorem…
Algorithm solves robust linear regression with block Lewis weights.
The paper introduces a penalized matrix estimation procedure aiming at solutions which are sparse and low-rank at the same time. Such structures arise in the context of social networks or protein interactions where underlying graphs have adjacency matrices which are block-diagonal in the appropriate basis. We introduce…
New method estimates sparse covariance matrices in logit mixtures.
In this paper, we study deep diagonal circulant neural networks, that is deep neural networks in which weight matrices are the product of diagonal and circulant ones. Besides making a theoretical analysis of their expressivity, we introduced principled techniques for training these models: we devise an initialization s…
We present a method based on the orthogonal symmetric non-negative matrix tri-factorization of the normalized Laplacian matrix for community detection in complex networks. While the exact factorization of a given order may not exist and is NP hard to compute, we obtain an approximate factorization by solving an optimiz…
New diagonal knots found with non-torus structure.
Ensembles of neural networks improve training dynamics and performance.