New algorithm for online optimization over symmetric cones, unifying previous methods.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The SCMU algorithm computes cone factorizations for symmetric cones, improving upon existing methods.
New algorithm for solving minimax problems over distributions converges to Nash equilibrium.
This short note reviews so-called Natural Gradient Descent (NGD) for multivariate Gaussians. The Fisher Information Matrix (FIM) is derived for several different parameterizations of Gaussians. Careful attention is paid to the symmetric nature of the covariance matrix when calculating derivatives. We show that there ar…
The present paper proposes generalized Gaussian kernel adaptive filtering, where the kernel parameters are adaptive and data-driven. The Gaussian kernel is parametrized by a center vector and a symmetric positive definite (SPD) precision matrix, which is regarded as a generalization of the scalar width parameter. These…
Symmetric CNNs improve sequential recommendation and protein structure prediction.
New algorithm provably converges to second-order stationary points in NMF.
This article discusses the existence problem of a compact quotient of a symmetric space by a properly discontinuous group with emphasis on the non-Riemannian case. Discontinuous groups are not always abundant in a homogeneous space if is non-compact. The first half of the article elucidates general machinery …
Efficiently approximates eigenspaces for symmetric and general matrices.
This work introduces a fixed-point optimization for variational inference.
Recurrent Neural Networks (RNNs) are designed to handle sequential data but suffer from vanishing or exploding gradients. Recent work on Unitary Recurrent Neural Networks (uRNNs) have been used to address this issue and in some cases, exceed the capabilities of Long Short-Term Memory networks (LSTMs). We propose a simp…
We propose a general model explanation system (MES) for "explaining" the output of black box classifiers. This paper describes extensions to Turner (2015), which is referred to frequently in the text. We use the motivating example of a classifier trained to detect fraud in a credit card transaction history. The key asp…
Olshausen and Field (OF) proposed that neural computations in the primary visual cortex (V1) can be partially modeled by sparse dictionary learning. By minimizing the regularized representation error they derived an online algorithm, which learns Gabor-filter receptive fields from a natural image ensemble in agreement …
The first part of this paper completes the classification of Whitney towers in the 4-ball that was started in three related papers. We provide an algebraic framework allowing the computations of the graded groups associated to geometric filtrations of classical link concordance by order n (twisted) Whitney towers in th…
New algorithm reduces complexity for SPD manifold optimization.
Efficient CD algorithms on matrix manifolds for optimization problems.
A new quasi-Newton method uses cubic regularization to avoid saddle points in deep learning.
Motivated by applications in neuroimaging analysis, we propose a new regression model, Sparse TensOr REsponse regression (STORE), with a tensor response and a vector predictor. STORE embeds two key sparse structures: element-wise sparsity and low-rankness. It can handle both a non-symmetric and a symmetric tensor respo…
Efficiently factorize tensors in streaming data with coreset selection.
Paper proposes a quasi-Newton method for nonlinear equations with global convergence guarantees.
Stochastic mirror descent (SMD) is a fairly new family of algorithms that has recently found a wide range of applications in optimization, machine learning, and control. It can be considered a generalization of the classical stochastic gradient algorithm (SGD), where instead of updating the weight vector along the nega…
We introduce the HSIC (Hilbert-Schmidt independence criterion) bottleneck for training deep neural networks. The HSIC bottleneck is an alternative to the conventional cross-entropy loss and backpropagation that has a number of distinct advantages. It mitigates exploding and vanishing gradients, resulting in the ability…
System allows users to critique explanations of recommendations.
Gradient clipping helps private SGD converge despite potential bias.
ITSPACE improves covariance alignment faster than other methods.
Newton's method tackles nonlinear mappings into vector bundles with connections and retractions.
Recently, there has been a trend to combine independent component analysis and canonical polyadic decomposition (ICA-CPD) for an enhanced robustness for the computation of CPD, and ICA-CPD could be further converted into CPD of a 5th-order partially symmetric tensor, by calculating the eigenmatrices of the 4th-order cu…
This is primarily a survey of the developments in the theory of harmonic maps of finite uniton number (or unitons) which have taken place since the introduction of extended solutions by Uhlenbeck. Such maps include all harmonic maps from the two-sphere to a compact Lie group or symmetric space. Extended solutions are e…
Paper proposes efficient weight updates for edge nodes with minimal communication.
A new update rule for deep reinforcement learning reduces learning variance and variance in reference signals.
We consider distributed optimization under communication constraints for training deep learning models. We propose a new algorithm, whose parameter updates rely on two forces: a regular gradient step, and a corrective direction dictated by the currently best-performing worker (leader). Our method differs from the param…
Paper proves Jeffrey's update rule minimizes relative entropy.
AMUSE uses reinforcement learning to predict optimal model updates.
We shed new insights on the two commonly used updates for the online -PCA problem, namely, Krasulina's and Oja's updates. We show that Krasulina's update corresponds to a projected gradient descent step on the Stiefel manifold of the orthonormal -frames, while Oja's update amounts to a gradient descent step using…
Efficiently updates posterior tree distributions over meta-trees.
SONIA optimizes machine learning problems with a novel algorithm.
Recently, the technique of local updates is a powerful tool in centralized settings to improve communication efficiency via periodical communication. For decentralized settings, it is still unclear how to efficiently combine local updates and decentralized communication. In this work, we propose an algorithm named as L…
The biological plausibility of the backpropagation algorithm has long been doubted by neuroscientists. Two major reasons are that neurons would need to send two different types of signal in the forward and backward phases, and that pairs of neurons would need to communicate through symmetric bidirectional connections. …
The paper examines how updates to probabilistic models influence behavior based on evidence.
Paper improves policy updates in reinforcement learning to speed up learning.
The paper develops a stationary-distribution theory for Random Forest ensemble size selection.
We analyze algorithms for approximating a function mapping to using deep linear neural networks, i.e. that learn a function parameterized by matrices and defined by . We focus on algorithms that learn through gradient descent on the population …
In this paper, we study the randomized distributed coordinate descent algorithm with quantized updates. In the literature, the iteration complexity of the randomized distributed coordinate descent algorithm has been characterized under the assumption that machines can exchange updates with an infinite precision. We con…
Proposes a new method for nonlinear Bayesian updates using ensemble kernel regression.
This paper analyzes how periodic and soft target updates stabilize linear Q-learning.
In this letter, we generalize the convolutional NMF by taking the -divergence as the contrast function and present the correct multiplicative updates for its factors in closed form. The new updates unify the -NMF and the convolutional NMF. We state why almost all of the existing updates are inexact and approximat…
EnKF's update is shown to be similar to Matheron's method in Gaussian process regression.
Jointly learns feature and sample relevancies for robust sparse recovery.