Gradient descent recovers principal components of overparametrized asymmetric matrices without explicit regularization.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper analyzes asymmetry in LoRA initialization for foundation models.
We initiate the rigorous study of classification in quasi-metric spaces. These are point sets endowed with a distance function that is non-negative and also satisfies the triangle inequality, but is asymmetric. We develop and refine a learning algorithm for quasi-metrics based on sample compression and nearest neighbor…
Gradient descent solves asymmetric low-rank matrix factorization efficiently.
Gradient descent solves asymmetric low-rank matrix sensing without balancing.
Tensor CANDECOMP/PARAFAC (CP) decomposition is an important tool that solves a wide class of machine learning problems. Existing popular approaches recover components one by one, not necessarily in the order of larger components first. Recently developed simultaneous power method obtains only a high probability recover…
The paper analyzes how over-parameterization affects GD convergence in matrix sensing problems.
We analyze the global convergence of gradient descent for deep linear residual networks by proposing a new initialization: zero-asymmetric (ZAS) initialization. It is motivated by avoiding stable manifolds of saddle points. We prove that under the ZAS initialization, for an arbitrary target matrix, gradient descent con…
Paper studies asymmetric matrix sensing, proving gradient descent converges to low-rank solutions.
The dying ReLU refers to the problem when ReLU neurons become inactive and only output 0 for any input. There are many empirical and heuristic explanations of why ReLU neurons die. However, little is known about its theoretical analysis. In this paper, we rigorously prove that a deep ReLU network will eventually die in…
We consider an American contingent claim on a financial market where the buyer has additional information. Both agents (seller and buyer) observe the same prices, while the information available to them may differ due to some extra exogenous knowledge the buyer has. The buyer's information flow is modeled by an initial…
The aim of this paper is to relate Thurston's metric on Teichmüller space to several ideas initiated by T. Sorvali on isomorphisms between Fuchsian groups. In particular, this will give a new formula for Thurston's asymmetric metric for surfaces with punctures. We also update some results of Sorvali on boundary isomorp…
We study surfaces evolving by mean curvature flow (MCF). For an open set of initial data that are -close to round, but without assuming rotational symmetry or positive mean curvature, we show that MCF solutions become singular in finite time by forming neckpinches, and we obtain detailed asymptotics of that singul…
This paper tackles learning Stackelberg equilibrium in asymmetric games efficiently from noisy samples.
Study of geometric analysis on asymmetric metric spaces, including heat flow and Sobolev spaces.
Study dynamic equilibrium with insider and general uninformed agent preferences.
New metrics for Anosov representations defined from Thurston's asymmetric metrics.
Generalizes Thurston's asymmetric metric to flat metrics.
This study examines asymmetric cross-correlations in cryptocurrency markets using fractal analysis.
Theoretical justification for asymmetric actor-critic algorithms in reinforcement learning.
This work presents deep asymmetric networks with a set of node-wise variant activation functions. The nodes' sensitivities are affected by activation function selections such that the nodes with smaller indices become increasingly more sensitive. As a result, features learned by the nodes are sorted by the node indices…
We consider the problem of designing locality sensitive hashes (LSH) for inner product similarity, and of the power of asymmetric hashes in this context. Shrivastava and Li argue that there is no symmetric LSH for the problem and propose an asymmetric LSH based on different mappings for query and database points. Howev…
New asymmetric kernel methods improve feature learning.
The article confirms two quasi-alternating surgeries for 9 asymmetric L-space knots.
Asymmetric expansion preserves convexity in hyperbolic geometry.
The paper improves asymmetric causality tests by addressing inefficiencies and statistical significance issues.
We propose Deep Asymmetric Multitask Feature Learning (Deep-AMTFL) which can learn deep representations shared across multiple tasks while effectively preventing negative transfer that may happen in the feature sharing process. Specifically, we introduce an asymmetric autoencoder term that allows reliable predictors fo…
In this paper we show how the study of asymmetric R&D alliances, that are those between young and small firms and large and MNEs firms for knowledge exploration and/or exploitation, requires the adoption of a coopetitive framework which consider both collaboration and competition. We draw upon the literature on asymmet…
Randomly initialized transformers show extreme token preferences.
Bayesian VI copula models capture asymmetric intraday equity dependence.
Gradient descent training of neural networks leads to solutions close to natural cubic splines.
This work describes compactifications of metric spaces and vector spaces using asymmetric norms.
Extends multidimensional scaling to analyze three-way asymmetric proximities.
Extends metric to Margulis spacetimes for convex properties.
In this paper, we provide local and global convergence guarantees for recovering CP (Candecomp/Parafac) tensor decomposition. The main step of the proposed algorithm is a simple alternating rank- update which is the alternating version of the tensor power iteration adapted for asymmetric tensors. Local convergence g…
In recent years, correntropy has been seccessfully applied to robust adaptive filtering to eliminate adverse effects of impulsive noises or outliers. Correntropy is generally defined as the expectation of a Gaussian kernel between two random variables. This definition is reasonable when the error between the two random…
This paper introduces constrained mixtures for continuous distributions, characterized by a mixture of distributions where each distribution has a shape similar to the base distribution and disjoint domains. This new concept is used to create generalized asymmetric versions of the Laplace and normal distributions, whic…
Enhances reinforcement learning with partial state information.
Recent theoretical work has demonstrated that deep neural networks have superior performance over shallow networks, but their training is more difficult, e.g., they suffer from the vanishing gradient problem. This problem can be typically resolved by the rectified linear unit (ReLU) activation. However, here we show th…
New methods train neural networks without changing weights, achieving similar or higher performance.
This paper studies gradient flows in asymmetric metric spaces and proves existence results.
Paper defines saddle points in asymmetric Dynkin games using martingale theory.
Maps and measures on surfaces link best Lipschitz and least gradient functions.
Study uncovers new phase transitions in asymmetric causal inference scenarios.
Modified asymmetric hidden Markov models for time series with autoregressive components.
Mixtures of multivariate contaminated shifted asymmetric Laplace distributions are developed for handling asymmetric clusters in the presence of outliers (also referred to as bad points herein). In addition to the parameters of the related non-contaminated mixture, for each (asymmetric) cluster, our model has one param…
In this paper, we propose a novel asymmetric -insensitive pinball loss function for quantile estimation. There exists some pinball loss functions which attempt to incorporate the -insensitive zone approach in it but, they fail to extend the -insensitive approach for quantile estimation in true sense. The propo…
Study proves value of non-Markovian games with partial, asymmetric info.