Paper analyzes asymmetry in LoRA initialization for foundation models.
problem Asymmetry in LoRA initialization affects generalization of foundation models.
method Theoretical analysis of asymmetric LoRA with frozen random factors.
result Upper bound on sample complexity of $ ilde{\mathcal{O}}\left(\frac{\sqrt{r}}{\sqrt{N}}
ight)$ with high probability.
The paper proves deep ReLU networks can die and proposes a new initialization method to prevent it.
problem Dying ReLU neurons in deep neural networks.
method The paper rigorously proves the dying ReLU problem and proposes a new randomized asymmetric initialization method.
result The new initialization method effectively prevents the dying ReLU problem.
Gradient descent recovers principal components of overparametrized asymmetric matrices without explicit regularization.
problem Asymmetric matrix factorization under overparametrization with minimal rank assumptions.
method Vanilla gradient descent with small random initialization and proper early stopping.
result Gradient descent produces the best low-rank approximation without explicit regularization.
Paper studies asymmetric matrix sensing, proving gradient descent converges to low-rank solutions.
problem Reconstructing asymmetric low-rank matrices from linear measurements.
method Factorized gradient descent with coupling and regularization properties.
result Gradient descent from small random initialization converges to globally optimal and generalizing solutions.
Randomly initialized transformers show extreme token preferences.
problem Structural biases in randomly initialized transformers.
method Dissection of transformer architecture at initialization.
result Initialization-induced biases persist throughout training.
A new asymmetric correntropy method improves robust adaptive filtering for asymmetric error distributions.
problem Inadequate handling of asymmetric error distributions in adaptive filtering.
method Proposes asymmetric correntropy using an asymmetric Gaussian kernel and develops a robust adaptive filtering algorithm.
result The proposed algorithm shows better steady-state convergence performance for asymmetric error distributions.
Study of classification in asymmetric quasi-metric spaces.
problem Classification in asymmetric quasi-metric spaces.
method Sample compression and nearest neighbor algorithm.
result Algorithm has favorable statistical properties.
Innovative extensions to option pricing models using asymmetric Brownian motion and random walk approaches.
problem Capturing empirical phenomena like return skewness, heavy tails, and volatility asymmetry in option pricing models.
method Developing the Geometric Asymmetric Brownian Motion (GABM) within the Bachelier--Black--Scholes--Merton framework.
result Deriving closed-form option pricing formulas and a discrete-time binomial tree algorithm that converges to the GABM limit.
The generalized correlation approach, which has been successfully used in statistical radio physics to describe non-Gaussian random processes, is proposed to describe stochastic financial processes. The generalized correlation approach has been used to describe a non-Gaussian random walk with independent, identically d…
Random walks on free groups reveal asymmetric expansion factors.
problem Understanding expansion factors in free groups.
method Random walks and BGIP on metric spaces.
result Generic outer automorphisms have different forward and backward expansion factors.
Study uncovers new phase transitions in asymmetric causal inference scenarios.
problem Understanding typical phase transitions in asymmetric causal inference.
method Combining Causal inference (C-inf) and Low-rank recovery (LRR) with Random duality - Free probability theory (RDT-FPT).
result Discovering a doubling low-rankness phenomenon in asymmetric scenarios.
Gradient descent solves asymmetric low-rank matrix factorization efficiently.
problem Optimizing asymmetric low-rank matrix factorization with non-convex and non-smoothness issues.
method Randomly initialized gradient descent with new symmetrization and perturbation techniques.
result Gradient descent converges to a global minimum of the asymmetric low-rank factorization problem.
Complex systems are typically represented by large ensembles of observations. Correlation matrices provide an efficient formal framework to extract information from such multivariate ensembles and identify in a quantifiable way patterns of activity that are reproducible with statistically significant frequency compared…
Gradient descent solves asymmetric low-rank matrix sensing without balancing.
problem Recovering asymmetric low-rank matrices from linear measurements.
method Gradient descent with spectral initialization, avoiding balancing term.
result Gradient descent converges linearly without balancing, factors stay balanced.
Paper shows robustness of gradient descent in matrix sensing despite perturbations.
problem Understanding robustness of gradient descent in matrix sensing.
method Developed perturbed gradient flow to capture noise and improve robustness.
result Gradient descent is robust to perturbations in matrix sensing.
New algorithms solve tensor problems with random components using SDP.
problem Exact tensor nuclear norm, decomposition, and completion for random tensors.
method Degree-4 Sum of Squares (SOS) semidefinite programs.
result Exact solutions for tensor nuclear norm, decomposition, and completion with random asymmetric components.
Investors with asymmetric information play a game to optimize their portfolios.
problem Two investors with different information levels compete in portfolio selection.
method Modelled as a Stackelberg game with entropy-regularized mean-variance objectives.
result Equilibria exist where follower's strategy depends on leader's actions.
The price impact for a single trade is estimated by the immediate response on an event time scale, i.e., the immediate change of midpoint prices before and after a trade. We work out the price impacts across a correlated financial market. We quantify the asymmetries of the distributions and of the market structures of …
In this paper, we provide local and global convergence guarantees for recovering CP (Candecomp/Parafac) tensor decomposition. The main step of the proposed algorithm is a simple alternating rank-1 update which is the alternating version of the tensor power iteration adapted for asymmetric tensors. Local convergence g…
AGD converges in polynomial iterations to optimal matrix factorization.
problem Matrix factorization optimization with alternating gradient descent.
method Alternating gradient descent with fixed step size, proving convergence in polynomial iterations.
result AGD reaches ε-optimal factorization in T iterations with high probability.
Tensor CANDECOMP/PARAFAC (CP) decomposition is an important tool that solves a wide class of machine learning problems. Existing popular approaches recover components one by one, not necessarily in the order of larger components first. Recently developed simultaneous power method obtains only a high probability recover…
New methods train neural networks without changing weights, achieving similar or higher performance.
problem Training neural networks efficiently with randomly initialized weights.
method Switching connections on and off, flipping weights' signs, minimizing changed connections.
result Achieves similar or higher performance with less computational cost than training all weights.
The paper analyzes how over-parameterization affects GD convergence in matrix sensing problems.
problem Matrix sensing problem with over-parameterized gradient descent.
method Analyzes symmetric and asymmetric parameterizations, provides lower bounds and convergence rates.
result Over-parameterization slows down GD convergence, but asymmetric parameterization can speed up convergence.
Study analyzes accuracy of tensor deflation in noisy conditions.
problem Analyzing accuracy of tensor deflation in noisy conditions.
method Asymptotic study of Hotelling-type tensor deflation in large tensor dimensions.
result Characterization of estimated singular values and singular vector alignments.
Gradient descent converges globally in deep linear residual networks with ZAS initialization.
problem Optimizing deep linear residual networks for convergence.
method Zero-asymmetric (ZAS) initialization for gradient descent.
result Gradient descent converges to an ε-optimal point in O(L^3 log(1/ε)) iterations.
Recently it was shown that the problem of Maximum Inner Product Search (MIPS) is efficient and it admits provably sub-linear hashing algorithms. Asymmetric transformations before hashing were the key in solving MIPS which was otherwise hard. In the prior work, the authors use asymmetric transformations which convert th…
New method finds rare dense clusters in asymmetric binary perceptrons, resolving algorithmic hardness.
problem Resolving algorithmic hardness in asymmetric binary perceptrons.
method Fully lifted random duality theory (fl RDT) and large deviation upgrade (sfl LD RDT).
result Local entropy breaks down for constraint densities in (0.77, 0.78) interval, matching current solver limits.
We consider an American contingent claim on a financial market where the buyer has additional information. Both agents (seller and buyer) observe the same prices, while the information available to them may differ due to some extra exogenous knowledge the buyer has. The buyer's information flow is modeled by an initial…
The aim of this paper is to relate Thurston's metric on Teichmüller space to several ideas initiated by T. Sorvali on isomorphisms between Fuchsian groups. In particular, this will give a new formula for Thurston's asymmetric metric for surfaces with punctures. We also update some results of Sorvali on boundary isomorp…
We examine random variables in the power law/regularly varying class with stochastic tail exponent, the exponent α having its own distribution. We show the effect of stochasticity of α on the expectation and higher moments of the random variable. For instance, the moments of a right-tailed or right-asymmetric varia…
Proposes a new neural head for asymmetric representation learning.
problem Asymmetric representation learning in directed relations.
method Role-aware neural convex divergence head.
result Role-aware projections improve directional accuracy over plain ICNN-Bregman heads.
We derive the exact form of the eigenvalue spectra of correlation matrices derived from a set of time-shifted, finite Brownian random walks (time-series). These matrices can be seen as random, real, asymmetric matrices with a special structure superimposed due to the time-shift. We demonstrate that the associated eigen…
DEUA detects diffusion-generated images by accounting for different types of uncertainty.
problem Detecting generated images with varying aleatoric and epistemic uncertainty.
method DEUA framework using Laplace approximation for DEU estimation and asymmetric loss function.
result DEUA achieves state-of-the-art performance on large-scale benchmarks.
A new UCB policy improves reward-cost ratio estimation in budgeted MAB.
problem Maximizing reward-cost ratio under budget constraints.
method ω-UCB policy with asymmetric confidence intervals.
result Logarithmic regret and superior performance in various settings.
We study surfaces evolving by mean curvature flow (MCF). For an open set of initial data that are C3-close to round, but without assuming rotational symmetry or positive mean curvature, we show that MCF solutions become singular in finite time by forming neckpinches, and we obtain detailed asymptotics of that singul…
AACC improves RL performance in changing environments.
problem Deterioration of RL performance in real-world tasks.
method Formalizes CMDPs, proposes AACC for contextual RL.
result AACC outperforms baselines in various simulated environments.
APGD algorithm reconstructs point set from partial distance measurements.
problem Reconstructing point set configuration from partial Euclidean distance measurements.
method Asymmetric Projected Gradient Descent (APGD) for EDMC problem.
result Global convergence and exact recovery with O(μ2r3κ2nlogn) observations. The performance of the Self-Organizing Map (SOM) algorithm is dependent on the initial weights of the map. The different initialization methods can broadly be classified into random and data analysis based initialization approach. In this paper, the performance of random initialization (RI) approach is compared to that…
Study of asymmetric rank-one tensor models with non-Gaussian noise.
problem Analyzing maximum-likelihood estimators for asymmetric rank-one tensor models.
method Spectrally separated branch analysis, resolvent methods, cumulant expansions, Efron-Stein-type variance bounds.
result Asymptotic singular value and mode-wise alignments are robust to non-Gaussian noise.
This paper tackles learning Stackelberg equilibrium in asymmetric games efficiently from noisy samples.
problem Learning Stackelberg equilibrium in asymmetric, general-sum games efficiently from noisy samples.
method The paper initiates the theoretical study of sample-efficient learning of the Stackelberg equilibrium in bandit feedback setting.
result Sharp positive results on sample-efficient learning of Stackelberg equilibrium with value optimal up to a fundamental gap identified.
Study dynamic equilibrium with insider and general uninformed agent preferences.
problem Analyzing asymmetric information and general utility functions in a continuous-time economy.
method Introducing a new method to prove existence of a partial communication equilibrium (PCE) for agents with general utility functions.
result Identify the equilibrium price in the small and large risk aversion limits for agents with power utility.
Study of geometric analysis on asymmetric metric spaces, including heat flow and Sobolev spaces.
problem Analysis of geometric properties on asymmetric metric measure spaces.
method Introduction of upper gradients, q-Laplacian, and q-heat flow in asymmetric settings. result Extension of concepts from symmetric to asymmetric metric measure spaces.
We consider a random walk on the mapping class group of a surface of finite type. We assume that the random walk is determined by a probability measure whose support is finite and generates a non-elementary subgroup H. We further assume that H is not consisting only of lifts with respect to any one covering. Then w…
GNNs with random node initialization are shown to be universally expressive.
problem Limitations of standard GNNs in distinguishing graphs.
method Random node initialization (RNI) to enhance GNNs' expressive power.
result GNNs with RNI are proven to be universally expressive.
New method shows random, diverse initializations are not essential for deep neural networks.
problem The necessity of random, diverse initializations in deep neural networks.
method Constructed a deep convolutional network with identical features by initializing weights to 0, enabling signal propagation and stable gradients.
result Random, diverse initializations are not necessary for training neural networks.
This paper examines the impact of random initialization in neural networks using NTK theory.
problem Understanding the impact of random initialization in neural networks using NTK theory.
method Analyzes the convergence of training dynamics and generalization error of wide neural networks with random initialization.
result The generalization error of wide neural networks trained by gradient descent is \( \Omega(n^{-\frac{3}{d+3}}) \), highlighting the benefits of mirror initialization and suggesting limitations of NTK theory.
We consider the occurrence of record-breaking events in random walks with asymmetric jump distributions. The statistics of records in symmetric random walks was previously analyzed by Majumdar and Ziff and is well understood. Unlike the case of symmetric jump distributions, in the asymmetric case the statistics of reco…
Bounds neural network output distribution to Gaussian for random initialization.
problem Quantifying the distribution of randomly initialized deep neural networks.
method Quantitative Gaussian approximation using quadratic Wasserstein distance.
result Explicit inequalities show how network sizes affect Gaussian behavior.