A well-known issue of Batch Normalization is its significantly reduced effectiveness in the case of small mini-batch sizes. When a mini-batch contains few examples, the statistics upon which the normalization is defined cannot be reliably estimated from it during a training iteration. To address this problem, we presen…
This paper resolves BIHT convergence, showing normalization is not necessary in noiseless settings but crucial for robustness.
problem Analyzing convergence and robustness of BIHT for 1-bit compressed sensing.
method Characterizes BIHT convergence and robustness, proving necessity of normalization for robustness under sign corruptions.
result Per-iteration normalization is not necessary for optimal recovery in noiseless settings but is crucial for robustness under sign corruptions.
Formal normal form created for real-smooth hypersurfaces.
problem Real-smooth hypersurfaces in complex spaces.
method Iterative normalization procedure.
result Formal normal form constructed for a large class of hypersurfaces.
Study symplectification of rank 2 distributions and their connections.
problem Understanding symplectification and Cartan prolongations of rank 2 distributions.
method Using Tanaka-Morimoto theory and symplectification procedure for rank 2 distributions.
result Demonstrates the existence of normal Cartan connections and iterated prolongations for rank 2 distributions.
Proves a theorem similar to Moser's using a normalization method.
problem Proving a theorem similar to Moser's in a specific context.
method Iterative normalization procedure based on Generalized Fischer Decompositions.
result An analogue of the Theorem of Moser proven.
Critical volatility triggers log-normal to power-law transitions in interconnected systems.
problem Understanding the transition from log-normal to power-law distributions in interconnected systems.
method Analyzing an infinite option-on-option chain model, deriving a critical volatility threshold.
result A critical volatility threshold of approximately 250.66% for unconditional cases, dropping to 125.3% with selective survival.
Paper analyzes normal approximation for two-timescale stochastic algorithms, revealing interaction between fast and slow timescales.
problem Non-asymptotic bounds for accuracy of normal approximation in linear two-timescale stochastic approximation algorithms.
method Established bounds for normal approximation in terms of convex distance, focusing on last iterate and Polyak-Ruppert averaging.
result Normal approximation rate for the last iterate improves with increased timescale separation, while it decreases in the averaged setting.
SINF models transform arbitrary PDFs to target PDFs using 1D slices.
problem Transforming arbitrary probability distributions to target distributions efficiently.
method Iterative Optimal Transport of 1D slices, maximizing Wasserstein distance.
result SINF models generate high-quality samples and competitive density estimates.
Paper proves convergence of bi-stochastically normalized graph Laplacian to manifold Laplacian and robustness to outlier noise.
problem Convergence of bi-stochastically normalized graph Laplacian to manifold Laplacian and robustness to outlier noise.
method Proves convergence of bi-stochastically normalized graph Laplacian to manifold Laplacian with rates, and proposes an approximate and constrained matrix scaling problem to achieve the same consistency rate.
result Graph Laplacian consistency rate matches the rate for clean manifold data plus an additional term proportional to the boundedness of the inner-products of the noise vectors.
SGD converges to critical points of normalized margin in late-stage training for homogeneous neural networks.
problem Analyzing the implicit bias of SGD on homogeneous neural networks.
method Interpreting SGD dynamics as an Euler-like discretization of a conservative field flow associated with the normalized classification margin.
result Normalized SGD iterates converge to the set of critical points of the normalized margin at late-stage training.
This paper provides a framework to analyze stochastic gradient algorithms in a mean squared error (MSE) sense using the asymptotic normality result of the stochastic gradient descent (SGD) iterates. We perform this analysis by taking the asymptotic normality result and applying it to the finite iteration case. Specific…
Stochastic algo learns from evolving data, achieving optimal performance.
problem Performative prediction and multiplayer extensions.
method Stochastic approximation with decision-dependent distributions.
result Asymptotic normality and optimality of the algorithm's performance.
We propose to improve trust region policy search with normalizing flows policy. We illustrate that when the trust region is constructed by KL divergence constraints, normalizing flows policy generates samples far from the 'center' of the previous policy iterate, which potentially enables better exploration and helps av…
BR-SNIS reduces bias in self-normalized IS without increasing variance.
problem Bias in self-normalized IS.
method Iterated sampling-importance resampling (ISIR) to form a bias-reduced estimator.
result Significant reduction in bias without increasing variance.
New Convolutional Unit improves Batch Whitening performance.
problem Improving the efficiency and effectiveness of Batch Whitening.
method Proposes a new Convolutional Unit that aligns with Batch Whitening theory and empirically analyzes the original Convolutional Unit.
result Significantly improved performance on multiple image classification datasets.
Study wSAA for contextual decisions, improving uncertainty quantification under computational constraints.
problem Uncertainty quantification limitations in wSAA for contextual stochastic optimization.
method Establish central limit theorems and asymptotic-normality-based confidence intervals for optimal costs.
result Over-optimizing can mitigate misspecification and preserve asymptotic normality, albeit at a slower convergence rate.
New method clusters matrix-variate data with outliers.
problem Clustering matrix-variate data with outliers.
method Iterative approach using subset log-likelihoods.
result Extends OCLUST algorithm to matrix-variate normal data.
Gaussianization flows transform any random vector into a Gaussian, enabling efficient computation and sample generation.
problem Transforming any random vector into a Gaussian for efficient computation and sample generation.
method Iterative Gaussianization and normalizing flow model.
result Gaussianization flows are universal approximators and achieve better performance on tabular datasets.
The Sinkhorn-Knopp algorithm converges quickly but the number of iterations is poorly understood.
problem Understanding the number of iterations required for the Sinkhorn-Knopp algorithm to converge.
method Analyzing the Sinkhorn-Knopp algorithm for matrices with a specific density threshold.
result The Sinkhorn-Knopp algorithm requires Ω(n1/2/ε) iterations for matrices with density γ<1/2. In every point of a Kähler manifold there exist special holomorphic coordinates well adapted to the underlying geometry. Comparing these Kähler normal coordinates with the Riemannian normal coordinates defined via the exponential map we prove that their difference is a universal power series in the curvature tensor and…
Weight normalization speeds up matrix sensing problems.
problem Matrix sensing with overparameterization.
method Generalized weight normalization with Riemannian optimization.
result WN achieves linear convergence, improving speed and complexity.
Inspired by recent advances in deep learning, we propose a novel iterative BP-CNN architecture for channel decoding under correlated noise. This architecture concatenates a trained convolutional neural network (CNN) with a standard belief-propagation (BP) decoder. The standard BP decoder is used to estimate the coded b…
Article studies symmetry in smooth vector bundles using advanced operations.
problem Symmetry phenomena in smooth vector bundles after two iterations of the normal functor.
method Developed theory of pullback and quotient for double vector bundles and morphisms, focusing on naturality of the normal functor.
result Expected symmetry is obtained through universal behavior and compatibility of operations.
We consider free and proper cotangent-lifted symmetries of Hamiltonian systems. For the special case of G = SO(3), we construct symplectic slice coordinates around an arbitrary point. We thus obtain a parametrisation of the phase space suitable for the study of dynamics near relative equilibria, in particular for the B…
In a recent paper, Darvas-Rubinstein proved a convergence result for the Kahler-Ricci iteration, which is a sequence of recursively defined complex Monge-Ampere equations. We introduce the Monge-Ampere iteration to be an analogous, but more general, sequence of recursively defined real Monge-Ampere second boundary valu…
This paper approximates SA iterates using Gaussian distributions for tail bounds.
problem Characterizing the distribution of stochastic approximation iterates in finite time.
method Approximating pre-limit distributions of SA iterates by Gaussian sequences with recursively defined covariances.
result Explicit bounds on the Wasserstein-1 distance between rescaled iterates and Gaussians.
RS-NSGD improves SGD convergence for heavy-tailed noise.
problem Nonconvex optimization with heavy-tailed noise.
method Integrates direction normalization into subspace updates.
result Achieves better oracle complexity than full-dimensional normalized SGD.
Paper proposes CoopFlow, a two-flow generator for energy-based models.
problem Training energy-based models with Langevin flow and normalizing flow.
method CoopFlow trains an energy-based model using a normalizing flow initialization and a short-run Langevin flow revision.
result CoopFlow converges to a moment matching estimator and synthesizes realistic images.
EWFM trains continuous flows with only energy evaluations, improving sample quality with fewer computations.
problem Efficiently sampling from complex, high-dimensional Boltzmann distributions using only energy evaluations.
method Energy-Weighted Flow Matching (EWFM) using importance sampling and iterative/annealed training.
result Improved sample quality with up to 3 orders of magnitude fewer energy evaluations compared to existing methods.
This paper introduces a novel theoretically sound approach for the celebrated CMA-ES algorithm. Assuming the parameters of the multi variate normal distribution for the minimum follow a conjugate prior distribution, we derive their optimal update at each iteration step. Not only provides this Bayesian framework a justi…
We describe the first term of the Λk−1C--spectral sequence (see math.DG/0610917) of the diffiety (E,C), E being the infinite prolongation of an l-normal system of partial differential equations, and C the Cartan distribution on it.
iEFM trains CNF models from unnormalized densities efficiently.
problem Training generators from energy functions or unnormalized densities.
method Iterated energy-based flow matching (iEFM) with simulation-free objective.
result iEFM outperforms existing methods in probabilistic modeling.
The paper improves matrix completion with auxiliary covariates using LS estimation.
problem Matrix completion with noisy data and auxiliary covariates.
method Iterative least squares estimation with statistical properties derived.
result Asymptotic normal distributions of estimators for low-rank matrix and coefficient matrix.
Two new rational formulae for normal implied volatility are presented.
problem Calculating normal implied volatility using iterative methods.
method Two explicit rational formulae that avoid iteration and logarithms.
result Accurate and fast formulae for normal implied volatility.
The paper extends NUP representations to factor graphs for better estimation.
problem Nontrivial model-based estimation problems.
method Augmenting factor graphs with convex-dual variables and NUP representations; proposing a new iterative algorithm.
result A new dual algorithm for state space problems.
New variational flows improve Monte Carlo and normalization tasks.
problem Intractable global optimum in expressive variational families.
method Constructing asymptotically exact variational flows from involutive MCMC kernels.
result Provable total variation convergence of new variational families.
Computer vision SSL methods show effectiveness on time series data.
problem Evaluate if computer vision SSL frameworks are effective on time series data.
method Evaluated on UCR and UEA archives, proposed a new method improving VICReg.
result Computer vision SSL frameworks can be effective on time series data.
The mean field variational Bayes method is becoming increasingly popular in statistics and machine learning. Its iterative Coordinate Ascent Variational Inference algorithm has been widely applied to large scale Bayesian inference. See Blei et al. (2017) for a recent comprehensive review. Despite the popularity of the …
Based on uniform CR Sobolev inequality and Moser iteration, this paper investigates the convergence of closed pseudo-Hermitian manifolds. In terms of the subelliptic inequality, the set of closed normalized pseudo-Einstein manifolds with some uniform geometric conditions is compact. Moreover, the set of closed normaliz…
Fast algorithm for rescaling vectors with clipping, improving training efficiency.
problem Efficiently rescale vectors to a desired length while maintaining them within a domain after clipping.
method Analytical solution for optimal rescaling using fast and differentiable algorithm.
result Optimal rescaling can be found analytically, improving training efficiency for neural networks.
We prove the existence and uniqueness of Kähler-Einstein metrics on Q-Fano varieties with log terminal singularities (and more generally on log Fano pairs) whose Mabuchi functional is proper. We study analogues of the works of Perelman on the convergence of the normalized Kähler-Ricci flow, and of Keller, Rubinstein on…
Estimates CDF over complex regions using normalizing flows.
problem Challenges in estimating CDF over complex regions using traditional methods.
method Leverages diffeomorphic properties of normalizing flows and divergence theorem.
result Improves sample efficiency in estimating CDF over complex regions.
Online learning to rank is a core problem in machine learning. In Lattimore et al. (2018), a novel online learning algorithm was proposed based on topological sorting. In the paper they provided a set of self-normalized inequalities (a) in the algorithm as a criterion in iterations and (b) to provide an upper bound for…
GD iterates for non-homogeneous deep nets increase margin and converge in direction.
problem Understanding implicit bias in non-homogeneous deep networks.
method Characterization of GD iterates' properties starting from small empirical risk.
result GD iterates converge in direction despite diverging norms, satisfying KKT conditions.
Paper develops a robust PP distributed quasi-Newton estimation for Byzantine machines.
problem Byzantine machines in distributed computing under Privacy Protection constraints.
method Robust PP distributed quasi-Newton estimation method that transmits only five vectors.
result Reduces privacy budgeting and transmission cost compared to gradient descent and Newton iteration.
We introduce a new routing algorithm for capsule networks, in which a child capsule is routed to a parent based only on agreement between the parent's state and the child's vote. The new mechanism 1) designs routing via inverted dot-product attention; 2) imposes Layer Normalization as normalization; and 3) replaces seq…
Study shows long-term flow on special manifolds with positive Yamabe constant.
problem Analyzing long-time behavior of Yamabe flow on singular spaces.
method Formulated axioms for long-time existence, used parabolic Moser iteration for bounds.
result Established long-time existence of normalized Yamabe flow on specified manifolds.
We provide an improved analysis of normalized SGD showing that adding momentum provably removes the need for large batch sizes on non-convex objectives. Then, we consider the case of objectives with bounded second derivative and show that in this case a small tweak to the momentum formula allows normalized SGD with mom…