Algorithm improves stochastic gradient optimization with normalized steps.
problem Stochastic and finite sum minimization problems.
method Trust region algorithm with normalized steps.
result Algorithm converges similarly to traditional stochastic gradient under certain conditions.
In this paper we study the main geometric properties of the Carnot-Carathéodory (abbreviated CC) distance $\dc$ in the setting of k k k -step sub-Riemannian Carnot groups from many different points of view. An extensive study of the so-called normal CC-geodesics is given. We state and prove some related variational formul…
Gradually Truncated Log-normal distribution - Size distribution of firms Abstract Many natural and economical phenomena are described through power law or log- normal distributions. In these cases, probability decreases very slowly with step size compared to normal distribution. Thus it is essential to cut-off these di…
Proposes adversarial normalization for multi-domain image segmentation.
problem Current image normalization is per-dataset, limiting multi-domain segmentation.
method Adversarial training to learn common normalizing functions across multiple datasets.
result Optimal normalizer improves segmentation accuracy and realism.
Paper classifies heart sound recordings as normal or abnormal.
problem Classifying normal/abnormal heart sound recordings.
method Four steps: preprocessing, feature extraction, training, validation. Back propagation neural network used.
result Optimal threshold determined for distinguishing normal and abnormal.
Proposes a neural network model for detecting collective anomalies in network security.
problem Traditional anomaly detection struggles with new, unknown intrusion types.
method Trains a Long Short-Term Memory Recurrent Neural Network (LSTM RNN) on normal data to predict anomalies and uses prediction errors over time to detect collective anomalies.
result The proposed model efficiently detects collective anomalies in network security.
MuonEq improves training of matrix-valued parameters by rebalancing momentum before orthogonalization.
problem Training matrix-valued parameters with orthogonalized-update optimizers like Muon.
method MuonEq introduces three lightweight pre-orthogonalization equilibration schemes: two-sided row/column normalization (RC), row normalization (R), and column normalization (C).
result Row/column normalization acts as a zeroth-order surrogate for whitening and improves the geometry seen by orthogonalization.
Study improves BN TTA under distribution shift using higher-order asymptotics.
problem Improving BN TTA for changing data distributions.
method Integrates Edgeworth expansion and saddlepoint approximation with one-step M-estimation.
result Derives optimal weighting parameter for minimized mean-squared error.
Adaptive step-size improves optimization in complex geometries.
problem Optimizing functions with non-Euclidean geometries.
method Adaptive step-size strategy for optimization algorithms.
result Guaranteed convergence for Adaptive Conditional Gradient Descent.
New method normalizes matrix features for robust low-rank approximation.
problem Robust feature normalization for low-rank matrix approximation.
method Learn quantile normalization operators jointly with matrix factorization.
result Improves quality of low-rank representation of data.
A deep learning method for probabilistic weather forecasting.
problem Probabilistic forecasting of weather.
method Two chained machine-learning steps: dimension reduction and density estimation using normalizing flows.
result The method produces accurate conditional forecast distributions for weather.
Model predicts stock price changes and forecasts using tokenized data.
problem Challenges in stock price forecasting and prediction due to dynamic data and statistical differences.
method Introduces PCIE model with tokenization to handle both forecasting and prediction.
result PCIE model outperforms state-of-the-art models in forecast and prediction tasks.
Proposes a new normalization method for deep neural networks in financial forecasting.
problem Deep neural networks are sensitive to input variable range and prone to numerical issues, especially with financial time-series.
method Bilinear input normalization method that handles high-frequency financial time-series without expert knowledge.
result Significant improvements in forecasting future stock price dynamics over other normalization techniques.
MR estimator simplifies causal inference by combining models without hyperparameter tuning.
problem Difficulty in choosing optimal hyperparameters for neural network models in causal inference.
method Multiply Robust (MR) estimator that combines multiple first-step models.
result MR estimator is n r n^r n r consistent and asymptotically normal under certain conditions. Study properties of sets with constant normal in Carnot groups.
problem Properties of sets with constant normal in Carnot groups.
method Analysis of subsets with intrinsic constant normal, proving regularity and structural results.
result Every constant-normal set in Carnot groups of step 4 or less is intrinsically rectifiable.
Batch normalization improves deep learning by enabling larger learning rates.
problem Improving accuracy and speeding up training in deep neural networks.
method Empirical experiments and analysis of gradient and activation behavior.
result Batch normalization primarily enables training with larger learning rates, leading to faster convergence and better generalization.
RC reduces neural network redundancy and improves performance through independent BN layers.
problem Improving neural network performance and reducing redundancy.
method Recurrent convolution with independent batch normalization layers for different unrolling steps.
result The proposed method improves RC networks' performance and achieves cost-adjustable inference.
SNF combines stochastic and deterministic steps to sample complex distributions.
problem Sampling complex probability distributions efficiently.
method Stochastic Normalizing Flows (SNF) - sequence of invertible functions and stochastic blocks.
result SNFs improve efficiency and representational power over pure MCMC/LD.
Proposes nAIPW for robust ATE estimation using neural networks.
problem Estimation of ATE with potential confounders and nonlinear relationships.
method Normalized AIPW (nAIPW) with neural networks and regularization.
result nAIPW maintains double-robustness and orthogonality properties.
Channel normalization prevents vanishing gradients in convolutional neural networks.
problem Vanishing gradients in convolutional neural networks during optimization.
method Channel normalization, which centers and normalizes each channel individually.
result Channel normalization avoids vanishing gradients, enabling efficient optimization.
In Carnot groups of step 3, all subriemannian geodesics are proved to be normal. The proof is based on a reduction argument and the Goh condition for minimality of singular curves. The Goh condition is deduced from a reformulation and a calculus of the end-point mapping which boils down to the graded structures of Carn…
Fine-tuning normalization layers can reconstruct smaller networks.
problem Understanding the expressive power of fine-tuning normalization layers.
method Random ReLU networks and sparsified networks were fine-tuned to reconstruct target networks.
result Fine-tuning normalization layers can reconstruct networks that are O ( e x t w i d t h ) O(\sqrt{ ext{width}}) O ( e x t w i d t h ) times smaller. Improves Graph Convolutional Network performance on citation datasets.
problem Improving Graph Convolutional Network performance on citation datasets.
method Exploring graph regularization and alternative graph convolution approaches.
result Explicit graph regularization was incorrectly rejected by Kipf & Welling (2016).
Paper improves CLT and bootstrap approximations for LSA with decreasing step size.
problem Improving normal approximation and bootstrap methods for LSA with decreasing step sizes.
method Refined Berry-Esseen bounds and multiplier bootstrap procedure for LSA.
result Approximation rates up to 1 / n 1/\sqrt{n} 1/ n for LSA rescaled error distribution. We propose a novel time discretization for the log-normal SABR model and derive its asymptotic properties.
problem Analyzing the log-normal SABR model's time-discretized behavior and implied volatility surface.
method We use the Euler-Maruyama scheme for time discretization and derive asymptotic properties in the limit of large number of time steps.
result We derive an exact representation of the implied volatility surface for arbitrary maturity and strike in the asymptotic regime.
Curves in Carnot groups avoid compact sets, growing at least t 1 / s t^{1/s} t 1/ s .
problem Existence of periodic normal geodesics in subFinsler Carnot groups.
method Analysis of curves satisfying Pontryagin Maximum Principle.
result Normal curves in subFinsler Carnot groups leave every compact set.
Study of curvature flow on complex Lie groups, leading to soliton convergence.
problem Characterizing long-time behavior of curvature flow on complex 2-step nilpotent Lie groups.
method Analyzing left-invariant metrics and using Cheeger-Gromov topology.
result Normalized solutions converge to a non-flat algebraic soliton.
The estimation of normalizing constants is a fundamental step in probabilistic model comparison. Sequential Monte Carlo methods may be used for this task and have the advantage of being inherently parallelizable. However, the standard choice of using a fixed number of particles at each iteration is suboptimal because s…
SGD's uncertainty quantified in non-convex learning problems.
problem Uncertainty quantification in non-convex learning problems.
method Asymptotic normality of SGD iterates and bias characterization.
result SGD iterates are asymptotically normally distributed around the expected value of the invariant distribution.
Batch-normalized RHN improves gradient control in recurrent networks.
problem Gradient vanishing or exploding in recurrent networks.
method Batch normalization applied at each recurrence loop in RHN.
result Batch-normalized RHN converges faster and performs better.
The paper analyzes the randomized midpoint method for Langevin diffusions, revealing biases and asymptotic properties.
problem Analyzing biases and asymptotic properties of the randomized midpoint method for Langevin diffusions.
method Characterization of stationary distribution and asymptotic normality for numerical integration.
result The step-size needs to go to zero for the method to be asymptotically unbiased.
Proposes a new algorithm for solving optimization problems with stochastic objectives and equality constraints.
problem Optimization problems with stochastic objectives and deterministic equality constraints.
method Trust-region stochastic sequential quadratic programming (TR-StoSQP) with adaptive relaxation techniques.
result Established a global almost sure convergence guarantee for TR-StoSQP.
Improves neural network performance by normalizing activation functions.
problem Improving convergence speed and robustness of neural networks.
method Transforming existing activation functions into ones with better properties.
result Significantly promotes convergence robustness, maximum training depth, and anytime performance.
Analyzes factors affecting flow VI performance.
problem Consistent performance of flow VI across studies.
method Step-by-step analysis of capacity, objectives, batchsize, estimators, and step-sizes.
result Specific recommendations and a flow VI recipe.
Study exact formula and geodesics for Carnot-Carathéodory distance on 2-step groups.
problem Exact formula and geodesics for Carnot-Carathéodory distance on 2-step groups.
method Combining Varadhan's formula, Loewner's theorem, and the method of stationary phase.
result Characterization of squared sub-Riemannian distance and cut locus on generalized Heisenberg-type groups and star graphs.
Training state-of-the-art, deep neural networks is computationally expensive. One way to reduce the training time is to normalize the activities of the neurons. A recently introduced technique called batch normalization uses the distribution of the summed input to a neuron over a mini-batch of training cases to compute…
Bayesian approach improves CMA-ES algorithm for faster convergence.
problem Improving the CMA-ES algorithm for faster convergence.
method Deriving optimal updates for CMA-ES parameters using conjugate priors.
result New versions of CMA-ES converge faster with normal-Wishart or normal-Inverse Wishart priors.
We show that in any triangulated 3-manifold, every index n topologically minimal surface can be transformed to a surface which has local indices (as computed in each tetrahedron) that sum to at most n. This generalizes classical theorems of Kneser and Haken, and more recent theorems of Rubinstein and Stocking, and is t…
Proves finitely generated graded rings for klt singularities.
problem Understanding the structure of klt singularities.
method Analyzes graded rings associated with minimizers of normalized volume functions.
result Graded rings are finitely generated for klt singularities.
We show that strictly abnormal geodesics arise in graded nilpotent Lie groups. We construct such a group, for which some Carnot geodesics are strictly abnormal; in fact, they are not normal in any subgroup. In the step-2 case we also prove that these geodesics are always smooth. Our main technique is based on the equat…
Proposes method for eliciting non-parametric joint priors using normalizing flows.
problem Learning complex non-parametric joint priors for model parameters.
method Expert elicitation combined with normalizing flows for generative modeling.
result Framework supports elicitation of both parametric and non-parametric priors.
Yau's Affine Normal Descent optimizes smooth unconstrained problems with geometrically adapted directions.
problem Optimizing smooth unconstrained problems with geometrically adapted directions.
method Yau's Affine Normal Descent (YAND) uses the equi-affine normal of level-set hypersurfaces as search directions.
result YAND converges globally under standard smoothness assumptions and locally quadratically near nondegenerate minimizers.
Unified framework for data-free sampling using Wasserstein gradient flows.
problem Efficient sampling from unnormalized distributions without data.
method Unified theoretical framework based on Wasserstein gradient flows.
result Unified form of velocity field for various f-divergences.
Unified approach to geometric structure equivalence problem.
problem Equivalence problem of geometric structures.
method Unified framework, step prolongation, structure function γ γ γ . result Unified scheme for equivalence problem of geometric structures.
General-purpose model learns visual reasoning without strong priors.
problem Achieving visual reasoning in image-related questions.
method Conditional Batch Normalization approach.
result 2.4% error rate on CLEVR Visual Reasoning benchmark.
SGD converges to critical points of normalized margin in late-stage training for homogeneous neural networks.
problem Analyzing the implicit bias of SGD on homogeneous neural networks.
method Interpreting SGD dynamics as an Euler-like discretization of a conservative field flow associated with the normalized classification margin.
result Normalized SGD iterates converge to the set of critical points of the normalized margin at late-stage training.
Paper introduces Categorical Normalizing Flows for better handling of categorical data.
problem Limited application of normalizing flows on categorical data due to lack of intrinsic order.
method Categorical Normalizing Flows use continuous transformations to model latent relations in categorical data, optimizing both continuous representation and model likelihood.
result GraphCNF, a permutation-invariant generative model, outperforms state-of-the-art on molecule generation.
Normalizing flow regression approximates posterior distributions without additional sampling.
problem Bayesian inference with computationally expensive likelihood evaluations.
method Normalizing flow regression (NFR) for offline inference.
result NFR yields a tractable posterior approximation through regression on existing log-density evaluations.