Doubly-stochastic normalization improves robustness to heteroskedastic noise.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Woodbury transformations improve deep generative models with efficient invertibility and determinant calculation.
New method estimates mutual information using normalizing flows.
The estimation of normalizing constants is a fundamental step in probabilistic model comparison. Sequential Monte Carlo methods may be used for this task and have the advantage of being inherently parallelizable. However, the standard choice of using a fixed number of particles at each iteration is suboptimal because s…
Network slicing promises to provision diversified services with distinct requirements in one infrastructure. Deep reinforcement learning (e.g., deep -learning, DQL) is assumed to be an appropriate algorithm to solve the demand-aware inter-slice resource management issue in network slicing by regarding the …
The normalized maximized likelihood (NML) provides the minimax regret solution in universal data compression, gambling, and prediction, and it plays an essential role in the minimum description length (MDL) method of statistical modeling and estimation. Here we show that the normalized maximum likelihood has a Bayes-li…
Bayesian inference for expensive likelihoods using Langevin Monte Carlo with NF.
Deep learning improves GW signal detection efficiency and robustness.
New algorithm tackles nonconvex machine learning problems with adaptive normalization and independent sampling.
Improves NF for complex data distributions with multiple modes.
Globally normalized neural sequence models are considered superior to their locally normalized equivalents because they may ameliorate the effects of label bias. However, when considering high-capacity neural parametrizations that condition on the whole input sequence, both model classes are theoretically equivalent in…
This paper improves normalizing flows by combining MLE and sliced-Wasserstein distance for better data fidelity.
Complete normal forms for specific real hypersurfaces in complex space are constructed.
SPO optimizes LLMs by eliminating group-based baselines and variance issues.
FF algorithm uses goodness as a likelihood-ratio test for scalar normalization.
A new method improves generative models by learning lower-dimensional representations.
This paper proposed a bias-compensated normalized maximum correntropy criterion (BCNMCC) algorithm charactered by its low steady-state misalignment for system identification with noisy input in an impulsive output noise environment. The normalized maximum correntropy criterion (NMCC) is derived from a correntropy based…
This paper develops a slice sampler for Bayesian linear regression models with arbitrary priors. The new sampler has two advantages over current approaches. One, it is faster than many custom implementations that rely on auxiliary latent variables, if the number of regressors is large. Two, it can be used with any prio…
Paper proposes a debiased estimator for adaptive linear regression.
Proposes group whitening to enhance deep learning models' performance.
We present sparse topical coding (STC), a non-probabilistic formulation of topic models for discovering latent representations of large collections of data. Unlike probabilistic topic models, STC relaxes the normalization constraint of admixture proportions and the constraint of defining a normalized likelihood functio…
Batch normalization with regularization turns deterministic autoencoders into generative models.
Efficiently computes optimal transport maps and Wasserstein barycenters using conditional normalizing flows.
Four new methods for computing generalized chi-square distribution.
Multiplicative stochasticity such as Dropout improves the robustness and generalizability of deep neural networks. Here, we further demonstrate that always-on multiplicative stochasticity combined with simple threshold neurons are sufficient operations for deep neural networks. We call such models Neural Sampling Machi…
This paper simplifies the Nash Bargaining Solution for use in intellectual property cases.
Vanna-Volga is a popular method for the interpolation/extrapolation of volatility smiles. The technique is widely used in the FX markets context, due to its ability to consistently construct the entire Lognormal smile using only three Lognormal market quotes. However, the derivation of the Vanna-Volga method itself is …
A recent strategy to circumvent the exploding and vanishing gradient problem in RNNs, and to allow the stable propagation of signals over long time scales, is to constrain recurrent connectivity matrices to be orthogonal or unitary. This ensures eigenvalues with unit norm and thus stable dynamics and training. However …
ButterflyFlow uses butterfly matrices for efficient invertible layers in normalizing flows.
Proposes online debiasing estimators for adaptive linear regression.
This paper puts forth a new formulation and algorithm for the elastic matching problem on unparametrized curves and surfaces. Our approach combines the frameworks of square root normal fields and varifold fidelity metrics into a novel framework, which has several potential advantages over previous works. First, our var…
This paper improves flow models to better handle perturbations in real-world data.
Computing partition functions, the normalizing constants of probability distributions, is often hard. Variants of importance sampling give unbiased estimates of a normalizer Z, however, unbiased estimates of the reciprocal 1/Z are harder to obtain. Unbiased estimates of 1/Z allow Markov chain Monte Carlo sampling of "d…
FF algorithm uses goodness as a measure of input quality, derived from likelihood-ratio tests.
In low-dimensional topology, many important decision algorithms are based on normal surface enumeration, which is a form of vertex enumeration over a high-dimensional and highly degenerate polytope. Because this enumeration is subject to extra combinatorial constraints, the only practical algorithms to date have been v…
In this paper we present an application of the use of autocopulas for modelling financial time series showing serial dependencies that are not necessarily linear. The approach presented here is semi-parametric in that it is characterized by a non-parametric autocopula and parametric marginals. One advantage of using au…
AlphaGrad optimizes memory usage in RL algorithms by normalizing gradients.
Regularization and normalization have become indispensable components in training deep neural networks, resulting in faster training and improved generalization performance. We propose the projected error function regularization loss (PER) that encourages activations to follow the standard normal distribution. PER rand…
The choice of approximate posterior distribution is one of the core problems in variational inference. Most applications of variational inference employ simple families of posterior approximations in order to allow for efficient inference, focusing on mean-field or other simple structured approximations. This restricti…
CIFs improve VI by providing flexible posteriors for complex topologies.
We prove that the marginal densities of a global probability mass function in a primal normal factor graph and the corresponding marginal densities in the dual normal factor graph are related via local mappings. The mapping depends on the Fourier transform of the local factors of the models. Details of the mapping, inc…
Paper proposes robust LAD estimators for 2D sinusoidal model, proving consistency and normality.
The paper analyzes the randomized midpoint method for Langevin diffusions, revealing biases and asymptotic properties.
This paper improves bandwidth selectors for SPBNs to enhance their performance.
The Information Bottleneck (IB) objective uses information theory to formulate a task-performance versus robustness trade-off. It has been successfully applied in the standard discriminative classification setting. We pose the question whether the IB can also be used to train generative likelihood models such as normal…
NVAE improves VAE performance on large image datasets.
TROLL improves RL for LLMs by replacing clipping with a trust region projection.
Generalized score matching for densities on general domains.