Training recurrent neural networks (RNNs) is a hard problem due to degeneracies in the optimization landscape, a problem also known as vanishing/exploding gradients. Short of designing new RNN architectures, previous methods for dealing with this problem usually boil down to orthogonalization of the recurrent dynamics,…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Causal normalizing flows recover causal models from observational data.
We solve image inverse problems using a flow-based noise model.
The presence of non linear instruments is responsible for the emergence of non Gaussian features in the price changes distribution of realistic portfolios, even for Normally distributed risk factors. This is especially true for the benchmark Delta Gamma Normal model, which in general exhibits exponentially damped power…
In recent work on both generative and discriminative score to log-likelihood-ratio calibration, it was shown that linear transforms give good accuracy only for a limited range of operating points. Moreover, these methods required tailoring of the calibration training objective functions in order to target the desired r…
Cash managers make daily decisions based on predicted monetary inflows from debtors and outflows to creditors. Usual assumptions on the statistical properties of daily net cash flow include normality, absence of correlation and stationarity. We provide a comprehensive study based on a real-world cash flow data set from…
AlphaGrad optimizes memory usage in RL algorithms by normalizing gradients.
INF-clip optimizes heavy-tailed MAB problems with improved performance.
Enhances RJMCMC efficiency with non-linear transport-based proposals.
Study high-dimensional Bayesian linear regression using variational inference.
Study shows depth improves trainability of neural networks by improving kernel conditioning.
Layer normalization with activations prevents Gram matrix rank collapse at initialization.
Training state-of-the-art, deep neural networks is computationally expensive. One way to reduce the training time is to normalize the activities of the neurons. A recently introduced technique called batch normalization uses the distribution of the summed input to a neuron over a mini-batch of training cases to compute…
Improved binning technique boosts nUV measure performance.
When approximating a black-box function, sampling with active learning focussing on regions with non-linear responses tends to improve accuracy. We present the FLOLA-Voronoi method introduced previously for deterministic responses, and theoretically derive the impact of output uncertainty. The algorithm automatically p…
We study the regret minimization problem in the novel setting of generalized kernelized bandits (GKBs), where we optimize an unknown function belonging to a reproducing kernel Hilbert space (RKHS) having access to samples generated by an exponential family (EF) reward model whose mean is a non-linear function $μ(…
We consider the group of sense-preserving diffeomorphisms $\Diff S^1$ of the unit circle and its central extension, the Virasoro-Bott group, with their respective horizontal distributions chosen to be Ehresmann connections with respect to a projection to the smooth universal Teichmüller space and the universal Teichmül…
This work improves manifold learning for multi-modal data.
This paper studies semiparametric contextual bandits, a generalization of the linear stochastic bandit problem where the reward for an action is modeled as a linear function of known action features confounded by an non-linear action-independent term. We design new algorithms that achieve regret …
We provide a pointwise confidence bound for non-linear least-squares with fixed design.
New measures detect asymmetries, non-linearity in stock returns.
The enumeration of normal surfaces is a key bottleneck in computational three-dimensional topology. The underlying procedure is the enumeration of admissible vertices of a high-dimensional polytope, where admissibility is a powerful but non-linear and non-convex constraint. The main results of this paper are significan…
Capitalizing on the need for addressing the existing challenges associated with gesture recognition via sparse multichannel surface Electromyography (sEMG) signals, the paper proposes a novel deep learning model, referred to as the XceptionTime architecture. The proposed innovative XceptionTime is designed by integrati…
Utilizing recently introduced concepts from statistics and quantitative risk management, we present a general variant of Batch Normalization (BN) that offers accelerated convergence of Neural Network training compared to conventional BN. In general, we show that mean and standard deviation are not always the most appro…
Normal distributions ensure asymptotic variance reduction in moment matching Monte Carlo.
Computer vision model automates residual plot assessment for diagnosing model assumptions.
The study uses CoDa to analyze family business financial ratios, highlighting methodological issues.
This paper offers a precise analytical characterization of the distribution of returns for a portfolio constituted of assets whose returns are described by an arbitrary joint multivariate distribution. In this goal, we introduce a non-linear transformation that maps the returns onto gaussian variables whose covariance …
High-risk domains require reliable confidence estimates from predictive models. Deep latent variable models provide these, but suffer from the rigid variational distributions used for tractable inference, which err on the side of overconfidence. We propose Stochastic Quantized Activation Distributions (SQUAD), which im…
We introduce a new distance metric for non-linear embeddings of Tempered Exponential Measures.
Batch Normalization (BN) is essential to effectively train state-of-the-art deep Convolutional Neural Networks (CNN). It normalizes inputs to the layers during training using the statistics of each mini-batch. In this work, we study BN from the viewpoint of Fisher kernels. We show that assuming samples within a mini-ba…
We construct a tangent bundle exponential map and locally autoparallel coordinates for geometries based on a general connection on the tangent bundle of a manifold. As concrete application we use these new coordinates for Finslerian geometries and obtain Finslerian geodesic coordinates. They generalise normal coordinat…
We address the problem of estimating statistics of hidden units in a neural network using a method of analytic moment propagation. These statistics are useful for approximate whitening of the inputs in front of saturating non-linearities such as a sigmoid function. This is important for initialization of training and f…
Paper improves basket option pricing for log-normal models.
We generalize the concept of sub-Riemannian geometry to infinite-dimensional manifolds modeled on convenient vector spaces. On a sub-Riemannian manifold , the metric is defined only on a sub-bundle $\calH$ of the tangent bundle , called the horizontal distribution. Similarly to the finite-dimensional case, we ar…
The paper proves conditions for curvature blow-up in quiescent big bang singularities.
The paper estimates CoVaR with various models for financial risk analysis.
TrajectoryNet models dynamic cellular trajectories using optimal transport.
Combines MCTM and NF for flexible multivariate density regression with interpretable marginals.
We study higher-order conservation laws of the non-linearizable elliptic Poisson equation as elements of the characteristic cohomology of the associated exterior differential system. The theory of characteristic cohomology determines a normal form for diffe…
Mixture of Experts (MoE) is a popular framework in the fields of statistics and machine learning for modeling heterogeneity in data for regression, classification and clustering. MoE for continuous data are usually based on the normal distribution. However, it is known that for data with asymmetric behavior, heavy tail…
Mixture of Experts (MoE) is a popular framework for modeling heterogeneity in data for regression, classification and clustering. For continuous data which we consider here in the context of regression and cluster analysis, MoE usually use normal experts, that is, expert components following the Gaussian distribution. …
LOCA learns standardized data coordinates from measurements.
Local Hebbian learning is believed to be inferior in performance to end-to-end training using a backpropagation algorithm. We question this popular belief by designing a local algorithm that can learn convolutional filters at scale on large image datasets. These filters combined with patch normalization and very steep …
Geometric Variational Inference improves efficiency in complex probability distributions.
We discuss the geometric foundation behind the use of stochastic processes in the frame bundle of a smooth manifold to build stochastic models with applications in statistical analysis of non-linear data. The transition densities for the projection to the manifold of Brownian motions developed in the frame bundle lead …
We propose a new way of thinking about deep neural networks, in which the linear and non-linear components of the network are naturally derived and justified in terms of principles in probability theory. In particular, the models constructed in our framework assign probabilities to uncertain realizations, leading to Ku…
We propose a unified methodology to input non-linear views from any number of users in fully general non-normal markets, and perform, among others, stress-testing, scenario analysis, and ranking allocation. We walk the reader through the theory and we detail an extremely efficient algorithm to easily implement this met…