RevDEQs improve performance on tasks with exact gradients and fewer function evaluations.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New ODE solvers improve training efficiency and accuracy.
New path-gradient estimator for continuous normalizing flows.
A new method for optimal filtration learning in time-series data analysis.
New method improves NN performance across various settings.
We consider the case of derivative-free algorithms for non-convex optimization, also known as zero order algorithms, that use only function evaluations rather than gradients. For a wide variety of gradient approximators based on finite differences, we establish asymptotic convergence to second order stationary points u…
There are found exact values of (Matveev) complexity for the 2-parameter family of hyperbolic 3-manifolds with boundary constructed by Paoluzzi and Zimmermann. Moreover, -invariants for these manifolds are calculated.
We propose a general framework for solving statistical mechanics of systems with finite size. The approach extends the celebrated variational mean-field approaches using autoregressive neural networks, which support direct sampling and exact calculation of normalized probability of configurations. It computes variation…
ShuffleNet is a state-of-the-art light weight convolutional neural network architecture. Its basic operations include group, channel-wise convolution and channel shuffling. However, channel shuffling is manually designed empirically. Mathematically, shuffling is a multiplication by a permutation matrix. In this paper, …
New method calculates sensitivity of system failure probability.
We propose to learn deep undirected graphical models (i.e., MRFs) with a non-ELBO objective for which we can calculate exact gradients. In particular, we optimize a saddle-point objective deriving from the Bethe free energy approximation to the partition function. Unlike much recent work in approximate inference, the d…
Expressive quantum circuits are harder to train due to flatter cost landscapes.
Quantum annealer speeds up RBM training for image classification.
Much of studies on neural computation are based on network models of static neurons that produce analog output, despite the fact that information processing in the brain is predominantly carried out by dynamic neurons that produce discrete pulses called spikes. Research in spike-based computation has been impeded by th…
Efficiently calculates PL model likelihood for partitioned preference data.
In multi-objective Bayesian optimization and surrogate-based evolutionary algorithms, Expected HyperVolume Improvement (EHVI) is widely used as the acquisition function to guide the search approaching the Pareto front. This paper focuses on the exact calculation of EHVI given a nondominated set, for which the existing …
We study the Wasserstein natural gradient in parametric statistical models with continuous sample spaces. Our approach is to pull back the -Wasserstein metric tensor in the probability density space to a parameter space, equipping the latter with a positive definite metric tensor, under which it becomes a Riemanni…
Thompson sampling (TS) is a class of algorithms for sequential decision-making, which requires maintaining a posterior distribution over a model. However, calculating exact posterior distributions is intractable for all but the simplest models. Consequently, efficient computation of an approximate posterior distributio…
Gradient descent achieves exact linear convergence rate for symmetric matrix completion.
TERA method speeds up derivative Gaussian processes in high dimensions.
In this paper, we study the inequality indices for some models of wealth exchange. We calculated Gini index and newly introduced k-index and compare the results with reported empirical data available for different countries. We have found lower and upper bounds for the indices and discuss the efficiencies of the models…
We analyze how an observer synchronizes to the internal state of a finite-state information source, using the epsilon-machine causal representation. Here, we treat the case of exact synchronization, when it is possible for the observer to synchronize completely after a finite number of observations. The more difficult …
We propose Kernel Hamiltonian Monte Carlo (KMC), a gradient-free adaptive MCMC algorithm based on Hamiltonian Monte Carlo (HMC). On target densities where classical HMC is not an option due to intractable gradients, KMC adaptively learns the target's gradient structure by fitting an exponential family model in a Reprod…
In this work, we consider the use of model-driven deep learning techniques for massive multiple-input multiple-output (MIMO) detection. Compared with conventional MIMO systems, massive MIMO promises improved spectral efficiency, coverage and range. Unfortunately, these benefits are coming at the cost of significantly i…
SGD reduces test error by decorrelating updates.
New method differentiates square-root Kalman filters robustly.
FastAMI efficiently approximates AMI and SMI for large datasets.
Analyzes the generalization and training errors of the random feature model over time.
We calculate the smooth structure set of , , for and . As a consequence we show that in general cannot admit a group structure such that the smooth surgery exact sequence is a long exact sequence of groups. We also show that the image of forgetful map $F:…
Gradient descent slows significantly in over-parameterized single neuron learning.
In an earlier paper (math.SG/0101206), we introduced Floer homology theories associated to closed, oriented three-manifolds Y and SpinC structures. In the present paper, we give calculations and study the properties of these invariants. The calculations suggest a conjectured relationship with Seiberg-Witten theory. The…
We use the theory of large deviations to study the pricing of investment-grade tranches of synthetic CDO's. In this paper, we consider a simplified model which will allow us to introduce some of the concepts and calculations.
In this paper, we generalize the notion of Serre fibration to the Morita category of topological groupoids and derive the associated long exact sequence of homotopy groups. We use this results for calculation of homotopy groups of various groupoids, such as the foliation groupoid of a Riemannian foliation.
Paper extends knowledge gradient for preferential BO, overcoming computational challenges.
In the presence of boundaries the integrated conformal anomaly is modified by the boundary terms so that the anomaly is non-vanishing in any (even or odd) dimension. The boundary terms are due to extrinsic curvature whose exact structure in and has recently been identified. In this note we present a hologra…
Gradient coding is a technique for straggler mitigation in distributed learning. In this paper we design novel gradient codes using tools from classical coding theory, namely, cyclic MDS codes, which compare favorably with existing solutions, both in the applicable range of parameters and in the complexity of the invol…
Contrastive Divergence (CD) and Persistent Contrastive Divergence (PCD) are popular methods for training the weights of Restricted Boltzmann Machines. However, both methods use an approximate method for sampling from the model distribution. As a side effect, these approximations yield significantly different biases and…
Various bias-correction methods such as EXTRA, gradient tracking methods, and exact diffusion have been proposed recently to solve distributed {\em deterministic} optimization problems. These methods employ constant step-sizes and converge linearly to the {\em exact} solution under proper conditions. However, their per…
Bilevel optimization has been recently revisited for designing and analyzing algorithms in hyperparameter tuning and meta learning tasks. However, due to its nested structure, evaluating exact gradients for high-dimensional problems is computationally challenging. One heuristic to circumvent this difficulty is to use t…
Deep neural networks achieve stellar generalisation even when they have enough parameters to easily fit all their training data. We study this phenomenon by analysing the dynamics and the performance of over-parameterised two-layer neural networks in the teacher-student setup, where one network, the student, is trained…
The paper studies the solution of stochastic optimization problems in which approximations to the gradient and Hessian are obtained through subsampling. We first consider Newton-like methods that employ these approximations and discuss how to coordinate the accuracy in the gradient and Hessian to yield a superlinear ra…
The article is devoted to models of financial markets with stochastic volatility, which is defined by a functional of Ornstein-Uhlenbeck process or Cox-Ingersoll-Ross process. We study the question of exact price of European option. The form of the density function of the random variable, which expresses the average of…
The original k-means clustering method works only if the exact vectors representing the data points are known. Therefore calculating the distances from the centroids needs vector operations, since the average of abstract data points is undefined. Existing algorithms can be extended for those cases when the sole input i…
Researchers calculate exact moduli for type II flux backgrounds using spectral sequences.
Improves meta-learning efficiency with mixed-mode differentiation.
The challenge of assigning importance to individual neurons in a network is of interest when interpreting deep learning models. In recent work, Dhamdhere et al. proposed Total Conductance, a "natural refinement of Integrated Gradients" for attributing importance to internal neurons. Unfortunately, the authors found tha…
Researchers use Gysin sequence to show sl(N) homology of T(2,m) is cohomology of SU(N) representations.
A new parallel BO method with exact gradients for multi-objective optimization.