Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

51101152202 · Jun 202019922001200920172026
48 results for epistemically-derived GF probability

Geometrical flows (GF) play an important role in modern mathematics and physics. In this letter we have considered some integrable isotropic GF -- Ricci flows (RF) and mean curvature flows (MCF) -- which are related with integrable Heisenberg ferromagnets. In 2+1 dimensions, these GF have a singularity at t=t0t=t_{0}.

2008-04-05abs ↗pdf ↗

Gradient flow on ReLU networks converges to a simple model with few regions.

problem Understanding the dynamics of gradient flow in shallow ReLU networks.
method Analysis of gradient flow dynamics on univariate ReLU neural networks.
result Gradient flow converges to a network with at most O(r) linear regions.

In this work, we propose a novel recurrent neural network (RNN) architecture. The proposed RNN, gated-feedback RNN (GF-RNN), extends the existing approach of stacking multiple recurrent layers by allowing and controlling signals flowing from upper recurrent layers to lower layers using a global gating unit for each pai…

2015-02-09abs ↗pdf ↗

Gradient descent and SGD achieve low test error in specific network weight regimes.

problem Optimizing two-layer ReLU networks with standard initialization.
method Gradient flow and stochastic gradient descent, analyzing margins and weight norms.
result Gradient descent and SGD can achieve globally maximal margins under certain constraints.

Stein variational gradient decent (SVGD) has been shown to be a powerful approximate inference algorithm for complex distributions. However, the standard SVGD requires calculating the gradient of the target density and cannot be applied when the gradient is unavailable. In this work, we develop a gradient-free variant …

2018-06-07abs ↗pdf ↗

Filtering is a general name for inferring the states of a dynamical system given observations. The most common filtering approach is Gaussian Filtering (GF) where the distribution of the inferred states is a Gaussian whose mean is an affine function of the observations. There are two restrictions in this model: Gaussia…

2018-11-14abs ↗pdf ↗

Wide neural networks converge linearly to zero loss with feature learning.

problem Optimizing wide neural networks with feature learning guarantees.
method Gradient flow analysis for wide shallow and multi-layer NNs.
result Training loss converges linearly to zero for wide NNs under GF, demonstrating feature learning and better generalization.

Let N^h be a hyperbolic 3-manifold of bounded geometry corresponding to a hyperbolic structure on a pared manifold (M,P). Further, suppose that (\partial{M} - P) is incompressible, i.e. the boundary of M is incompressible away from cusps. Further, suppose that M_{gf} is a geometrically finite hyperbolic structure on (M…

2005-03-25abs ↗pdf ↗

Gradient descent converges to a global minimum in nonlinear ReLU implicit networks with linear width.

problem Understanding convergence of gradient methods in nonlinear, infinitely deep ReLU networks.
method Introduced a scaling constant to ensure well-posedness of the equilibrium equation, proving convergence to a global minimum for linear width networks.
result Gradient descent converges to a global minimum at a linear rate for nonlinear ReLU implicit networks with linear width.

Graph neural networks (GNNs) have been shown to replicate convolutional neural networks' (CNNs) superior performance in many problems involving graphs. By replacing regular convolutions with linear shift-invariant graph filters (LSI-GFs), GNNs take into account the (irregular) structure of the graph and provide meaning…

2018-10-29abs ↗pdf ↗

Develops a new framework to analyze gradient flow regimes and derive explicit solutions.

problem Analyzing scaling regimes and deriving explicit analytic solutions for gradient flow in large learning problems.
method Formal power series expansion of the loss evolution with coefficients encoded by diagrams.
result Reveals different learning phases and obtains explicit solutions in some cases.

We derive bounds on the path length ζζ of gradient descent (GD) and gradient flow (GF) curves for various classes of smooth convex and nonconvex functions. Among other results, we prove that: (a) if the iterates are linearly convergent with factor (1c)(1-c), then ζζ is at most O(1/c)\mathcal{O}(1/c); (b) under the Polyak-K…

2019-08-02abs ↗pdf ↗

ReLU networks implicitly favor low-rank solutions, but not as strongly as linear networks.

problem Understanding implicit regularization in ReLU networks for rank minimization.
method Analysis of gradient flow on ReLU networks, empirical testing.
result Gradient flow on ReLU networks does not necessarily minimize ranks, unlike in linear networks.

GF-Net learns Green's functions for linear reaction-diffusion equations.

problem Learning Green's functions for linear reaction-diffusion equations on arbitrary domains.
method GF-Net, a neural network, learns Green's functions in an unsupervised manner using physics-informed approach and symmetry.
result GF-Net efficiently solves linear reaction-diffusion equations under various boundary conditions and sources.

Derives EoM for DNNs to describe GD dynamics precisely.

problem Gaps between differential equations and actual DNN learning dynamics due to discretization error.
method Starts from GF, derives counter term to cancel discretization error, obtains EoM.
result EoM precisely describes GD dynamics of DNNs, highlights differences between continuous and discrete GD.

The paper is divided in 2 parts. The first part is the original paper of the second and third authors arXiv:1202.5442v2. The second part is an erratum/addendum written in english and concatenated at the end of the former paper. In the erratum/addentum, we amend Theorems 1.3 and 1.11 of arXiv:1202.5442v2: Finitude géomé…

2012-02-24abs ↗pdf ↗

Proposes a method to improve surrogate modeling and design optimization using latent variables.

problem Improving efficiency in multi-fidelity adaptive sampling without hierarchical assumptions.
method A framework using a latent variable Gaussian process to capture correlations between different fidelity models and optimize adaptive sampling.
result Demonstrates superior performance in convergence rate and robustness compared to existing methods.

Despite the advantages of all-weather and all-day high-resolution imaging, SAR remote sensing images are much less viewed and used by general people because human vision is not adapted to microwave scattering phenomenon. However, expert interpreters can be trained by compare side-by-side SAR and optical images to learn…

2019-01-08abs ↗pdf ↗

Funar algebra K=K(α,β;k)K_\infty=K_\infty(α,β;k) is the quotient of the group algebra over a ring kk of the braid group BB_\infty by two cubic relations: σ13ασ12+βσ11=0σ_1^3-ασ_1^2+βσ_1-1=0 and another one which involves σ1σ_1 and σ2σ_2. The universal Markov trace on KK_\infty is the quotient map tt of K(α,β,k[u,v])K_\infty(α,β,k[u,v]) to its qu…

2012-06-04abs ↗pdf ↗

We extend the notion of link colorings with values in an Alexander quandle to link colorings with values in a module MM over the Laurent polynomial ring Λμ=Z[t1±1,,tμ±1]Λ_μ=\mathbb{Z}[t_1^{\pm1},\dots,t_μ^{\pm1}]. If DD is a diagram of a link LL with μμ components, then the colorings of DD with values in MM form a ΛμΛ_μ-module…

2018-05-06abs ↗pdf ↗

New concept of attitude towards probability introduced in risk sharing problems.

problem Risk sharing problems and attitudes towards probability.
method Generalized definition of probability premium, local approximation, rank-dependent utility model, dual theory.
result Attitude towards probability can be first-order or second-order, depending on the model.

Identifies conditions for multiple invariant probabilities in Markov kernels.

problem Global irreducibility and recurrence do not guarantee uniqueness of invariant probabilities.
method Uses Jordan decomposition of the difference of two invariant probabilities.
result A Markov kernel has more than one invariant probability if and only if it admits a visible absorbing decomposition.

This work improves deep neural network probability estimation methods.

problem Estimating probabilities from high-dimensional data with inherent uncertainty.
method Investigates and compares methods for probability estimation using deep neural networks, proposing a new method that promotes consistent probabilities.
result The new method outperforms existing approaches on most metrics on simulated and real-world data.

This work presents a new classifier that is specifically designed to be fully interpretable. This technique determines the probability of a class outcome, based directly on probability assignments measured from the training data. The accuracy of the predicted probability can be improved by measuring more probability es…

2017-10-27abs ↗pdf ↗

Investigates statistical properties of perturb-softmax and perturb-argmax distributions.

problem Underexplored statistical properties of Gumbel-Softmax and Gumbel-Argmax distributions.
method Investigates convexity and differentiability to determine completeness and minimality of these distributions.
result Identifies parameters that admit complete and minimal representation of probability distributions.

Paper constructs unfaithful probability distributions in binary causal graphs.

problem Unfaithful probability distributions in binary causal graphs.
method Constructs unfaithful probability distributions in binary causal graphs.
result Examples of unfaithful probability distributions in binary causal graphs.