New method improves deep policy gradient algorithms by learning relative state values.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Parastatistic distribution of a total debt owed to a large number of creditors considered in relation to the duration of these debts. The process of debt calculation depends on the fractal dimension of economic system in which this process takes place. Two actual variants of these dimensions are investigated. Critical …
Smaller actor-critic models lead to performance degradation and overfitting, highlighting the critic's role in value underestimation.
Decouples critic chunk length from policy to improve policy reactivity and performance.
We give the classification, up to homeomorphisms, of reduced complex polynomials with 2 variables with one critical value.
New groups found with critical exponents close to but less than max.
Polynomials with distinct critical values have braid monodromy groups equal to braid groups.
Smooth manifolds have functions with exactly two critical values.
We study the problem of off-policy critic evaluation in several variants of value-based off-policy actor-critic algorithms. Off-policy actor-critic algorithms require an off-policy critic evaluation step, to estimate the value of the new policy after every policy gradient update. Despite enormous success of off-policy …
We describe how to compute topological objects associated to a polynomial map of several complex variables with isolated singularities. These objects are: the affine critical values, the affine Milnor numbers for all irregular fibers, the critical values at infinity, and the Milnor numbers at infinity for all irregular…
MAGE optimizes policies using action gradients from model-based learning.
We give a global version of Le-Ramanujam mu-constant theorem for polynomials. Let f_t, (t in [0,1]), be a family of polynomials of n complex variables with isolated singularities, whose coefficients are polynomials in t. We consider the case where some numerical invariants are constant (the affine Milnor number, the Mi…
ESAC improves reinforcement learning by lookahead and intuition.
Short note on upper bounds for loop homology classes.
A novel Q-learning variant reduces underestimation bias in deep actor-critic methods for reinforcement learning.
WAVE improves stability in reinforcement learning by adaptively weighting critic's loss.
Study magnetic geodesics on half-Lie groups, proving Hopf-Rinow theorem for energies above critical value.
In value-based reinforcement learning methods such as deep Q-learning, function approximation errors are known to lead to overestimated value estimates and suboptimal policies. We show that this problem persists in an actor-critic setting and propose novel mechanisms to minimize its effects on both the actor and the cr…
Study uses actor-critic method for continuous-time mean-field control with entropy regularisation.
We use noncommutative localization to construct a chain complex which counts the critical points of a circle-valued Morse function on a manifold, generalizing the Novikov complex. As a consequence we obtain new topological lower bounds on the minimum number of critical points of a circle-valued Morse function within a …
In traditional reinforcement learning, an agent maximizes the reward collected during its interaction with the environment by approximating the optimal policy through the estimation of value functions. Typically, given a state s and action a, the corresponding value is the expected discounted sum of rewards. The optima…
We introduce a new critical value for Tonelli Lagrangians on the tangent bundle of the 2-sphere without minimizing measures supported on a point. We show that is strictly larger than the Mañé critical value , and on every energy level there exist infinitely…
Formula for critical points of chi fields on manifolds.
Paper proves gradient estimates for Lagrangian mean curvature equation.
Recently many efforts have been made to incorporate persistence diagrams, one of the major tools in topological data analysis (TDA), into machine learning pipelines. To better understand the power and limitation of persistence diagrams, we carry out a range of experiments on both graph data and shape data, aiming to de…
When f : R power n to R power p, is a surjective real analytic map with isolated critical value, we prove that the (m)-regularity condition (in a sense we define) ensures that f ||f|| is a fibration on small spheres, f induces a fibration on the tubes and both fibrations are equivalent. In particular, we make the state…
Proposes a value-based method for continuous control without an actor.
PBVFs generalize across policies using learned value functions.
Study uncovers complex critical points in tensor decomposition.
We present the first provably convergent two-timescale off-policy actor-critic algorithm (COF-PAC) with function approximation. Key to COF-PAC is the introduction of a new critic, the emphasis critic, which is trained via Gradient Emphasis Learning (GEM), a novel combination of the key ideas of Gradient Temporal Differ…
We establish a new connection between value and policy based reinforcement learning (RL) based on a relationship between softmax temporal value consistency and policy optimality under entropy regularization. Specifically, we show that softmax consistent action values correspond to optimal entropy regularized policy pro…
We study the phenomenon that some modules of deep neural networks (DNNs) are more critical than others. Meaning that rewinding their parameter values back to initialization, while keeping other modules fixed at the trained parameters, results in a large drop in the network's performance. Our analysis reveals interestin…
The paper finds infinitely many magnetic geodesics on non-compact manifolds.
The paper proves extremal black holes form at a critical point of gravitational collapse.
Study magnetic geodesics on odd spheres, computing critical energy values.
Off-policy stochastic actor-critic methods rely on approximating the stochastic policy gradient in order to derive an optimal policy. One may also derive the optimal policy by approximating the action-value gradient. The use of action-value gradients is desirable as policy improvement occurs along the direction of stee…
This work, dealt with the classical mean value theorem and took advantage of it in the fractional calculus. The concept of a fractional critical point is introduced. Some sufficient conditions for the existence of a critical point is studied and an illustrative example rele- vant to the concept of the time dilation eff…
Solves local minima problems on smooth manifolds.
The z-transform technique is used to investigate the model for distribution of high-tax payers, which is proposed by two of the authors (K. Y and S. M) and others. Our analysis shows an asymptotic power-law of this model with the exponent -5/2 when a total ``mass'' has a certain critical value. Below the critical value…
This article introduces a framework to estimate the value of evidence-based decision making.
The paper defines a critical value for a magnetic system and extends solutions beyond blow-up.
Band-limited SAC improves learning efficiency and stability in simulated environments.
New algorithm solves mean-field control problems using actor-critic learning with moment neural networks.
Given any n-tuple of complex numbers, one can canonically define a polynomial of degree n+1 that has the entries of this n-tuple as its critical points. In 2002, Beardon, Carne, and Ng studied a map which outputs the critical values of the canonical polynomial constructed from the…
We study the space of smooth Riemannian structures on compact three-manifolds with boundary that satisfies a critical point equation associated with a boundary value problem, for simplicity, Miao-Tam critical metrics. We provide an estimate to the area of the boundary of Miao-Tam critical metrics on compact three-manif…
We consider the estimation of the policy gradient in partially observable Markov decision processes (POMDP) with a special class of structured policies that are finite-state controllers. We show that the gradient estimation can be done in the Actor-Critic framework, by making the critic compute a "value" function that …
USAC balances pessimism and optimism in actor-critic training for better exploration and performance.
Many policy gradient methods are variants of Actor-Critic (AC), where a value function (critic) is learned to facilitate updating the parameterized policy (actor). The update to the actor involves a log-likelihood update weighted by the action-values, with the addition of entropy regularization for soft variants. In th…