Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,982 papers · 148 categories

Trend · papers per month

151302452603 · Jun 202019922001200920172026
48 results for Parameterized Action Space

New algorithms tackle multi-agent problems with hybrid action spaces.

problem Applying deep reinforcement learning to multi-agent problems with discrete-continuous hybrid action spaces.
method Proposed two novel algorithms: Deep MAPQN and Deep MAHHQN, using centralized training and decentralized execution.
result Empirical results show both algorithms significantly outperform existing methods.

We consider reinforcement learning in parameterized Markov Decision Processes (MDPs), where the parameterization may induce correlation across transition probabilities or rewards. Consequently, observing a particular state transition might yield useful information about other, unobserved, parts of the MDP. We present a…

2014-06-29abs ↗pdf ↗

Hybrid SAC improves RL for video games with discrete, continuous actions.

problem Improving RL performance in video games with practical constraints.
method Extension of Soft Actor-Critic (SAC) for handling discrete, continuous, and parameterized actions.
result Hybrid SAC successfully solves a high-speed driving task and is competitive on parameterized actions benchmarks.

We propose a new sample-efficient methodology, called Supervised Policy Update (SPU), for deep reinforcement learning. Starting with data generated by the current policy, SPU formulates and solves a constrained optimization problem in the non-parameterized proximal policy space. Using supervised regression, it then con…

2018-05-29abs ↗pdf ↗

Paper presents an efficient exploration method for reinforcement learning.

problem Efficient exploration in reinforcement learning with uncertainty quantification.
method Parameterized Indexed Value Function (PIV) using index sampling.
result Proves the regret bound for learning PIV in a tabular setting and proposes PINs for computational learning.

The recently proposed option-critic architecture Bacon et al. provide a stochastic policy gradient approach to hierarchical reinforcement learning. Specifically, they provide a way to estimate the gradient of the expected discounted return with respect to parameters that define a finite number of temporally extended ac…

2018-12-04abs ↗pdf ↗

It has long been assumed that high dimensional continuous control problems cannot be solved effectively by discretizing individual dimensions of the action space due to the exponentially large number of bins over which policies would have to be learned. In this paper, we draw inspiration from the recent success of sequ…

2017-05-14abs ↗pdf ↗

In this short note, we prove that the space of all admissible piecewise linear metrics parameterized by length square on a triangulated manifolds is a convex cone. We further study Regge's Einstein-Hilbert action and give a much more reasonable definition of discrete Einstein metric than our former version in \cite{G}.…

2015-08-25abs ↗pdf ↗

hyperSBINN improves drug cardiosafety assessment by efficiently modeling cardiac action potentials.

problem Complexity and limited data in modeling cardiac effects of drugs.
method Combining meta-learning with SBINNs to solve parameterized cardiac action potential models.
result hyperSBINN outperforms traditional solvers in speed and accuracy for predicting APD90 values.

In this paper we describe the space of maximal components of the character variety of surface group representations into PSp(4,R) and Sp(4,R). For every rank 2 real Lie group of Hermitian type, we construct a mapping class group invariant complex structure on the maximal components. For the groups PSp(4,R) and Sp(4,R),…

2017-08-17abs ↗pdf ↗

Policy gradient converges linearly with Hadamard parameterization in tabular settings.

problem Convergence of policy gradient methods under Hadamard parameterization.
method Studied convergence rate and established linear convergence after k0k_0 iterations.
result Algorithm converges linearly with rate $O( rac{1}{k})$ and faster locally after k0k_0.

This is a survey of the theory of complex projective (CP^1) structures on compact surfaces. After some preliminary discussion and definitions, we concentrate on three main topics: (1) Using the Schwarzian derivative to parameterize the moduli space (2) Thurston's parameterization of the moduli space using grafting (3) …

2009-02-11abs ↗pdf ↗

Continuous control imitation learning fails if expert actions are smooth.

problem Continuous control imitation learning fails if expert actions are smooth.
method Study of imitation learning in discrete-time, continuous state-and-action control systems.
result Any smooth, deterministic imitator policy suffers exponentially larger error than the expert.

This paper improves reinforcement learning efficiency for large-scale MDPs.

problem High sample complexity in tabular RL settings with large state and action spaces.
method Model-based approach and Q-learning with linearly parameterized features.
result Provably efficient learning with sample complexity bounds.

Develops a new reinforcement learning framework for complex control problems.

problem Continuous-time extended mean field control with deterministic policies.
method Model-free sensitivity formula, deterministic policy gradient, local value and advantage-rate representations.
result Demonstrates efficiency, stability, and robustness in solving complex control problems.

Local PCA detects intrinsic parameterization of complex thermo-chemical state-spaces.

problem Detecting intrinsic parameterization of complex thermo-chemical state-spaces.
method Local PCA applied to local clusters of data.
result Local PCA finds meaningful parameterization linked to local stoichiometry, reaction progress, and soot formation processes.

This paper identifies drift Lipschitz budget K as key to diffusion policy expressivity and statistical trade-offs.

problem Understanding and maximizing the expressivity of diffusion policies while managing statistical limitations.
method Identifying drift Lipschitz budget K as central, quantifying expressivity and statistical behavior, proving lower bounds, and providing practical implementation guidelines.
result Balancing expressivity and statistical complexity yields a finite-sample performance gap, with rates depending on sample size and drift type.

Study on rotational surfaces in de Sitter space with specific curvature conditions.

problem Characterizing rotational surfaces in de Sitter space with Weingarten conditions.
method Analyzing spacelike and timelike rotational surfaces in 3D de Sitter space, determining profile curves, and classifying surfaces based on curvature relations.
result Classification of Weingarten rotational surfaces in de Sitter space with specific curvature relations.

We prove generic regularity and Uhlenbeck-type compactification theorems for the moduli spaces of PU(2)-monopoles. Generic regularity is NOT obtained in the usual way (by applying Sard theorem to a smooth parameterized moduli space), since the parameterized moduli space can be a priori singular. We explain why, using t…

1999-06-24abs ↗pdf ↗

We show that the real-valued function SαS_α on the moduli space M0,n\mathcal{M}_{0,n} of pointed rational curves, defined as the critical value of the Liouville action functional on a hyperbolic 2-sphere with n3n\geq 3 conical singularities of arbitrary orders α={α1,...,αn}α=\{α_1,...,α_n\}, generates accessory parameters of the as…

2001-12-17abs ↗pdf ↗

Geometric Occam's Razor shapes deep learning solutions.

problem Understanding the regularization in over-parameterized neural networks.
method Analyzing the geometric model complexity and Dirichlet energy in neural networks.
result Over-parameterized neural networks are implicitly regularized by geometric model complexity.

We study a set MK,N\mathcal{M}_{K,N} parameterizing filtered SL(K)SL(K)-Higgs bundles over CP1\mathbb{CP}^1 with an irregular singularity at z=z = \infty, such that the eigenvalues of the Higgs field grow like λzN/Kdz\lvert λ\rvert \sim \lvert z ^{N/K} \mathrm{d} z \rvert, where KK and NN are coprime. MK,N\mathcal{M}_{K,N} carrie…

2017-09-18abs ↗pdf ↗

Method converts neural networks to function space for better uncertainty quantification.

problem Lack of uncertainty estimates and difficulty in incorporating new data in deep neural networks.
method Dual parameterization to convert from weight space to function space, enabling sparse representation.
result Compact and principled way to capture uncertainty and incorporate new data.

New geometric interpretation explains over-parameterized models and adversarial perturbations.

problem Geometric understanding of over-parameterized regression and adversarial perturbations.
method Alternative geometric interpretation of regression in feature space.
result Adversarial perturbations are a natural feature of biased models due to underlying geometry.

Visualizes movement control optimization landscapes to understand why it's hard and how to make it easier.

problem Understanding and optimizing movement control problems in animation research.
method Novel visualizations of high-dimensional control optimization landscapes.
result Trajectory optimization becomes increasingly ill-conditioned with longer trajectories, while parameterizing control as partial target states can act as an efficient preconditioner.

The paper studies holomorphic curves in a pseudo-Riemannian space and their moduli space.

problem Understanding the moduli space of holomorphic curves in a pseudo-Riemannian space.
method Using Frenet framing and G2G_2'-Higgs bundles, the paper describes the moduli space of equivariant alternating holomorphic curves.
result Equivariant alternating holomorphic curves are infinitesimally rigid.

Paper proposes an efficient RL algorithm for discounted MDPs using feature mapping.

problem Efficient reinforcement learning for large state and action spaces.
method Uses feature mapping to represent states and actions in a low-dimensional space, proposing a novel algorithm with polynomial regret bound.
result Achieves a O(dT/(1γ)2)O(d\sqrt{T}/(1-γ)^2) regret bound, near-optimal up to a (1γ)0.5(1-γ)^{-0.5} factor.

Kirchhoff energy is a classical functional on the space of arclength-parameterized framed curves whose critical points approximate configurations of springy elastic rods. We introduce a generalized functional on the space of framed curves of arbitrary parameterization, which model rods with axial stretch or cross-secti…

2017-08-30abs ↗pdf ↗

We establish the existence of an integer degree for the natural projection map from the space of parameterizations of asymptotically conical self-expanders to the space of parameterizations of the asymptotic cones when this map is proper. As an application we show that there is an open set in the space of cones in the …

2018-07-17abs ↗pdf ↗

HOPE improves SSMs for long-memory tasks with robust initialization and training.

problem Improving state-space models for long-memory tasks with robust initialization and training.
method Developed a new parameterization scheme called HOPE using Hankel operators and Markov parameters.
result HOPE improves SSMs' performance on Long-Range Arena tasks and demonstrates non-decaying memory.

We present a classification of compact Kaehler manifolds admitting a hamiltonian 2-form (which were classified locally in part I of this work). This involves two components of independent interest. The first is the notion of a rigid hamiltonian torus action. This natural condition, for torus actions on a Kaehler manifo…

2004-01-23abs ↗pdf ↗