A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
We consider actions of Z^k, k \ge 2, by Anosov diffeomorphisms which are uniformly quasiconformal on each coarse Lyapunov distribution. These actions generalize Cartan actions for which coarse Lyapunov distributions are one-dimensional. We show that, under certain non-resonance assumptions on the Lyapunov exponents, a …
DRL agents perform poorly at high decision frequencies, but a new algorithm improves performance.
problem DRL agents struggle at high decision frequencies, leading to poor performance.
method Proved that DRL agents' action-conditioned return distributions collapse to their policy's return distribution as decision frequency increases. Defined superiority as a probabilistic generalization of advantage for high-frequency value-based RL.
result Proper modeling of superiority distribution improves performance of controllers at high decision frequencies.
Network slicing is a key technology in 5G communications system. Its purpose is to dynamically and efficiently allocate resources for diversified services with distinct requirements over a common underlying physical infrastructure. Therein, demand-aware resource allocation is of significant importance to network slicin…
Stochastic games provide a framework for interactions among multiple agents and enable a myriad of applications. In these games, agents decide on actions simultaneously, the state of every agent moves to the next state, and each agent receives a reward. However, finding an equilibrium (if exists) in this game is often …
As Computer Vision moves from a passive analysis of pixels to active analysis of semantics, the breadth of information algorithms need to reason over has expanded significantly. One of the key challenges in this vein is the ability to identify the information required to make a decision, and select an action that will …
Let F be a Lie foliation on a closed manifold M with structural Lie group G. Its transverse Lie structure can be considered as a transverse action Φ of G on (M,F); i.e., an ``action'' which is defined up to leafwise homotopies. This Φ induces an action Φ∗ of G on the reduced leafwis…
We consider a totally nonsymplectic Anosov action of Z^k which is either uniformly quasiconformal or pinched on each coarse Lyapunov distribution. We show that such an action on a torus is C^\infty--conjugate to an action by affine automorphisms. We also obtain similar global rigidity results for actions on an arbitrar…
Estimation of importance sampling weights for off-policy evaluation of contextual bandits often results in imbalance - a mismatch between the desired and the actual distribution of state-action pairs after weighting. In this work we present balanced off-policy evaluation (B-OPE), a generic method for estimating weights…
We develop variation formulas on almost-product (e.g. foliated) pseudo-Riemannian manifolds, and we consider variations of metric preserving orthogonality of the distributions. These formulae are applied to Einstein-Hilbert type actions: the total mixed scalar curvature and the total extrinsic scalar curvature of a dis…
The results of R^2 dynamical random surface model (2-dimensional quantum gravity with a R2 term) are applied to explain the personal income distribution. A scale invariance exists if there is not the R2 term in the action. The R^2 term provides a typical scale and breaks the scale invariance explicitly in the low…
In traditional reinforcement learning, an agent maximizes the reward collected during its interaction with the environment by approximating the optimal policy through the estimation of value functions. Typically, given a state s and action a, the corresponding value is the expected discounted sum of rewards. The optima…
In this paper we propose a novel framework for decentralized, online learning by many learners. At each moment of time, an instance characterized by a certain context may arrive to each learner; based on the context, the learner can select one of its own actions (which gives a reward and provides information) or reques…
We identify a fundamental problem in policy gradient-based methods in continuous control. As policy gradient methods require the agent's underlying probability distribution, they limit policy representation to parametric distribution classes. We show that optimizing over such sets results in local movement in the actio…
We show that the cone associated with a moment map for an action of a torus on a contact compact connected manifold is a convex polyhedral cone and that the moment map has connected fibers provided the dimension of the torus is bigger than 2 and that no orbit is tangent to the contact distribution. This may be consider…
Empowerment quantifies the influence an agent has on its environment. This is formally achieved by the maximum of the expected KL-divergence between the distribution of the successor state conditioned on a specific action and a distribution where the actions are marginalised out. This is a natural candidate for an intr…
We introduce a method for learning the dynamics of complex nonlinear systems based on deep generative models over temporal segments of states and actions. Unlike dynamics models that operate over individual discrete timesteps, we learn the distribution over future state trajectories conditioned on past state, past acti…
We introduce equivariant Liouville forms and Duistermaat-Heckman distributions for Hamiltonian group actions with group valued moment maps. The theory is illustrated by applications to moduli spaces of flat connections on 2-manifolds.