A novel framework interprets driving patterns using Action phases clustering.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Describes reconstructing Poisson structures from Lie group actions.
Detects anomalies in product health metrics at eBay for better alerts.
Bandit algorithms have various application in safety-critical systems, where it is important to respect the system constraints that rely on the bandit's unknown parameters at every round. In this paper, we formulate a linear stochastic multi-armed bandit problem with safety constraints that depend (linearly) on an unkn…
Poisson and symplectic structures discussed in lecture notes.
We explore a computational model of an incompressible fluid with a multi-phase field in three-dimensional Euclidean space. By investigating an incompressible fluid with a two-phase field geometrically, we reformulate the expression of the surface tension for the two-phase field found by Lafaurie, Nardone, Scardovelli, …
The stability of money value is an important requisite for a functioning economy, yet it critically depends on the actions of participants in the market themselves. Here we model the value of money as a dynamical variable that results from trading between agents. The basic trading scenario can be recast into an Ising t…
Formula calculates Riemann-Roch number for singular symplectic quotients.
Observer learns optimal policy from learner's actions without rewards.
Maximizes Rényi entropy for efficient exploration in reward-free RL.
New algorithm tackles nonstationary linear bandits with latent dynamics.
Safe Gaussian Process Bandit Optimization with sub-linear regret bounds.
New method uses image registration to recover complex signals from amplitude data.
Improved regret bounds for bandit phase retrieval.
In this paper, we investigate cost-aware joint learning and optimization for multi-channel opportunistic spectrum access in a cognitive radio system. We investigate a discrete time model where the time axis is partitioned into frames. Each frame consists of a sensing phase, followed by a transmission phase. During the …
Reinforcement learning (RL) in discrete action space is ubiquitous in real-world applications, but its complexity grows exponentially with the action-space dimension, making it challenging to apply existing on-policy gradient based deep RL algorithms efficiently. To effectively operate in multidimensional discrete acti…
The Lie algebroids are generalization of the Lie algebras. They arise, in particular, as a mathematical tool in investigations of dynamical systems with the first class constraints. Here we consider canonical symmetries of Hamiltonian systems generated by a special class of Lie algebroids. The ``coordinate part'' of th…
In this paper we propose new algorithm to reduce autocorrelation in Markov chain Monte-Carlo algorithms for euclidean field theories on the lattice. Our proposing algorithm is the Hybrid Monte-Carlo algorithm (HMC) with restricted Boltzmann machine. We examine the validity of the algorithm by employing the phi-fourth t…
Theory explains how deep nets learn features from data.
We consider the classical stochastic multi-armed bandit problem with a constraint that limits the total cost incurred by switching between actions to be no larger than a given switching budget. For this problem, we prove matching upper and lower bounds on the optimal (i.e., minimax) regret, and provide efficient rate-o…
New algorithm for reward-free RL with linear function approximation, reducing sample complexity.
PHASE dataset simulates complex social interactions in physical environments.
Robust algorithm optimizes corrupted Gaussian process bandits.
This paper focuses on energy management in buildings with phase change material (PCM), which is primarily used to improve thermal performance, but can also serve as an energy storage system. In this setting, optimal scheduling of an HVAC system is challenging because of the nonlinear and non-convex characteristics of t…
An artificial stock market is established based on multi-agent . Each agent has a limit memory of the history of stock price, and will choose an action according to his memory and trading strategy. The trading strategy of each agent evolves ceaselessly as a result of self-teaching mechanism. Simulation results exhibit …
New method deforms function algebras on manifolds using spectral decomposition.
There are two variants of the classical multi-armed bandit (MAB) problem that have received considerable attention from machine learning researchers in recent years: contextual bandits and simple regret minimization. Contextual bandits are a sub-class of MABs where, at every time step, the learner has access to side in…
In this paper we analyze the obstructions to the existence of global action-angle variables for regular non-commutative integrable systems (NCI systems) on Poisson manifolds. In contrast with local action-angle variables, which exist as soon as the fibers of the momentum map of such an integrable system are compact, gl…
Reward hacking exploits misspecified rewards, affecting agent capabilities and true performance.
In 1974, Berezin proposed a quantum theory for dynamical systems having a Kähler manifold as their phase space. The system states were represented by holomorphic functions on the manifold. For any homogeneous Kähler manifold, the Lie algebra of its group of motions may be represented either by holomorphic differential …
The paper tackles fair sequential decision making with biased linear bandit feedback.
We study the problem of designing AI agents that can robustly cooperate with people in human-machine partnerships. Our work is inspired by real-life scenarios in which an AI agent, e.g., a virtual assistant, has to cooperate with new users after its deployment. We model this problem via a parametric MDP framework where…
Reinforcement learning is considered to be a strong AI paradigm which can be used to teach machines through interaction with the environment and learning from their mistakes, but it has not yet been successfully used for automotive applications. There has recently been a revival of interest in the topic, however, drive…
The paper analyzes the current state of the world economy and offers a short-term forecast of its development. Our analysis of log-periodic oscillations in the DJIA dynamics suggests that in the second half of 2017 the United States and other more developed countries could experience a new recession, due to the third p…
In this paper we study the global behavior of the Ricci flow equation for two classes of homogeneous manifolds with two isotropy summands. Using methods of the qualitative theory of differential equations, we present the global phase portrait of such systems and derive some geometrical consequences on the structure of …
We prove a theorem on singular symplectic cotangent bundle reduction in the Fréchet setting and apply it to Yang-Mills-Higgs theory with special emphasis on the Higgs sector of the Glashow-Weinberg-Salam model. For the latter model we give a detailed description of the reduced phase space and show that the singular str…
Foliate systems are those which preserve some (possibly singular) foliation of phase space, such as systems with integrals, systems with continuous symmetries, and skew product systems. We study numerical integrators which also preserve the foliation. The case in which the foliation is given by the orbits of an action …
We extend the AKSZ formulation of the Poisson sigma model to more general target spaces, and we develop the general theory of graded geometry for poly-symplectic and poly-Poisson structures. In particular we prove a Schwarz-type theorem and transgression for graded poly-symplectic structures, recovering the action func…
A new method learns to stop with minimal data, outperforming traditional approaches.
Reduces equations for contact mechanical systems on Lie groups by exploiting symmetries.
Paper uses deep reinforcement learning for better control of rocket engines during start-up phases.
Neural Architecture Search (NAS) has emerged as a promising technique for automatic neural network design. However, existing MCTS based NAS approaches often utilize manually designed action space, which is not directly related to the performance metric to be optimized (e.g., accuracy), leading to sample-inefficient exp…
Quantum CNNs can be efficiently simulated classically on simple datasets.
New algorithm learns policies without explicit rewards for MDPs.
Improved formulation of spinfoam quantum gravity with cosmological constant, ensuring all amplitudes are finite and providing semiclassical asymptotics.
We discuss a class of (local and non-local) theories of gravity that share same properties: i) they admit the Einstein spacetime with arbitrary cosmological constant as a solution; ii) the on-shell action of such a theory vanishes and iii) any (cosmological or black hole) horizon in the Einstein spacetime with a positi…
The use of Reinforcement Learning (RL) is still restricted to simulation or to enhance human-operated systems through recommendations. Real-world environments (e.g. industrial robots or power grids) are generally designed with safety constraints in mind implemented in the shape of valid actions masks or contingency con…
A Poisson realization of the simple real Lie algebra on the phase space of each -Kepler problem is exhibited. As a consequence one obtains the Laplace-Runge-Lenz vector for each classical -Kepler problem. The verification of these Poisson realizations is greatly s…