New algorithm for multi-player bandits with collision-dependent rewards.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Algorithm reduces regret in multi-player bandits with unknown collision rewards.
A new algorithm RESYNC for defenders against malicious attackers in multi-player bandits.
The paper addresses the Multiplayer Multi-Armed Bandit (MMAB) problem, where decision makers or players collaborate to maximize their cumulative reward. When several players select the same arm, a collision occurs and no reward is collected on this arm. Players involved in a collision are informed about this collis…
New algorithm for multi-player bandits without needing lower bounds or scaling inversely.
We consider a fully decentralized multi-player stochastic multi-armed bandit setting where the players cannot communicate with each other and can observe only their own actions and rewards. The environment may appear differently to different players, , the reward distributions for a given arm are heterog…
We study multiplayer stochastic multi-armed bandit problems in which the players cannot communicate and if two or more players pull the same arm, a collision occurs and the involved players receive zero reward. We consider two feedback models: a model in which the players can observe whether a collision has occurred an…
The decentralized stochastic multi-player multi-armed bandit (MP-MAB) problem, where the collision information is not available to the players, is studied in this paper. Building on the seminal work of Boursier and Perchet (2019), we propose error correction synchronization involving communication (EC-SIC), whose regre…
Motivated by cognitive radios, stochastic multi-player multi-armed bandits gained a lot of interest recently. In this class of problems, several players simultaneously pull arms and encounter a collision - with 0 reward - if some of them pull the same arm at the same time. While the cooperative case where players maxim…
A multi-user multi-armed bandit (MAB) framework is used to develop algorithms for uncoordinated spectrum access. The number of users is assumed to be unknown to each user. A stochastic setting is first considered, where the rewards on a channel are the same for each user. In contrast to prior work, it is assumed that t…
A multi-player bandit system resists adversarial attacks with near-optimal regret.
This work examines the role of reinforcement learning in reducing the severity of on-road collisions by controlling velocity and steering in situations in which contact is imminent. We construct a model, given camera images as input, that is capable of learning and predicting the dynamics of obstacles, cars and pedestr…
No communication allows optimal instance-dependent regret guarantees in multi-player bandits.
We study a multiplayer stochastic multi-armed bandit problem in which players cannot communicate, and if two or more players pull the same arm, a collision occurs and the involved players receive zero reward. We consider the challenging heterogeneous setting, in which different arms may have different means for differe…
We consider the stochastic multi-armed bandit (MAB) problem in a setting where a player can pay to pre-observe arm rewards before playing an arm in each round. Apart from the usual trade-off between exploring new arms to find the best one and exploiting the arm believed to offer the highest reward, we encounter an addi…
Residual neural networks improve collision prediction in planetary simulations.
Study nonholonomic systems with collisions using variational principles.
Paper analyzes dynamics of nonholonomic systems with collisions using variational techniques.
Model forecasts motor vehicle collision rates with high accuracy.
Machine learning competition predicts spacecraft collision risks.
New algorithms estimate and test collision probability with near-optimal sample complexity.
One approach to designing decision making logic for an aircraft collision avoidance system frames the problem as a Markov decision process and optimizes the system using dynamic programming. The resulting collision avoidance strategy can be represented as a numeric table. This methodology has been used in the developme…
No-collision maps improve manifold learning for image data.
Reduces necessary conditions for collision avoidance on curved spaces.
A new metric for uncertainty quantification using class collisions.
An important application of intelligent vehicles is advance detection of dangerous events such as collisions. This problem is framed as a problem of optimal alarm choice given predictive models for vehicle location and motion. Techniques for real-time collision detection are surveyed and grouped into three classes: ran…
New algorithm for multi-player bandits in decentralized, asynchronous systems.
Bayesian deep learning predicts satellite collisions.
We study the multi-player stochastic multiarmed bandit (MAB) problem in an abruptly changing environment. We consider a collision model in which a player receives reward at an arm if it is the only player to select the arm. We design two novel algorithms, namely, Round-Robin Sliding-Window Upper Confidence Bound\# (RR-…
New algorithms tackle adversarial multi-player bandits with forced-collision communication.
New strategy achieves optimal regret without communication or collisions in multi-player bandit.
New algorithm tackles multi-player bandit problems with limited access to arms.
A model used for velocity control during car following was proposed based on deep reinforcement learning (RL). To fulfil the multi-objectives of car following, a reward function reflecting driving safety, efficiency, and comfort was constructed. With the reward function, the RL agent learns to control vehicle speed in …
Study shows how transformers classify symbols without naming them, proving a margin-versus-collision criterion.
The configuration manifold of a mechanical system consisting of two unconstrained rigid bodies in , , is a manifold with boundary (typically with singularities.) A complete description of the system requires boundary conditions that specify how orbits should be continued after collisions. A b…
The Kepler-Heisenberg problem is that of determining the motion of a planet around a sun in the Heisenberg group, thought of as a three-dimensional sub-Riemannian manifold. The sub-Riemannian Hamiltonian provides the kinetic energy, and the gravitational potential is given by the fundamental solution to the sub-Laplaci…
Study motion planning for points avoiding obstacles in a plane.
Centrality, as a geometrical property of the collision, is crucial for the physical interpretation of nucleus-nucleus and proton-nucleus experimental data. However, it cannot be directly accessed in event-by-event data analysis. Common methods for centrality estimation in A-A and p-A collisions usually rely on a single…
Recovering manifold geometry from geodesic intersections.
We consider the non-stochastic version of the (cooperative) multi-player multi-armed bandit problem. The model assumes no communication at all between the players, and furthermore when two (or more) players select the same action this results in a maximal loss. We prove the first -type regret guarantee for th…
Unified approach detects traffic conflicts across various interactions.
Rolling systems limit to billiard models with no-slip collisions.
Motion planning for robots of high degrees-of-freedom (DOFs) is an important problem in robotics with sampling-based methods in configuration space C as one popular solution. Recently, machine learning methods have been introduced into sampling-based motion planning methods, which train a classifier to distinguish coll…
Up to symmetries, the orbits of three equal masses under an inverse cube force with zero angular momentum and constant moment of inertia can be reparametrized as the geodesics of a complete, negatively curved metric on a pair of pants. The ends of the pants represent binary collisions. Here we will examine the visibili…
Stochastic approach improves neural network training for kinetic simulations.
Multipeakons are special solutions to the Camassa-Holm equation described by an integrable geodesic flow on a Riemannian manifold. We present a bi-Hamiltonian formulation of the system explicitly and write down formulae for the associated first integrals. Then we exploit the first integrals and present a novel approach…
Generalizes Landau-Ginzburg mirrors for Frobenius manifolds in Dynkin type A.
This work improves motion planning for quadcopters by learning and reasoning about controller performance.