Deep RL learns to play Pong from frames alone.
problem Scaling up RL problems leads to computational bottlenecks.
method End-to-end DRL approach using ANN and Policy Gradients.
result Successfully learned to play Pong from frames.
Improved DRL performance with novel pre-training method.
problem Data inefficiency in DRL algorithms.
method Jointly pre-training with supervised, autoencoder, and value losses.
result Significantly improved learning performance in Atari games.
Study ping-pong dynamics in hyperbolic-like groups with non-simple points.
problem Investigate the ping-pong dynamics of hyperbolic-like groups.
method Explicitly provide a proper ping-pong partition for any pair of non-cyclic point stabilizers.
result Existence of a proper ping-pong partition for any pair of non-cyclic point stabilizers.
New algebra pong algebra computed for knot Floer homology.
problem Computing A-infinity structure on knot Floer homology.
method Introduced differential graded algebra, pong algebra.
result Computed A-infinity structure on pong algebra's homology.
While deep reinforcement learning has successfully solved many challenging control tasks, its real-world applicability has been limited by the inability to ensure the safety of learned policies. We propose an approach to verifiable reinforcement learning by training decision tree policies, which can represent complex p…
Study estimates gaps in semigroup products, proving embedding properties.
problem Estimating singular value gaps in semigroup products.
method Lower estimates for singular value gaps of free products of semigroups in ping-pong position.
result Groups generated by semigroups in ping-pong position are quasi-isometrically embedded.
New groups discovered with unique properties in a specific space.
problem Finding new discrete subgroups with special properties in a mathematical space.
method Proved by showing groups play ping-pong on cones, related to crooked surfaces.
result Infinite family of discrete subgroups with remarkable properties in Sp4(R). Geometric group theory explores groups through their geometric properties.
problem Understanding groups via geometric properties.
method Cayley and Schreier graphs, ping-pong lemma, quasi-isometries, growth of groups, hyperbolicity.
result Gromov's theorem on groups of polynomial growth and amenability.
Paper introduces a technique to simplify RNN policies for better understanding and analysis.
problem Difficulty in explaining and analyzing RNN policies due to continuous-valued memory vectors and observation features.
method Quantized Bottleneck Insertion technique to learn finite representations of RNN vectors and features.
result Finite representations of RNN policies can be as small as 3 discrete memory states and 10 observations, improving interpretability.
Improved off-policy reinforcement learning by discounting and soft normalization.
problem Divergence issues in off-policy reinforcement learning.
method Introducing a discount factor and a soft normalization penalty into COP-TD.
result Discounted COP-TD is better behaved both theoretically and empirically.
Study weightings from singular Lie filtrations.
problem Generalize constructions for singular Lie filtrations.
method Study weightings arising from singular Lie filtrations.
result Generalizes constructions for (regular) Lie filtrations.
Algorithm determines discrete, free subgroups of SL2 over non-archimedean fields.
problem Identifying discrete, free subgroups of SL2 over non-archimedean fields.
method Ping Pong Lemma applied to Bruhat-Tits tree action.
result Algorithm determines if subgroup is discrete and free of rank two.
The study examines subgroups of torus mapping class group generated by Dehn twists powers.
problem Characterizing subgroups generated by powers of Dehn twists.
method Using the ping pong lemma and geometric intersection numbers.
result Subgroups can be free groups, direct products, or have specific ranks.
We prove that all atoroidal automorphisms of Out(FN) act on the space of projectivized geodesic currents with generalized north-south dynamics. As an application, we produce new examples of non virtually cyclic, free and purely atoroidal subgroups of Out(FN) such that the corresponding free group extension is hyp…
In this paper, we prove a quantitative version of the Tits alternative for negatively pinched manifolds X. Precisely, we prove that a nonelementary discrete isometry subgroup of Isom(X) generated by two non-elliptic isometries g, f contains a free subgroup of rank 2 generated by isometries fN,h …
We prove that if φ,ψ∈Out(FN) are hyperbolic iwips (irreducible with irreducible powers) such that <φ,ψ>≤Out(FN) is not virtually cyclic then some high powers of φ and ψ generate a free subgroup of rank two, all of whose nontrivial elements are again hyperbolic iwips. Being a hyperbolic iwip element of $…
Local-to-global principle for Morse actions on symmetric spaces.
problem Recognizing Morse actions on symmetric spaces.
method Equivariant Morse quasiisometric embeddings of trees into symmetric spaces.
result Algorithmic recognizability of Morse actions and construction of Morse Schottky subgroups.
We prove that every acylindrically hyperbolic group that has no non-trivial finite normal subgroup satisfies a strong ping pong property, the Pnaive property: for any finite collection of elements h1,…,hk, there exists another element γ=1 such that for all i, $\langle h_i, γ\rangle = \langle h_i …
The paper models and deforms A-infinity structures for bordered knot algebras.
problem Understanding A-infinity structures for bordered knot algebras.
method Combinatorial model and weighted deformation of A-infinity structures.
result Explicit combinatorial model for bordered knot algebras' A-infinity structure.
Researchers discover all affinely homogeneous models for surfaces in 4D space.
problem Identifying all affinely homogeneous models for surfaces in 4D space.
method Improved power series method of equivalence, capturing invariants at the origin, creating branches, and infinitesimalizing calculations.
result Find several inequivalent terminal branches yielding each to some nonempty moduli space of homogeneous models.
Defines Wodzicki residue using groupoids and fibered distributions.
problem Defining and understanding the Wodzicki residue in noncommutative geometry.
method Using groupoid language and filtered manifolds, defining the residue and showing its properties.
result The groupoidal residue is a trace on pseudodifferential operators and matches the usual residue in certain cases.
WMPG reduces policy gradient variance using world models.
problem Reducing variance in policy gradient estimates.
method Trains a world model online to estimate policy gradients and uses imagined trajectories as a baseline.
result WMPG achieves better sample efficiency compared to AC and MAC.
Visual analogies help transfer knowledge between Atari games.
problem Can visual analogies transfer knowledge between Atari games?
method Created visual analogies between pairs of Atari games and used them to train policies for one game using data from another.
result Visual analogies can be used to transfer knowledge between Atari games.
Paper presents content-based models for game recommendation in cold start scenarios.
problem Cold start problem in game recommendation where new games and players have no historical data.
method Uses survey data to develop content-based interaction models that generalize to new games, players, and both.
result Content models outperform collaborative filtering in predicting new interactions.
Potential games, originally introduced in the early 1990's by Lloyd Shapley, the 2012 Nobel Laureate in Economics, and his colleague Dov Monderer, are a very important class of models in game theory. They have special properties such as the existence of Nash equilibria in pure strategies. This note introduces graphical…
A new framework for playing and learning board games.
problem Tackling the tedious and repetitive aspects of coding for board game AI.
method Developed a generic TD(λ)-n-tuple agent for arbitrary board games. result TD(λ)-n-tuple outperforms other generic agents on various games. IGGP learns game rules from varying quality game play, finding no overall trend.
problem Learn game rules from varying quality game play.
method Used Sancho's intelligent game traces and ILP systems (Metagol, Aleph, ILASP) to induce game rules from traces of varying quality and volume.
result No overall trend in accuracy of learned game rules from varying quality and volume of training data.
This note removes technical assumptions and characterizes relatively dominated representations.
problem Geometrically finiteness and Anosov conditions in higher-rank settings.
method Characterization using eigenvalue gaps and limit maps.
result Relatively dominated representations are characterized using eigenvalue gaps and limit maps.
New game introduces linking-unlinking strategy for two-component links.
problem Tackling the linking and unlinking of two-component links.
method Introducing and analyzing the Linking-Unlinking Game on various link shadows.
result Winning strategies for specific link shadows are presented.
A game on diagrams switches crossing directions to achieve connectedness.
problem Achieving connectedness in diagrams through crossing switches.
method Players switch crossing directions on regions of a diagram to achieve connectedness.
result Connectedness can be achieved through strategic crossing switches.
The paper introduces MDP homomorphic networks for faster reinforcement learning.
problem Current reinforcement learning approaches do not exploit symmetries in the joint state-action space.
method Equivariant neural networks with group-structured symmetries (reflections, rotations).
result MDP homomorphic networks converge faster than unstructured baselines on various tasks.
Introduces SM-games to analyze machine learning interactions.
problem Lack of understanding and control in n-player games.
method Introduces SM-games with pairwise zero-sum interactions.
result SM-games are amenable to first-order optimization methods.
Just as war is sometimes fallaciously represented as a zero sum game -- when in fact war is a negative sum game - stock market trading, a positive sum game over time, is often erroneously represented as a zero sum game. This is called the "zero sum fallacy" -- the erroneous belief that one trader in a stock market exch…
The existence of stationary Markov perfect equilibria in stochastic games is shown under a general condition called "(decomposable) coarser transition kernels". This result covers various earlier existence results on correlated equilibria, noisy stochastic games, stochastic games with finite actions and state-independe…
TextWorld is a Python library for RL agents in text-based games.
problem Training RL agents on text-based games with varying challenges and sparse rewards.
method Developed a Python library with backend functions for state tracking and reward assignment. Enables users to create new games with precise control over difficulty and scope.
result Demonstrated the effectiveness of TextWorld in training RL agents on a curated list of games and generated sets of games.
Game theory helps analyze ESOs/EBIs in production and service sectors.
problem Economic incentives affect traditional production/service functions and create intangible capital.
method Uses game theory to analyze interactions in ESO/EBI transactions.
result No perfect Nash Equilibria for two-stage games involving many participants.
We start briefly surveying research on optimal stopping games since their introduction by E.B.Dynkin more than 40 years ago. Recent renewed interest to dynkin's games is due, in particular, to the study of Israeli (game) options introduced in 2000. We discuss the work on these options and related derivative securities …
The paper explores how regularization can lead to convergence in imperfect information games.
problem Finding equilibrium in imperfect information games with imperfect information.
method Investigates Follow the Regularized Leader dynamics and how adding a regularization term can lead to strong convergence guarantees.
result The approach leads to algorithms that converge exactly to the Nash equilibrium in imperfect information games.
Educational game on crypto investment helps students grasp macroeconomics.
problem Weak connections between microeconomic decision-making and macroeconomic concepts in classroom games.
method Design and study of an educational game on cryptocurrency investment.
result Engages students in understanding macroeconomics through incentivized individual investment decisions.
Simplified NFT games discussed with methods for extracting value.
problem Issues influencing NFT games' structure and stability.
method Three methods for extracting value from NFT games.
result Various design constraints and mutual beneficial games.
The paper proposes a method to learn continuous-action graphical games from perturbed equilibria.
problem Learning the exact structure of continuous-action graphical games from limited data.
method A ℓ12− block regularized method to recover the graphical game structure. result The method recovers the exact structure of the graphical game under certain conditions.
Gradient Descent Ascent converges to von-Neumann solution in hidden zero-sum games.
problem Understanding dynamics of zero-sum games with hidden structure.
method Gradient Descent Ascent applied to hidden zero-sum games with specific convex-concave structure.
result Gradient Descent Ascent converges to von-Neumann solution in strictly convex-concave hidden games.
Paper tackles hidden game problem in AI alignment and language games.
problem Hidden game problem in AI alignment and language games.
method Developed a composition of regret minimization techniques to discover and exploit hidden structures.
result Achieved optimal external and swap regret bounds for rapid convergence to correlated equilibria.
AEC Games model represents software MARL environments better than POSGs.
problem POSGs are conceptually unsuitable for software MARL environments.
method Introduced AEC Games model as an equivalent to POSGs.
result AEC Games model is more representative of software MARL environments.
We introduce CSE for MLSF games and devise online learning algorithms for achieving no-external Stackelberg-regret.
problem Learning equilibrium in leader-follower games with noisy bandit feedback.
method Proposed Correlated Stackelberg Equilibrium (CSE) and online learning algorithms balancing exploration and exploitation.
result Achieves no-external Stackelberg-regret, converging to approximate CSE.
Generalizes region select game to k-colored knot diagrams.
problem Play game on knot diagrams with multiple colors.
method Generalize region select game to k-colored knot diagrams. result Generalization of the region select game to k-colored knot diagrams. Deep Reinforcement Learning automates match-3 game testing.
problem Reducing human effort in testing match-3 video games.
method Dueling Deep Q-Network paradigm applied to Jelly Juice game.
result The network outperforms random player and adapts to game difficulty.
Unified framework for Bayesian and Frequentist statistics.
problem Embedding Bayesian statistics within a broader decision-making framework.
method Game theory and statistical analysis.
result Statistical games unify Bayesian and Frequentist statistics.