Paper tackles hidden game problem in AI alignment and language games.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
KG-A2C agent learns natural language IF games by reasoning and constraining action spaces.
The paper explores game-theoretic alignment of LLMs with human preferences, finding limitations and conditions.
Graphs help agents learn emergent communication.
While Reinforcement Learning (RL) approaches lead to significant achievements in a variety of areas in recent history, natural language tasks remained mostly unaffected, due to the compositional and combinatorial nature that makes them notoriously hard to optimize. With the emerging field of Text-Based Games (TBGs), re…
We introduce TextWorld, a sandbox learning environment for the training and evaluation of RL agents on text-based games. TextWorld is a Python library that handles interactive play-through of text games, as well as backend functions like state tracking and reward assignment. It comes with a curated list of games whose …
Deep CapsNet improves sign language recognition from wearable IMUs.
New method for evaluating LLMs reduces bias in open-ended evaluations.
There has been an increasing interest in the area of emergent communication between agents which learn to play referential signalling games with realistic images. In this work, we consider the signalling game setting of Havrylov and Titov and investigate the effect of the feature extractor's weights and of the task bei…
Deep learning models generate languages that lack abstract reasoning.
Graph neural networks help AI agents learn more complex language.
The General Video Game AI (GVGAI) competition and its associated software framework provides a way of benchmarking AI algorithms on a large number of games written in a domain-specific description language. While the competition has seen plenty of interest, it has so far focused on online planning, providing a forward …
The paper investigates how supervised learning and self-play improve sample efficiency in teaching AI to communicate.
This work simplifies data valuation for LLMs using Shapley value computation.
Scaling up model and data size improves imitation learning in single-agent games.
New algorithms achieve logarithmic regret in KL-regularized Markov games.
Understanding procedural text requires tracking entities, actions and effects as the narrative unfolds. We focus on the challenging real-world problem of action-graph extraction from material science papers, where language is highly specialized and data annotation is expensive and scarce. We propose a novel approach, T…
A new framework quantifies how model explanations influence each other.
In this work we present a technique to use natural language to help reinforcement learning generalize to unseen environments. This technique uses neural machine translation, specifically the use of encoder-decoder networks, to learn associations between natural language behavior descriptions and state-action informatio…
To be successful in real-world tasks, Reinforcement Learning (RL) needs to exploit the compositional, relational, and hierarchical structure of the world, and learn to transfer it to the task at hand. Recent advances in representation learning for language make it possible to build models that acquire world knowledge f…
BIG-bench benchmarks language models, revealing their strengths and weaknesses.
BED-LLM uses Bayesian experimental design to improve LLMs' information gathering.
SPPO optimizes language model alignment by treating preferences as a game and achieving state-of-the-art performance.
Recent reinforcement learning (RL) approaches have shown strong performance in complex domains such as Atari games, but are often highly sample inefficient. A common approach to reduce interaction time with the environment is to use reward shaping, which involves carefully designing reward functions that provide the ag…
MAXMINLCB optimizes unknown target functions with preference feedback using a Stackelberg game approach.
Q*BERT learns to navigate text-based games by building a knowledge graph.
CausalGame benchmarks LLM agents' causal thinking in games.
Interpersonal relations are fickle, with close friendships often dissolving into enmity. In this work, we explore linguistic cues that presage such transitions by studying dyadic interactions in an online strategy game where players form alliances and break those alliances through betrayal. We characterize friendships …
Dynamic meta-learning improves multi-agent communication with natural language.
SLHF uses sequential game theory to optimize preferences from human feedback.
Study combines chit-chat and goal-oriented dialogue in fantasy games.
Some optimization or equilibrium problems involving somehow the concept of optimal transport are presented in these notes, mainly devoted to applications to economic and game theory settings. A variant model of transport, taking into account traffic congestion effects is the first topic, and it shows various links with…
We consider the issue of multiple agents learning to communicate through reinforcement learning within partially observable environments, with a focus on information asymmetry in the second part of our work. We provide a review of the recent algorithms developed to improve the agents' policy by allowing the sharing of …
Assessing the impact of the individual actions performed by soccer players during games is a crucial aspect of the player recruitment process. Unfortunately, most traditional metrics fall short in addressing this task as they either focus on rare actions like shots and goals alone or fail to account for the context in …
PropFair algorithm ensures fair performance in federated learning.
We study the problem of interpreting trained classification models in the setting of linguistic data sets. Leveraging a parse tree, we propose to assign least-squares based importance scores to each word of an instance by exploiting syntactic constituency structure. We establish an axiomatic characterization of these i…
Deep neural networks have excelled on a wide range of problems, from vision to language and game playing. Neural networks very gradually incorporate information into weights as they process data, requiring very low learning rates. If the training distribution shifts, the network is slow to adapt, and when it does adapt…
Standard economic theory makes an allowance for the agency problem, but not the compounding of moral hazard in the presence of informational opacity, particularly in what concerns high-impact events in fat tailed domains (under slow convergence for the law of large numbers). Nor did it look at exposure as a filter that…
Improved language models learn complex distributions using Fourier series.
We examine the effects of instantiating Lewis signaling games within a population of speaker and listener agents with the aim of producing a set of general and robust representations of unstructured pixel data. Preliminary experiments suggest that the set of representations associated with languages generated within a …
A novel memory mechanism for reinforcement learning agents that stores past events in human-readable language.
This research integrates attention into XAI frameworks for better model explanations.
NAMEx merges experts using Nash bargaining for improved performance.
NetHack Learning Environment (NLE) tests RL algorithms, offering scalable, complex, and challenging gameplay.
Paper presents content-based models for game recommendation in cold start scenarios.
In this work, we ask the following question: Can visual analogies, learned in an unsupervised way, be used in order to transfer knowledge between pairs of games and even play one game using an agent trained for another game? We attempt to answer this research question by creating visual analogies between a pair of game…
Potential games, originally introduced in the early 1990's by Lloyd Shapley, the 2012 Nobel Laureate in Economics, and his colleague Dov Monderer, are a very important class of models in game theory. They have special properties such as the existence of Nash equilibria in pure strategies. This note introduces graphical…
IGGP learns game rules from varying quality game play, finding no overall trend.