Proposes DEXP3.M for unknown delay in multi-arm bandit with multiple play.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We discuss a multiple-play multi-armed bandit (MAB) problem in which several arms are selected at each round. Recently, Thompson sampling (TS), a randomized algorithm with a Bayesian spirit, has attracted much attention for its empirically excellent performance, and it is revealed to have an optimal regret bound in the…
We investigate the adversarial bandit problem with multiple plays under semi-bandit feedback. We introduce a highly efficient algorithm that asymptotically achieves the performance of the best switching -arm strategy with minimax optimal regret bounds. To construct our algorithm, we introduce a new expert advice alg…
We study a generalization of the multi-armed bandit problem with multiple plays where there is a cost associated with pulling each arm and the agent has a budget at each time that dictates how much she can expect to spend. We derive an asymptotic regret lower bound for any uniformly efficient algorithm in our setting. …
We study the multi-armed bandit problem with multiple plays and a budget constraint for both the stochastic and the adversarial setting. At each round, exactly out of possible arms have to be played (with ). In addition to observing the individual rewards for each arm played, the player also lea…
New algorithm for shareable arms with load-dependent rewards in stochastic bandits.
Proof of Donaldson's theorem using Seiberg-Witten equations for multiple spinors.
This paper proposes using a linear function approximator, rather than a deep neural network (DNN), to bias a Monte Carlo tree search (MCTS) player for general games. This is unlikely to match the potential raw playing strength of DNNs, but has advantages in terms of generality, interpretability and resources (time and …
A new bandit algorithm for web page item display.
We prove several results on Almgren's multiple valued functions and their links to integral currents. In particular, we give a simple proof of the fact that a Lipschitz multiple valued map naturally defines an integer rectifiable current; we derive explicit formulae for the boundary, the mass and the first variations a…
Investigates the fundamental components of attention mechanisms.
Polynomial algorithm for multiplication on one-hole torus skein algebra.
We study an extension of the classic stochastic multi-armed bandit problem which involves multiple plays and Markovian rewards in the rested bandits setting. In order to tackle this problem we consider an adaptive allocation rule which at each stage combines the information from the sample means of all the arms, with t…
Unified framework denoises data and abstains from uncertain predictions.
We analyze a notion of multiple valued sections of a vector bundle over an abstract smooth Riemannian manifold, which was suggested by W. Allard in the unpublished note "Some useful techniques for dealing with multiple valued functions" and generalizes Almgren's -valued functions. We study some relevant properties o…
Adversarial self-play in two-player games has delivered impressive results when used with reinforcement learning algorithms that combine deep neural networks and tree search. Algorithms like AlphaZero and Expert Iteration learn tabula-rasa, producing highly informative training data on the fly. However, the self-play t…
In this paper we analyze Gresham's Law, in particular, how the rate of inflow or outflow of currencies is affected by the demand elasticity of arbitrage and the difference in face value ratios inside and outside of a country under a bimetallic system. We find that these equations are very similar to those used to descr…
During the development of AlphaGo, its many hyper-parameters were tuned with Bayesian optimization multiple times. This automatic tuning process resulted in substantial improvements in playing strength. For example, prior to the match with Lee Sedol, we tuned the latest AlphaGo agent and this improved its win-rate from…
Optimizes financial decisions with illiquid assets using Kelly criterion.
FGNNs improve game-playing AI by exploiting symmetries.
This paper relaxes the common prior assumption in the public and private information game of Morris and Shin (2000, 2004). For the generalized game, where the agent's prior expectations are heterogenous, it derives a sharp condition for the emergence of unique/multiple equilibria. This condition indicates that unique e…
In a series of papers, including the present one, we give a new, shorter proof of Almgren's partial regularity theorem for area minimizing currents in a Riemannian manifold, with a slight improvement on the regularity assumption for the latter. This note establishes a new a priori estimate on the excess measure of an a…
The banking systems that deal with risk management depend on underlying risk measures. Following the Basel II accord, there are two separate methods by which banks may determine their capital requirement. The Value at Risk measure plays an important role in computing the capital for both approaches. In this paper we an…
A fundamental operation in many vision tasks, including motion understanding, stereopsis, visual odometry, or invariant recognition, is establishing correspondences between images or between images and data from other modalities. We present an analysis of the role that multiplicative interactions play in learning such …
Persona2vec learns multiple node roles in graphs.
This work introduces a method to learn dynamical systems from noisy sensor measurements using multiple shooting.
The classification work [5], [9] left unsettled only those anomalous isoparametric hypersurfaces with four principal curvatures and multiplicity pair or in the sphere. By systematically exploring the ideal theory in commutative algebra in conjunction with the geometry of isoparametric hypers…
We consider off-policy policy evaluation when the trajectory data are generated by multiple behavior policies. Recent work has shown the key role played by the state or state-action stationary distribution corrections in the infinite horizon context for off-policy policy evaluation. We propose estimated mixture policy …
BiHRNN predicts inflation by leveraging hierarchical structure and bidirectional RNNs.
We propose a deep neural network-based algorithm to identify the Markovian Nash equilibrium of general large -player stochastic differential games. Following the idea of fictitious play, we recast the -player game into decoupled decision problems (one for each player) and solve them iteratively. The individua…
JAMPI improves matrix multiplication in Spark, boosting performance by up to 24%.
Novel method uses information theory to measure causal influences during transient neural events.
Memory is an important aspect of intelligence and plays a role in many deep reinforcement learning models. However, little progress has been made in understanding when specific memory systems help more than others and how well they generalize. The field also has yet to see a prevalent consistent and rigorous approach f…
The study examines how hyperparameters affect prediction discrepancies in machine learning models.
Fictitious play is a simple and widely studied adaptive heuristic for playing repeated games. It is well known that fictitious play fails to be Hannan consistent. Several variants of fictitious play including regret matching, generalized regret matching and smooth fictitious play, are known to be Hannan consistent. In …
Deep learning models designed for visual classification tasks on natural images have become prevalent in medical image analysis. However, medical images differ from typical natural images in many ways, such as significantly higher resolutions and smaller regions of interest. Moreover, both the global structure and loca…
Despite the fact that nonlinear subspace learning techniques (e.g. manifold learning) have successfully applied to data representation, there is still room for improvement in explainability (explicit mapping), generalization (out-of-samples), and cost-effectiveness (linearization). To this end, a novel linearized subsp…
Factorial moments are convenient tools in nuclear physics to characterize the multiplicity distributions when phase-space resolution () becomes small. For uncorrelated particle production within , Gaussian statistics holds and factorial moments are equal to unity for all orders . Correlations between par…
Self-play is an unsupervised training procedure which enables the reinforcement learning agents to explore the environment without requiring any external rewards. We augment the self-play setting by providing an external memory where the agent can store experience from the previous tasks. This enables the agent to come…
Optimizes mobile notifications for multiple objectives using reinforcement learning.
The Normal Means problem plays a fundamental role in many areas of modern high-dimensional statistics, both in theory and practice. And the Empirical Bayes (EB) approach to solving this problem has been shown to be highly effective, again both in theory and practice. However, almost all EB treatments of the Normal Mean…
Grimaldi-Pansu metrics are constructed for manifolds with multiple ends.
Enhanced feedback model improves sample-efficiency in POMDPs.
A new framework for controllable generation of discrete masked models.
The standard theory of coherent risk measures fails to consider individual institutions as part of a system which might itself experience instability and spread new sources of risk to the market participants. In compliance with an approach adopted by Shapley and Shubik (1969), this paper proposes a cooperative market g…
Improved multi-task averaging reduces mean squared error in high-dimensional data.
Kernel methods are among the most popular techniques in machine learning. From a frequentist/discriminative perspective they play a central role in regularization theory as they provide a natural choice for the hypotheses space and the regularization functional through the notion of reproducing kernel Hilbert spaces. F…
Recent progress in artificial intelligence through reinforcement learning (RL) has shown great success on increasingly complex single-agent environments and two-player turn-based games. However, the real-world contains multiple agents, each learning and acting independently to cooperate and compete with other agents, a…