This work proves win rate is key to understanding preference learning.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Calculates winning probability for three candidates based on support rates and information timing.
New bid shading algorithm reduces costs by 55%.
Paper confirms winning tickets have sharp minima, useful for generalization.
Faster WIND accelerates iterative BOND for LLM alignment.
The Lucas critique has exposed the problem of the trade-off between changes in monetary policy and structural breaks in economic time series. The search for and characterisation of such breaks has been a major econometric task ever since. We have developed an integral technique similar to CUSUM using an empirical model…
Improved MACD trading strategies with other indicators for better performance.
The success of lottery ticket initializations (Frankle and Carbin, 2019) suggests that small, sparsified networks can be trained so long as the network is initialized appropriately. Unfortunately, finding these "winning ticket" initializations is computationally expensive. One potential solution is to reuse the same wi…
DR-MCTS improves decision quality and sample efficiency in complex environments.
The lottery ticket hypothesis finds multiple winning sub-networks in neural networks.
We find a remarkable agreement between the statistics of a randomly divided interval and the observed statistical patterns and distributions found in horse racing betting markets. We compare the distribution of implied winning odds, the average true winning probabilities, the implied odds conditional on a win, and the …
Adversarial policies beat superhuman Go AI systems.
E-LTH finds winning tickets scalable across different network architectures.
The lottery ticket hypothesis proposes that over-parameterization of deep neural networks (DNNs) aids training by increasing the probability of a "lucky" sub-network initialization being present rather than by helping the optimization process (Frankle & Carbin, 2019). Intriguingly, this phenomenon suggests that initial…
In programmatic advertising, ad slots are usually sold using second-price (SP) auctions in real-time. The highest bidding advertiser wins but pays only the second-highest bid (known as the winning price). In SP, for a single item, the dominant strategy of each bidder is to bid the true value from the bidder's perspecti…
(Frankle & Carbin, 2019) shows that there exist winning tickets (small but critical subnetworks) for dense, randomly initialized networks, that can be trained alone to achieve comparable accuracies to the latter in a similar number of iterations. However, the identification of these winning tickets still requires the c…
Injecting human knowledge is an effective way to accelerate reinforcement learning (RL). However, these methods are underexplored. This paper presents our discovery that an abstract forward model (thought-game (TG)) combined with transfer learning (TL) is an effective way. We take StarCraft II as our study environment.…
Greedy pruning reduces neural networks by a logarithmic number of tickets, improving accuracy.
In-game win probability models, which provide a sports team's likelihood of winning at each point in a game based on historical observations, are becoming increasingly popular. In baseball, basketball and American football, they have become important tools to enhance fan experience, to evaluate in-game decision-making,…
We present a powerful new loss function and training scheme for learning binary hash functions. In particular, we demonstrate our method by creating for the first time a neural network that outperforms state-of-the-art Haar wavelets and color layout descriptors at the task of automated scene matching. By accurately rel…
Framework for real-time win probability and player ability in sports.
This paper is part of an ongoing investigation of "pragmatic information", defined in Weinberger (2002) as "the amount of information actually used in making a decision". Because a study of information rates led to the Noiseless and Noisy Coding Theorems, two of the most important results of Shannon's theory, we begin …
StarCraft II poses a grand challenge for reinforcement learning. The main difficulties of it include huge state and action space and a long-time horizon. In this paper, we investigate a hierarchical reinforcement learning approach for StarCraft II. The hierarchy involves two levels of abstraction. One is the macro-acti…
In the online multiple testing problem, p-values corresponding to different null hypotheses are observed one by one, and the decision of whether or not to reject the current hypothesis must be made immediately, after which the next p-value is observed. Alpha-investing algorithms to control the false discovery rate (FDR…
We explore a new way to evaluate generative models using insights from evaluation of competitive games between human players. We show experimentally that tournaments between generators and discriminators provide an effective way to evaluate generative models. We introduce two methods for summarizing tournament outcomes…
This paper discusses the gambling contest introduced in Seel & Strack (Gambling in contests, Discussion Paper Series of SFB/TR 15 Governance and the Efficiency of Economic Systems 375, Mar 2012.) and considers the impact of adding a penalty associated with failure to follow a winning strategy. The Seel & Strack model c…
The thesis explores IMP, a process that identifies winning tickets in DNNs, and its universality.
Paper tackles learning win-win solutions in aggregation systems.
Model analyzes RFQ markets using stochastic control to optimize dealer performance and inventory.
New method schedules learning rate without stopping time, outperforming existing methods.
TabPFN-2.5 boosts tabular AI performance, especially for large datasets.
Model selection for time series forecasting can be biased by the distribution of scores.
Gamblers lose in long bets despite casino claims, study shows.
New algorithm for identifying Condorcet team in noisy comparisons.
New algorithms minimize regret in repeated auctions by estimating values and optimizing bids.
The paper calculates gap distributions for translation surfaces, focusing on the double heptagon.
Different models to study the wealth distribution in an artificial society have considered a transactional dynamics as the driving force. Those models include a risk aversion factor, but also a finite probability of favoring the poorer agent in a transaction. Here we study the case where the partners in the transaction…
This paper studies the correlations of the average winnings of agents and the volatilities of systems based on mix-game model which is an extension of minority game (MG). In mix-game, there are two groups of agents; group1 plays the majority game, but the group2 plays the minority game. The results show that the correl…
The study analyzes games and social hierarchies, incorporating luck and depth of competition.
In online advertising, display ads are increasingly being placed based on real-time auctions where the advertiser who wins gets to serve the ad. This is called real-time bidding (RTB). In RTB, auctions have very tight time constraints on the order of 100ms. Therefore mechanisms for bidding intelligently such as clickth…
During the development of AlphaGo, its many hyper-parameters were tuned with Bayesian optimization multiple times. This automatic tuning process resulted in substantial improvements in playing strength. For example, prior to the match with Lee Sedol, we tuned the latest AlphaGo agent and this improved its win-rate from…
MOVDA improves skill ratings by considering margin of victory deviations.
Stochastic Gradient TreeBoost is often found in many winning solutions in public data science challenges. Unfortunately, the best performance requires extensive parameter tuning and can be prone to overfitting. We propose PaloBoost, a Stochastic Gradient TreeBoost model that uses novel regularization techniques to guar…
Elo rating outperforms complex models in skill estimation despite model misspecification.
Proposes a new financial model capturing winning and losing streaks.
Uniform AMMs control loss in prediction markets.
We rebias estimates to improve interval calibration and prediction accuracy.
We consider the problem of high-level strategy selection in the adversarial setting of real-time strategy games from a reinforcement learning perspective, where taking an action corresponds to switching to the respective strategy. Here, a good strategy successfully counters the opponent's current and possible future st…