Algorithm identifies Copeland winners in dueling bandits with ternary feedback.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We study the K-armed dueling bandit problem, a variation of the standard stochastic bandit problem where the feedback is limited to relative comparisons of a pair of arms. The hardness of recommending Copeland winners, the arms that beat the greatest number of other arms, is characterized by deriving an asymptotic regr…
Multi-armed bandit(MAB) problem is a reinforcement learning framework where an agent tries to maximise her profit by proper selection of actions through absolute feedback for each action. The dueling bandits problem is a variation of MAB problem in which an agent chooses a pair of actions and receives relative feedback…
In this paper, we propose a Double Thompson Sampling (D-TS) algorithm for dueling bandit problems. As indicated by its name, D-TS selects both the first and the second candidates according to Thompson Sampling. Specifically, D-TS maintains a posterior distribution for the preference matrix, and chooses the pair of arms…
A new method corrects for bias in selecting the best candidate.
New algorithm reduces dynamic regret in non-stationary dueling bandits using a weighted Borda score.
New method identifies Condorcet winner in dueling bandits with improved sample complexity.
This paper tackles combinatorial pure exploration for dueling bandits, aiming to find the best candidate-position match.
New method uses geometric properties for better density estimation.
A game-theoretic approach to multi-criteria ranking from ordinal data.
SIREN protocol corrects optimistic winner's scores in LLM evaluation.
Study shows big winner stocks significantly impact passive and active investment strategies.
This paper reexamines the profitability of loser, winner and contrarian portfolios in the Chinese stock market using monthly data of all stocks traded on the Shanghai Stock Exchange and Shenzhen Stock Exchange covering the period from January 1997 to December 2012. We find evidence of short-term and long-term contraria…
Data-driven decision-making often overestimates benefits due to the winner's curse.
TimeMCL forecasts diverse time series futures using neural networks and WTA loss.
Stochastic LWTA networks resist adversarial attacks while maintaining accuracy.
Paper proposes a method to predict MOBA game winners with calibrated confidence.
aMCL uses annealing to improve hypothesis diversity in ambiguous tasks.
New method optimizes treatment policies to avoid winner's curse.
This case study tests the possibility of prediction for "success" (or "winner") components of four stock & shares market indices in a time period of three years from 02-Jul-2009 to 29-Jun-2012.We compare their performance ain two time frames: initial frame three months at the beginning (02/06/2009-30/09/2009) and the f…
In our model, traders interact with each other and with a central bank; they are taxed on the money they make, some of which is dissipated away by corruption. A generic feature of our model is that the richest trader always wins by 'consuming' all the others: another is the existence of a threshold wealth, below wh…
Sentiment Analysis of microblog feeds has attracted considerable interest in recent times. Most of the current work focuses on tweet sentiment classification. But not much work has been done to explore how reliable the opinions of the mass (crowd wisdom) in social network microblogs such as twitter are in predicting ou…
Public debates are a common platform for presenting and juxtaposing diverging views on important issues. In this work we propose a methodology for tracking how ideas flow between participants throughout a debate. We use this approach in a case study of Oxford-style debates---a competitive format where the winner is det…
DPLTM uses private winners to improve privacy and utility.
We consider the problem of evaluating the quality of startup companies. This can be quite challenging due to the rarity of successful startup companies and the complexity of factors which impact such success. In this work we collect data on tens of thousands of startup companies, their performance, the backgrounds of t…
Algorithm minimizes regret in non-stationary dueling bandits with unknown parameters.
Unified framework for best arm identification and dueling bandits regret minimization.
New algorithms improve dueling bandit performance in multiplayer settings.
We consider data in the form of pairwise comparisons of n items, with the goal of precisely identifying the top k items for some value of k < n, or alternatively, recovering a ranking of all the items. We analyze the Copeland counting algorithm that ranks the items in order of the number of pairwise comparisons won, an…
We study the effect of altruism in two simple asset exchange models: the yard sale model (winner gets a random fraction of the poorer player's wealth) and the theft and fraud model (winner gets a random fraction of the loser's wealth). We also introduce in these models the concept of bargaining efficiency, which makes …
We propose a simple change to existing neural network structures for better defending against gradient-based adversarial attacks. Instead of using popular activation functions (such as ReLU), we advocate the use of k-Winners-Take-All (k-WTA) activation, a C0 discontinuous function that purposely invalidates the neural …
AI predicts stock winners with 2.43 Sharpe ratio, but returns are highly concentrated.
Improved algorithm for adaptive dueling bandits with near-optimal regret bound.
LoRA-MCL improves language models by generating diverse sentence continuations.
New algorithms for batched dueling bandits with improved regret bounds.
Inspired by the advances in biological science, the study of sparse binary projection models has attracted considerable recent research attention. The models project dense input samples into a higher-dimensional space and output sparse binary data representations after the Winner-Take-All competition, subject to the co…
A fast ML method solves complex combinatorial auction problems.
Linear memory stores associations up to a logarithmic scale, but listwise retrieval can handle a quadratic scale.
New deep learning model robust to adversarial attacks using stochastic LWTA units.
We look at how asset exchange models can be mapped to random iterated function systems (IFS) giving new insights into the dynamics of wealth accumulation in such models. In particular, we focus on the "yard-sale" (winner gets a random fraction of the poorer players wealth) and the "theft-and-fraud" (winner gets a rando…
Visual Question Answering (VQA) requires AI models to comprehend data in two domains, vision and text. Current state-of-the-art models use learned attention mechanisms to extract relevant information from the input domains to answer a certain question. Thus, robust attention mechanisms are essential for powerful VQA mo…
The goal of this paper is to explore the relationship between momentum effects and liquidity in cryptocurrency markets. Portfolios based on momentum-liquidity bivariate sorts are formed and rebalanced on a varying number of cryptocurrencies through time. We find a strong momentum effect in the most liquid cryptocurrenc…
Recent years have witnessed the successful marriage of finance innovations and AI techniques in various finance applications including quantitative trading (QT). Despite great research efforts devoted to leveraging deep learning (DL) methods for building better QT strategies, existing studies still face serious challen…
Some online advertising offers pay only when an ad elicits a response. Randomness and uncertainty about response rates make showing those ads a risky investment for online publishers. Like financial investors, publishers can use portfolio allocation over multiple advertising offers to pursue revenue while controlling r…
Two approaches integrate qualitative views into portfolio optimization, showing aggregation methods outperform robust optimization.
Punctuated Equilibrium (PE) states that after long periods of evolutionary quiescence, species evolution can take place in short time intervals, where sudden differentiation makes new species emerge and some species extinct. In this paper, we introduce and study the effect of punctuated equilibrium on two different ass…
A novel algorithm for actively trading stocks is presented. While traditional expert advice and "universal" algorithms (as well as standard technical trading heuristics) attempt to predict winners or trends, our approach relies on predictable statistical relations between all pairs of stocks in the market. Our empirica…
We consider PAC-learning a good item from -subsetwise feedback information sampled from a Plackett-Luce probability model, with instance-dependent sample complexity performance. In the setting where subsets of a fixed size can be tested and top-ranked feedback is made available to the learner, we give an algorithm w…