New algorithm reduces dynamic regret in non-stationary dueling bandits using a weighted Borda score.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper tackles combinatorial pure exploration for dueling bandits, aiming to find the best candidate-position match.
The dueling bandit problem is a variation of the classical multi-armed bandit in which the allowable actions are noisy comparisons between pairs of arms. This paper focuses on a new approach for finding the "best" arm according to the Borda criterion using noisy comparisons. We prove that in the absence of structural a…
New algorithm for minimizing regret in adversarial dueling bandits.
The paper minimizes Borda regret in dueling bandits models.
Research on the multi-armed bandit problem has studied the trade-off of exploration and exploitation in depth. However, there are numerous applications where the cardinal absolute-valued feedback model (e.g. ratings from one to five) is not suitable. This has motivated the formulation of the duelling bandits problem, w…
The paper proposes methods to identify and sample from mixtures of Mallows models for top-k rankings.
New framework for reinforcement learning with adversarial preferences in tabular MDPs.
A new method corrects for bias in selecting the best candidate.
We consider the problem of learning over non-stationary ranking streams. The rankings can be interpreted as the preferences of a population and the non-stationarity means that the distribution of preferences changes over time. Our goal is to learn, in an online manner, the current distribution of rankings. The bottlene…
New method identifies Condorcet winner in dueling bandits with improved sample complexity.
New method uses geometric properties for better density estimation.
Algorithm identifies Copeland winners in dueling bandits with ternary feedback.
A game-theoretic approach to multi-criteria ranking from ordinal data.
SIREN protocol corrects optimistic winner's scores in LLM evaluation.
Study shows big winner stocks significantly impact passive and active investment strategies.
This paper reexamines the profitability of loser, winner and contrarian portfolios in the Chinese stock market using monthly data of all stocks traded on the Shanghai Stock Exchange and Shenzhen Stock Exchange covering the period from January 1997 to December 2012. We find evidence of short-term and long-term contraria…
A new method aggregates generative classifiers to resist adversarial attacks.
Data-driven decision-making often overestimates benefits due to the winner's curse.
TimeMCL forecasts diverse time series futures using neural networks and WTA loss.
Stochastic LWTA networks resist adversarial attacks while maintaining accuracy.
Paper proposes a method to predict MOBA game winners with calibrated confidence.
aMCL uses annealing to improve hypothesis diversity in ambiguous tasks.
New method optimizes treatment policies to avoid winner's curse.
This case study tests the possibility of prediction for "success" (or "winner") components of four stock & shares market indices in a time period of three years from 02-Jul-2009 to 29-Jun-2012.We compare their performance ain two time frames: initial frame three months at the beginning (02/06/2009-30/09/2009) and the f…
In our model, traders interact with each other and with a central bank; they are taxed on the money they make, some of which is dissipated away by corruption. A generic feature of our model is that the richest trader always wins by 'consuming' all the others: another is the existence of a threshold wealth, below wh…
Sentiment Analysis of microblog feeds has attracted considerable interest in recent times. Most of the current work focuses on tweet sentiment classification. But not much work has been done to explore how reliable the opinions of the mass (crowd wisdom) in social network microblogs such as twitter are in predicting ou…
Public debates are a common platform for presenting and juxtaposing diverging views on important issues. In this work we propose a methodology for tracking how ideas flow between participants throughout a debate. We use this approach in a case study of Oxford-style debates---a competitive format where the winner is det…
DPLTM uses private winners to improve privacy and utility.
We consider the problem of evaluating the quality of startup companies. This can be quite challenging due to the rarity of successful startup companies and the complexity of factors which impact such success. In this work we collect data on tens of thousands of startup companies, their performance, the backgrounds of t…
Algorithm minimizes regret in non-stationary dueling bandits with unknown parameters.
With the growing interest on Network Analysis, Relational Data Mining is becoming an emphasized domain of Data Mining. This paper addresses the problem of extracting representative elements from a relational dataset. After defining the notion of degree of representativeness, computed using the Borda aggregation procedu…
We study the K-armed dueling bandit problem, a variation of the standard stochastic bandit problem where the feedback is limited to relative comparisons of a pair of arms. The hardness of recommending Copeland winners, the arms that beat the greatest number of other arms, is characterized by deriving an asymptotic regr…
Unified framework for best arm identification and dueling bandits regret minimization.
New algorithms improve dueling bandit performance in multiplayer settings.
We study the effect of altruism in two simple asset exchange models: the yard sale model (winner gets a random fraction of the poorer player's wealth) and the theft and fraud model (winner gets a random fraction of the loser's wealth). We also introduce in these models the concept of bargaining efficiency, which makes …
We propose a simple change to existing neural network structures for better defending against gradient-based adversarial attacks. Instead of using popular activation functions (such as ReLU), we advocate the use of k-Winners-Take-All (k-WTA) activation, a C0 discontinuous function that purposely invalidates the neural …
AI predicts stock winners with 2.43 Sharpe ratio, but returns are highly concentrated.
Improved algorithm for adaptive dueling bandits with near-optimal regret bound.
LoRA-MCL improves language models by generating diverse sentence continuations.
New algorithms for batched dueling bandits with improved regret bounds.
Inspired by the advances in biological science, the study of sparse binary projection models has attracted considerable recent research attention. The models project dense input samples into a higher-dimensional space and output sparse binary data representations after the Winner-Take-All competition, subject to the co…
A fast ML method solves complex combinatorial auction problems.
Linear memory stores associations up to a logarithmic scale, but listwise retrieval can handle a quadratic scale.
New deep learning model robust to adversarial attacks using stochastic LWTA units.
We look at how asset exchange models can be mapped to random iterated function systems (IFS) giving new insights into the dynamics of wealth accumulation in such models. In particular, we focus on the "yard-sale" (winner gets a random fraction of the poorer players wealth) and the "theft-and-fraud" (winner gets a rando…
Visual Question Answering (VQA) requires AI models to comprehend data in two domains, vision and text. Current state-of-the-art models use learned attention mechanisms to extract relevant information from the input domains to answer a certain question. Thus, robust attention mechanisms are essential for powerful VQA mo…
Neural Machine Translation has lately gained a lot of "attention" with the advent of more and more sophisticated but drastically improved models. Attention mechanism has proved to be a boon in this direction by providing weights to the input words, making it easy for the decoder to identify words representing the prese…