Paper explores limits and possibilities of aligning LLMs with human preferences.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
RLHF performs well despite violating social choice theory axioms.
New method identifies Condorcet winner in dueling bandits with improved sample complexity.
The paper offers a simple proof of Condorcet's jury theorem.
New algorithm for identifying Condorcet team in noisy comparisons.
New algorithm reduces dynamic regret in non-stationary dueling bandits using a weighted Borda score.
Condorcet's Jury Theorem has been invoked for ensemble classifiers to indicate that the combination of many classifiers can have better predictive performance than a single classifier. Such a theoretical underpinning is unknown for consensus clustering. This article extends Condorcet's Jury Theorem to the mean partitio…
This paper tackles combinatorial pure exploration for dueling bandits, aiming to find the best candidate-position match.
The famous Banach-Tarski paradox claims that the three dimensional rotation group acts on the two dimensional sphere paradoxically. In this paper, we generalize their result to show that the classical group acts on the flag manifold paradoxically.
New algorithm achieves near-optimal performance in dueling bandit problem.
New pricing theory solves St. Petersburg paradox.
In this paper, we propose a Double Thompson Sampling (D-TS) algorithm for dueling bandit problems. As indicated by its name, D-TS selects both the first and the second candidates according to Thompson Sampling. Specifically, D-TS maintains a posterior distribution for the preference matrix, and chooses the pair of arms…
New algorithms improve dueling bandit performance in multiplayer settings.
Paper revisits five IF paradoxes using differential geometry.
Karl Menger's 1934 paper on the St. Petersburg paradox contains mathematical errors that invalidate his conclusion that unbounded utility functions, specifically Bernoulli's logarithmic utility, fail to resolve modified versions of the St. Petersburg paradox.
This paper develops a cohomological hierarchy for bistable visual paradoxes.
Improved algorithm for adaptive dueling bandits with near-optimal regret bound.
New algorithms for batched dueling bandits with improved regret bounds.
Siegel's paradox is a fundamental question in international finance about exchange rates for futures contracts and has puzzled many scholars for over forty years. The unorthodox approach presented in this article leads to an arbitrage-free solution which is invariant under currency re-denominations and is symmetric, as…
In this article, I will present a paradox whose purpose is to draw your attention to an important topic in finance, concerning the non-independence of the financial returns (non-ergodic hypothesis). In this paradox, we have two people sitting at a table separated by a black sheet so that they cannot see each other and …
The paper explores game-theoretic alignment of LLMs with human preferences, finding limitations and conditions.
The Allais and Ellsberg paradoxes show that the expected utility hypothesis and Savage's Sure-Thing Principle are violated in real life decisions. The popular explanation in terms of 'ambiguity aversion' is not completely accepted. On the other hand, we have recently introduced a notion of 'contextual risk' to mathemat…
Thompson Sampling shows polynomial regret for combinatorial semi-bandits with subgaussian rewards.
Algorithm minimizes regret in non-stationary dueling bandits with unknown parameters.
Compact models match or exceed GPT's performance in financial news sentiment analysis.
A resolution of the St. Petersburg paradox is presented. In contrast to the standard resolution, utility is not required. Instead, the time-average performance of the lottery is computed. The final result can be phrased mathematically identically to Daniel Bernoulli's resolution, which uses logarithmic utility, but is …
The paper tackles individualized decision-making under unmeasured confounding, providing a novel minimax solution and a paradox.
A mathematical paradox shows secant planes don't always form a tangent plane, but some analogies hold with a specific vector product.
Study on tracking preference shifts in dueling bandits problems.
Unified framework for best arm identification and dueling bandits regret minimization.
We consider strategies of investments into options and diffusion market model. It is shown that there exists a correct proportion between "put" and "call" in the portfolio such that the average gain is almost always positive for a generic Black and Scholes model. This gain is zero if and only if the market price of ris…
New methods resolve conflicting treatment effect estimates in health tech assessments.
Multi-armed bandit(MAB) problem is a reinforcement learning framework where an agent tries to maximise her profit by proper selection of actions through absolute feedback for each action. The dueling bandits problem is a variation of MAB problem in which an agent chooses a pair of actions and receives relative feedback…
Partial covariance factorizes in path diagrams, simplifying analysis.
Buying or selling assets leads to transaction costs for the investor. On one hand, it is well know to all market practionaires that the transaction costs are positive on average and present therefore systematic loss. On the other hand, for every trade, there is a buy side and a sell side, the total amount of asset and …
Researchers solved a geometry paradox for creased tubes.
Constant and symmetric price impact functions, most commonly used in agent-based market modelling, are shown to give rise to paradoxical and inconsistent outcomes in the simplest case of arbitrage exploitation when open-hold-close actions are considered. The solution of the paradox lies in the non-constant nature of re…
In this article we will propose a completely new point of view for solving one of the most important paradoxes concerning game theory. The solution develop shifts the focus from the result to the strategy s ability to operate in a cognitive way by exploiting useful information about the system. In order to determine fr…
GenAI adoption paradoxically lowers ROE for U.S. banks, with spillovers but systemic risk concerns.
Prior work finds a diversity paradox: diversity breeds innovation, and yet, underrepresented groups that diversify organizations have less successful careers within them. Does the diversity paradox hold for scientists as well? We study this by utilizing a near-population of ~1.2 million US doctoral recipients from 1977…
This paper highlights the role of risk neutral investors in generating endogenous bubbles in derivatives markets. We find that a market for derivatives, which has all the features of a perfect market except completeness and has some risk neutral investors, can exhibit extreme price movements which represent a violation…
This paper describes Simpson's paradox, and explains its serious implications for randomised control trials. In particular, we show that for any number of variables we can simulate the result of a controlled trial which uniformly points to one conclusion (such as 'drug is effective') for every possible combination of t…
We briefly review our recent studies on stochastic processes modelling internet on-line trading. We present a way to evaluate the average waiting time between the observation of the price in financial markets and the next price change, especially in an on-line foreign exchange trading service for individual customers v…
The paradox of the energy transition is that the low marginal costs of new renewable energy sources (RES) drag electricity prices down and discourage investments in flexible productions that are needed to compensate for the lack of dispatchability of the new RES. The energy transition thus discourages the investments t…
Osborne's paradox explained via Bayesian inference and superstatistics.
Simplified NFT games discussed with methods for extracting value.
New method uses correlation-ratio for transfer learning, improving target model inference.
Pachner move 3 ->3 deals with triangulations of four-dimensional manifolds. We present an algebraic relation corresponding in a natural way to this move and based, a bit paradoxically, on three-dimensional geometry.