Generalized statistical arbitrage concepts are introduced corresponding to trading strategies which yield positive gains on average in a class of scenarios rather than almost surely. The relevant scenarios or market states are specified via an information system given by a -algebra and so this notion contains classi…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
In recent publications, the authors have considered inverse statistics of the Dow Jones Industrial Averaged (DJIA) [1-3]. Specifically, we argued that the natural candidate for such statistics is the investment horizons distribution. This is the distribution of waiting times needed to achieve a predefined level of retu…
REGAIN learns optimal auxiliary directions for forecast reconciliation.
Inverse statistics in economics is considered. We argue that the natural candidate for such statistics is the investment horizons distribution. This distribution of waiting times needed to achieve a predefined level of return is obtained from (often detrended) historic asset prices. Such a distribution typically goes t…
A new method, InfoGuide, improves automatic clustering analysis.
As regulators pay more attentions to losses rather than gains, we are able to derive a new class of risk statistics, named regulator-based risk statistics with scenario analysis in this paper. This new class of risk statistics can be considered as a kind of risk extension of risk statistics introduced by Kou et al. \ci…
We develop a tractable model of realization utility that studies the role of reference-dependent S-shaped preferences in a dynamic investment setting with reinvestment. Our model generates both voluntarily realized gains and losses. It makes specific predictions about the volume of gains and losses, the holding periods…
An econometric or statistical model may undergo a marginal gain if we admit a new variable to the model, and a marginal loss if we remove an existing variable from the model. Assuming equality of opportunity among all candidate variables, we derive a valuation framework by the expected marginal gain and marginal loss i…
Improved deep learning performance in financial markets by using rank space.
Paper proposes a method to use in silico experiments with foundation models to reduce sample size.
Bayesian calibration for BCP self-assembly models using image data and measure transport.
Bayesian analysis reveals asymmetry in financial data.
Bayesian optimal design of experiments (BODE) has been successful in acquiring information about a quantity of interest (QoI) which depends on a black-box function. BODE is characterized by sequentially querying the function at specific designs selected by an infill-sampling criterion. However, most current BODE method…
Generative model improves intraday electricity price forecasting.
We describe how the market-based average and volatility of the "actual" return, which the investors gain within their market sales, depend on the statistical moments, volatilities, and correlations of the current and past market trade values. We describe three successive approximations. First, we derive the dependence …
The study of record statistics of correlated series is gaining momentum. In this work, we study the records statistics of the time series of select stock market data and the geometric random walk, primarily through simulations. We show that the distribution of the age of records is a power law with the exponent lyi…
Active inference framework improves -statistic estimation efficiency.
The inverse statistics is the distribution of waiting times needed to achieve a predefined level of return obtained from (detrended) historic asset prices \cite{optihori,gainloss}. Such a distribution typically goes through a maximum at a time coined the {\em optimal investment horizon}, , which defines the most…
New method tightens variational representations of divergences for faster learning.
Improved DOA estimation with distributed sensors across multiple frequencies.
We explore a simple lattice field model intended to describe statistical properties of high frequency financial markets. The model is relevant in the cross-disciplinary area of econophysics. Its signature feature is the emergence of a self-organized critical state. This implies scale invariance of the model, without tu…
Breiman's two cultures reconciled through blending statistical thinking.
Relying on recent advances in statistical estimation of covariance distances based on random matrix theory, this article proposes an improved covariance and precision matrix estimation for a wide family of metrics. The method is shown to largely outperform the sample covariance matrix estimate and to compete with state…
A new method models financial returns by separating sign and magnitude, improving forecasting accuracy.
Expands Bayesian experiment design framework to account for model discrepancies.
FisherSFT selects informative examples to fine-tune LLMs efficiently.
Generative AI amplifies data without increasing information, with a mathematical limit.
Kempe discusses NTK approach to machine learning problems.
Artificial intelligence (AI) is intrinsically data-driven. It calls for the application of statistical concepts through human-machine collaboration during generation of data, development of algorithms, and evaluation of results. This paper discusses how such human-machine collaboration can be approached through the sta…
Hybrid method uses LLM to filter lead-lag relationships in prediction markets.
The gain-loss asymmetry, observed in the inverse statistics of stock indices is present for logarithmic return levels that are over , and it is the result of the non-Pearson type auto-correlations in the index. These non-Pearson type correlations can be viewed also as functionally dependent daily volatilities, ext…
Transform non-private e-values into differentially private ones.
Teaches deep learning to statisticians.
Unified framework for high-dimensional online learning with non-divergent error bounds and adaptive gains.
The scalability of statistical estimators is of increasing importance in modern applications. One approach to implementing scalable algorithms is to compress data into a low dimensional latent space using dimension reduction methods. In this paper we develop an approach for dimension reduction that exploits the assumpt…
Enhances U-statistics for semi-supervised datasets using unlabeled data.
Enhances robustness in experimental design through Generalised Bayesian inference.
Starting from the requirement that risk measures of financial portfolios should be based on their losses, not their gains, we define the notion of loss-based risk measure and study the properties of this class of risk measures. We characterize loss-based risk measures by a representation theorem and give examples of su…
The study calculates the risk of semi-supervised multitask learning on Gaussian mixtures.
It is common to subsample Markov chain output to reduce the storage burden. Geyer (1992) shows that discarding out of every observations will not improve statistical efficiency, as quantified through variance in a given computational budget. That observation is often taken to mean that thinning MCMC output ca…
We investigate the adversarial bandit problem with multiple plays under semi-bandit feedback. We introduce a highly efficient algorithm that asymptotically achieves the performance of the best switching -arm strategy with minimax optimal regret bounds. To construct our algorithm, we introduce a new expert advice alg…
Data science enhances knot theory by analyzing invariant relations.
Improves survey sampling with unbiased machine learning methods.
This paper quantifies privacy loss in exploratory data analysis.
This study applies variational inference to improve music emotion recognition.
Paper proposes a statistical test for transfer learning in linear regression.
Paper proposes a statistical test for feature selection pipelines using selective inference.
The complex networks approach has been gaining popularity in analysing investor behaviour and stock markets, but within this approach, initial public offerings (IPO) have barely been explored. We fill this gap in the literature by analysing investor clusters in the first two years after the IPO filing in the Helsinki S…