Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

80160240320 · Jun 202019922001200920172026
48 results for statistical gains

Generalized statistical arbitrage concepts are introduced corresponding to trading strategies which yield positive gains on average in a class of scenarios rather than almost surely. The relevant scenarios or market states are specified via an information system given by a σσ-algebra and so this notion contains classi…

2019-07-22abs ↗pdf ↗

In recent publications, the authors have considered inverse statistics of the Dow Jones Industrial Averaged (DJIA) [1-3]. Specifically, we argued that the natural candidate for such statistics is the investment horizons distribution. This is the distribution of waiting times needed to achieve a predefined level of retu…

2005-11-10abs ↗pdf ↗

REGAIN learns optimal auxiliary directions for forecast reconciliation.

problem Forecast reconciliation from fixed systems; identifying useful auxiliary directions.
method REGAIN learns normalized auxiliary directions, forecasts induced series, and selects directions by loss reduction.
result Gain-selected auxiliary directions improve forecast quality, especially for residual uncertainty.

Inverse statistics in economics is considered. We argue that the natural candidate for such statistics is the investment horizons distribution. This distribution of waiting times needed to achieve a predefined level of return is obtained from (often detrended) historic asset prices. Such a distribution typically goes t…

2002-11-02abs ↗pdf ↗

As regulators pay more attentions to losses rather than gains, we are able to derive a new class of risk statistics, named regulator-based risk statistics with scenario analysis in this paper. This new class of risk statistics can be considered as a kind of risk extension of risk statistics introduced by Kou et al. \ci…

2019-04-16abs ↗pdf ↗

We develop a tractable model of realization utility that studies the role of reference-dependent S-shaped preferences in a dynamic investment setting with reinvestment. Our model generates both voluntarily realized gains and losses. It makes specific predictions about the volume of gains and losses, the holding periods…

2014-08-12abs ↗pdf ↗

Improved deep learning performance in financial markets by using rank space.

problem High volatility and low signal-to-noise ratio in equity market dynamics.
method Transformed equity market data from name space to rank space, enabling better learning by DNNs.
result DNNs achieve superior performance in statistical arbitrage in rank space compared to name space.

Paper proposes a method to use in silico experiments with foundation models to reduce sample size.

problem Costly and uncertain randomized experiments.
method Integrates predictions from multiple foundation models with experimental data.
result Estimator offers substantial precision gains, equivalent to a 20% reduction in sample size.

Bayesian calibration for BCP self-assembly models using image data and measure transport.

problem Calibrating models of BCP self-assembly from image data with aleatory uncertainty.
method Likelihood-free inference via measure transport and summary statistics.
result Expected information gains can be computed efficiently for model calibration.

Generative model improves intraday electricity price forecasting.

problem Intraday electricity price forecasting for improved trading strategies.
method Generative neural network model for probabilistic path forecasts.
result Generative model leads to higher profit gains than benchmark methods.

We describe how the market-based average and volatility of the "actual" return, which the investors gain within their market sales, depend on the statistical moments, volatilities, and correlations of the current and past market trade values. We describe three successive approximations. First, we derive the dependence …

2023-04-02abs ↗pdf ↗

The study of record statistics of correlated series is gaining momentum. In this work, we study the records statistics of the time series of select stock market data and the geometric random walk, primarily through simulations. We show that the distribution of the age of records is a power law with the exponent αα lyi…

2014-06-24abs ↗pdf ↗

The inverse statistics is the distribution of waiting times needed to achieve a predefined level of return obtained from (detrended) historic asset prices \cite{optihori,gainloss}. Such a distribution typically goes through a maximum at a time coined the {\em optimal investment horizon}, τρτ^*_ρ, which defines the most…

2006-01-02abs ↗pdf ↗

New method tightens variational representations of divergences for faster learning.

problem Improving tightness of variational representations of divergences for faster statistical estimation.
method Improved objective functionals constructed via an auxiliary optimization problem, leveraging neural network approximation.
result Tighter variational representations can result in significantly faster learning and more accurate estimation of divergences.

Improved DOA estimation with distributed sensors across multiple frequencies.

problem Sensor gain uncertainties and directional perturbations in multi-frequency scenarios.
method Distributed optimization with local coherence models and iterative exchange of information.
result Advantages in statistical and computational efficiency through parallel iterative technique.

A new method models financial returns by separating sign and magnitude, improving forecasting accuracy.

problem Capturing nonlinear predictability in financial return dynamics.
method Decomposes returns into sign and magnitude components, using a joint distribution model.
result Significantly outperforms traditional linear models in forecasting U.S. stock market returns.

Expands Bayesian experiment design framework to account for model discrepancies.

problem Model misspecification in Bayesian optimal experiment design.
method Introduces Expected General Information Gain and Expected Discriminatory Information criteria.
result Demonstrates improved robustness and detection capabilities in experiment design.

FisherSFT selects informative examples to fine-tune LLMs efficiently.

problem Adapting large language models to new domains efficiently.
method Selects examples maximizing information gain using Hessian of log-likelihood.
result Empirically demonstrates improved performance with reduced computational cost.

Artificial intelligence (AI) is intrinsically data-driven. It calls for the application of statistical concepts through human-machine collaboration during generation of data, development of algorithms, and evaluation of results. This paper discusses how such human-machine collaboration can be approached through the sta…

2017-12-08abs ↗pdf ↗

Hybrid method uses LLM to filter lead-lag relationships in prediction markets.

problem Challenges in discovering robust lead-lag relationships in prediction markets due to spurious correlations.
method Two-stage approach: statistical Granger causality followed by LLM semantic re-ranking.
result LLM-based method outperforms statistical baseline, increasing win rate and reducing average loss magnitude.

The gain-loss asymmetry, observed in the inverse statistics of stock indices is present for logarithmic return levels that are over 2%2\%, and it is the result of the non-Pearson type auto-correlations in the index. These non-Pearson type correlations can be viewed also as functionally dependent daily volatilities, ext…

2016-08-16abs ↗pdf ↗

Unified framework for high-dimensional online learning with non-divergent error bounds and adaptive gains.

problem Divergence of error bounds in high-dimensional online learning as data batches increase.
method Asynchronous decomposition framework with summary statistics and dynamic regularization.
result Non-divergent error bounds and adaptive gains in sparse online optimization.

The scalability of statistical estimators is of increasing importance in modern applications. One approach to implementing scalable algorithms is to compress data into a low dimensional latent space using dimension reduction methods. In this paper we develop an approach for dimension reduction that exploits the assumpt…

2015-04-13abs ↗pdf ↗

Enhances robustness in experimental design through Generalised Bayesian inference.

problem Poor inference and estimates of information gain when statistical model is incorrectly specified.
method Generalised Bayesian (Gibbs) inference framework applied to experimental design.
result GBOED enhances robustness to outliers and incorrect assumptions about noise distribution.

Starting from the requirement that risk measures of financial portfolios should be based on their losses, not their gains, we define the notion of loss-based risk measure and study the properties of this class of risk measures. We characterize loss-based risk measures by a representation theorem and give examples of su…

2011-10-07abs ↗pdf ↗

The study calculates the risk of semi-supervised multitask learning on Gaussian mixtures.

problem Understanding the risk in semi-supervised multitask learning on Gaussian mixtures.
method Statistical physics methods applied to Gaussian mixture models.
result The study evaluates the performance gain of learning tasks together versus separately.

It is common to subsample Markov chain output to reduce the storage burden. Geyer (1992) shows that discarding k1k-1 out of every kk observations will not improve statistical efficiency, as quantified through variance in a given computational budget. That observation is often taken to mean that thinning MCMC output ca…

2015-10-27abs ↗pdf ↗

Paper proposes a statistical test for feature selection pipelines using selective inference.

problem Assessing the significance of feature selection pipelines in data analysis.
method Selective inference technique applied to feature selection pipelines composed of various algorithms.
result The proposed statistical test controls false positive feature selection probabilities.

The complex networks approach has been gaining popularity in analysing investor behaviour and stock markets, but within this approach, initial public offerings (IPO) have barely been explored. We fill this gap in the literature by analysing investor clusters in the first two years after the IPO filing in the Helsinki S…

2019-05-31abs ↗pdf ↗