New method prevents cherry-picking in machine learning reports.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Estimates peeking effects in p-values to correct bias.
Study improves risk evaluation timing with right-censored reporting delays.
Deep learning model improves corporate distress prediction using text data.
Paper develops a BERT-based classifier to reduce pathology report annotation workload.
We examine the out-of-equilibrium phase reported by Plerou {\it et. al.} in Nature, {\bf 421}, 130 (2003) using the data of the New York stock market (NYSE) between the years 2001 --2002. We find that the observed two phase phenomenon is an artifact of the definition of the control parameter coupled with the nature of …
Improved language model for French clinical reports achieves state-of-the-art performance in medical NLP tasks.
PAQ8 is an open source lossless data compression algorithm that currently achieves the best compression rates on many benchmarks. This report presents a detailed description of PAQ8 from a statistical machine learning perspective. It shows that it is possible to understand some of the modules of PAQ8 and use this under…
Paper proposes a new method to compute cryptocurrency prices securely.
AI generates a sequence of death causes from hospital records.
New algorithm reduces privacy cost of LDP to central privacy model.
Private Generative Bootstrap protects privacy in statistical reporting.
We report the results of fifteen sets of portfolio selection simulations using stocks in the ASX200 index for the period May 2000 to December 2013. We investigated five portfolio selection methods, randomly and from within industrial groups, and three based on neighbor-Net phylogenetic networks. We report that using ra…
Deep RL agents vary significantly in Atari environments.
Recently we reported on an application of the Tsallis non-extensive statistics to the S&P500 stock index. There we argued that the statistics are applicable to a broad range of markets and exchanges where anamolous (super) diffusion and 'heavy' tails of the distribution are present, as they are in the S&P500. We have c…
We show that the behaviour of Bitcoin has interesting similarities to stock and precious metal markets, such as gold and silver. We report that whilst Litecoin, the second largest cryptocurrency, closely follows Bitcoin's behaviour, it does not show all the reported properties of Bitcoin. Agreements between apparently …
This paper examines risks and uncertainties of changing data sources in machine learning for official statistics.
Paper fine-tunes a language model to predict long-term stock buy signals.
The paper reports the construction of artificial stock market that emerges the similar statistical facts with real data in Indonesian stock market. We use the individual but dominant data, i.e.: PT TELKOM in hourly interval. The artificial stock market shows standard statistical facts, e.g.: volatility clustering, the …
New metric improves topic model evaluation.
This review compares GAMs and neural networks on real-world tabular data.
Conventional SVM-based image coding methods are founded on independently restricting the distortion in every image coefficient at some particular image representation. Geometrically, this implies allowing arbitrary signal distortions in an -dimensional rectangle defined by the -insensitivity zone in eac…
In this paper we study automatically recognized trends and investigate their statistics. To do that we introduce the notion of a wavelength for time series via cross correlation and use this wavelength to calibrate the 1-2-3 trend indicator of Maier-Paape [Automatic One Two Three, Quantitative Finance, 2013] to automat…
GANs improve stochastic dynamics prediction by selecting randomly between models.
Detects out-of-distribution inputs in deep generative models.
Model traffic congestion events using multi-modal data and attention-based neural networks.
New framework tackles deep financial reporting bottleneck by improving hallucination and coherence.
In this report we describe a tool for comparing the performance of graphical causal structure learning algorithms implemented in the TETRAD freeware suite of causal analysis methods. Currently the tool is available as package in the TETRAD source code (written in Java). Simulations can be done varying the number of run…
National statistical systems are the enterprises tasked with collecting, validating and reporting societal attributes. These data serve many purposes - they allow governments to improve services, economic actors to traverse markets, and academics to assess social theories. National statistical systems vary in quality, …
Paper shows comparing single performance scores is insufficient for non-deterministic systems, proposing to compare score distributions.
We analyze high-resolution foreign exchange data consisting of 20 million data points of USD-JPY for 13 years to report firm statistical laws in distributions and correlations of exchange rate fluctuations. A conditional probability density analysis clearly shows the existence of trend-following movements at time scale…
Paper explores DNNs in modulation recognition with channel effects and adversarial attacks.
The study formalizes temporal precision and recall for anomaly detection in sequences.
Inverse statistics in economics is considered. We argue that the natural candidate for such statistics is the investment horizons distribution. This distribution of waiting times needed to achieve a predefined level of return is obtained from (often detrended) historic asset prices. Such a distribution typically goes t…
Study confirms eurozone interbank market stability but finds higher collateral reuse.
AI governance lagging in finance despite widespread use.
This report concerns the problem of dimensionality reduction through information geometric methods on statistical manifolds. While there has been considerable work recently presented regarding dimensionality reduction for the purposes of learning tasks such as classification, clustering, and visualization, these method…
Study shows how financial report sentiment impacts bank profitability.
Deep RL evaluation underestimates uncertainty, leading to misleading conclusions.
Paper explains why plateau phenomenon is rare in modern deep learning.
This paper reports empirical evidence that a neural networks model is applicable to the statistically reliable prediction of foreign exchange rates. Time series data and technical indicators such as moving average, are fed to neural nets to capture the underlying "rules" of the movement in currency exchange rates. The …
This technical report is the union of two contributions to the discussion of the Read Paper "Riemann manifold Langevin and Hamiltonian Monte Carlo methods" by B. Calderhead and M. Girolami, presented in front of the Royal Statistical Society on October 13th 2010 and to appear in the Journal of the Royal Statistical Soc…
We investigate relationship between annual electric power consumption per capita and gross domestic production (GDP) per capita for 131 countries. We found that the relationship can be fitted with a power-law function. We examine the relationship for 47 prefectures in Japan. Furthermore, we investigate values of annual…
We report a statistical analysis of the Island ECN (NASDAQ) order book. We determine the static and dynamic properties of this system, and then analyze them from a physicist's viewpoint using an equivalent particle system obtained by treating orders as massive particles and price as position. We identify the fundamenta…
In the space of cubic forms of surfaces, regarded as a -space and endowed with a natural invariant metric, the ratio of the volumes of those representing umbilic points with negative to those with positive indexes is evaluated in terms of the asymmetry of the metric, defined here. A connection of this …
CLARA generates clinical reports from raw inputs, improving accuracy and efficiency.
We describe a method that infers whether statistical dependences between two observed variables X and Y are due to a "direct" causal link or only due to a connecting causal path that contains an unobserved variable of low complexity, e.g., a binary variable. This problem is motivated by statistical genetics. Given a ge…
Statistical mechanics models node-perturbation learning with noisy baselines.