Study predicts price predictability in ultra-high frequency financial data using entropy tests.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study uses Hawkes and diffusion models to analyze stock price dynamics.
A streaming algorithm estimates quadratic covariation from financial data efficiently.
Paper uses TCN with attention to predict UHF stock price changes.
This study examines how financial tick data becomes more random with time aggregation.
Social and economic systems are complex adaptive systems, in which heterogenous agents interact and evolve in a self-organized manner, and macroscopic laws emerge from microscopic properties. To understand the behaviors of complex systems, computational experiments based on physical and mathematical models provide a us…
A detailed analysis of correlation between stock returns at high frequency is compared with simple models of random walks. We focus in particular on the dependence of correlations on time scales - the so-called Epps effect. This provides a characterization of stochastic models of stock price returns which is appropriat…
We study the distributions of event-time returns and clock-time returns at different microscopic timescales using ultra-high-frequency data extracted from the limit-order books of 23 stocks traded in the Chinese stock market in 2003. We find that the returns at the one-trade timescale obey the inverse cubic law. For la…
We present a large-scale study of commonality in liquidity and resilience across assets in an ultra high-frequency (millisecond-timestamped) Limit Order Book (LOB) dataset from a pan-European electronic equity trading facility. We first show that extant work in quantifying liquidity commonality through the degree of ex…
At the ultra high frequency level, the notion of price of an asset is very ambiguous. Indeed, many different prices can be defined (last traded price, best bid price, mid price,...). Thus, in practice, market participants face the problem of choosing a price when implementing their strategies. In this work, we propose …
Using ultra-high-frequency data extracted from the order flows of 23 stocks traded on the Shenzhen Stock Exchange, we study the empirical regularities of order placement in the opening call auction, cool period and continuous auction. The distributions of relative logarithmic prices against reference prices in the thre…
Paper forecasts financial trading durations using a new point process model.
VOLARE provides standardized realized volatility measures from financial data.
When stock prices are observed at high frequencies, more information can be utilized in estimation of parameters of the price process. However, high-frequency data are contaminated by the market microstructure noise which causes significant bias in parameter estimation when not taken into account. We propose an estimat…
Study compares Fourier estimators to mitigate asynchrony effects in finance.
Through the analysis of a dataset of ultra high frequency order book updates, we introduce a model which accommodates the empirical properties of the full order book together with the stylized facts of lower frequency financial data. To do so, we split the time interval of interest into periods in which a well chosen r…
By studying all the trades and best bids/asks of ultra high frequency snapshots recorded from the order books of a basket of 10 futures assets, we bring qualitative empirical evidence that the impact of a single trade depends on the intertrade time lags. We find that when the trading rate becomes faster, the return var…
We have analyzed the statistical probabilities of limit-order book (LOB) shape through building the book using the ultra-high-frequency data from 23 liquid stocks traded on the Shenzhen Stock Exchange in 2003. We find that the averaged LOB shape has a maximum away from the same best price for both buy and sell LOBs. Th…
We propose a mathematical procedure for finding informed traders in ultra-high frequency trading. We wrote it as Vector ARMA and found condition of its stationarity. For the price exposure complied with ARMA(1,2) we proved that underlying asset price difference can be derived as ARMA(1,1) process. For validation of the…
Long-range correlation in financial time series reflects the complex dynamics of the stock markets driven by algorithms and human decisions. Our analysis exploits ultra-high frequency order book data from NASDAQ Nordic over a period of three years to numerically estimate the power-law scaling exponents using detrended …
Study uses multi-kernel Hawkes models to analyze high-frequency price dynamics.
Motivated by a zero-intelligence approach, the aim of this paper is to connect the microscopic (discrete price and volume), mesoscopic (discrete price and continuous volume) and macroscopic (continuous price and volume) frameworks for the modelling of limit order books, with a view to providing a natural probabilistic …
We study the statistical regularities of opening call auction using the ultra-high-frequency data of 22 liquid stocks traded on the Shenzhen Stock Exchange in 2003. The distribution of the relative price, defined as the relative difference between the order price in opening call auction and the closing price of last tr…
New method for spot volatility estimation with reduced microstructure noise.
We propose a limit order book (LOB) model with dynamics that account for both the impact of the most recent order and the shape of the LOB. We present an empirical analysis showing that the type of the last order significantly alters the submission rate of immediate future orders, even after accounting for the state of…
Modeling price clustering in financial markets using discrete distributions.
We study the dynamics of order flows around large intraday price changes using ultra-high-frequency data from the Shenzhen Stock Exchange. We find a significant reversal of price for both intraday price decreases and increases with a permanent price impact. The volatility, the volume of different types of orders, the b…
Bayesian model predicts mid-price dynamics in financial markets.
Study predicts stock transaction durations using LSTM and attention mechanism.
Data preprocessing improves data quality for robust data mining.
Big data sets must be carefully partitioned into statistically similar data subsets that can be used as representative samples for big data analysis tasks. In this paper, we propose the random sample partition (RSP) data model to represent a big data set as a set of non-overlapping data subsets, called RSP data blocks,…
A new method for handling imbalanced big data using ensembles and smart data.
Prevents sensitive data generation in diffusion models using labeled and unlabeled data.
Study reveals Data Shapley's inconsistent performance in data selection tasks.
PRRO generates synthetic tabular data that improves SL performance and class distribution.
Defines data science as a natural ecosystem with challenges and missions.
Synthetic data enhances analytics but requires careful volume management.
This paper introduces C-DSL to improve data mining outcomes by considering context.
Proposes using probabilistic models for privacy-preserving synthetic data.
New test ensures quality of shared data in machine learning.
Paper creates fair synthetic data ensuring equal predictions across sensitive attributes.
A new method classifies multiple correlated data streams simultaneously.
DPA preserves data distribution in reduced dimensions.
Efficient synthetic data generation improves model performance on tabular data.
For most problems in science and engineering we can obtain data sets that describe the observed system from various perspectives and record the behavior of its individual components. Heterogeneous data sets can be collectively mined by data fusion. Fusion can focus on a specific target relation and exploit directly ass…
DAERNN models censored data using neural networks with data augmentation.
Data preprocessing techniques are devoted to correct or alleviate errors in data. Discretization and feature selection are two of the most extended data preprocessing techniques. Although we can find many proposals for static Big Data preprocessing, there is little research devoted to the continuous Big Data problem. A…
Data collection is a major bottleneck in machine learning and an active research topic in multiple communities. There are largely two reasons data collection has recently become a critical issue. First, as machine learning is becoming more widely-used, we are seeing new applications that do not necessarily have enough …