A novel framework extracts essential factors from order flow data for high-frequency trading.
problem Challenges in extracting and utilizing order flow data due to its large volume and limitations of traditional techniques.
method Proposes a Context Encoder and Factor Extractor for unsupervised learning of important signals from order flow data.
result Extracts superior factors from order flow data, improving stock trend prediction and order execution tasks.
Study integrates deep learning with financial data for improved trading strategies.
problem Enhancing predictive performance in algorithmic trading and portfolio optimization.
method Developed embedding techniques to treat limit order book snapshots as image-based input channels.
result Achieved state-of-the-art performance in high-frequency trading algorithms.
Generates realistic stock market order streams using GANs.
problem Creating high-fidelity stock market data.
method Conditional Wasserstein GAN with auction mechanism and order-book augmentation.
result Generated data is close to real market data.
A method to estimate high order derivatives of data distributions from samples.
problem Estimating high order derivatives of data distributions efficiently and accurately.
method Generalizing denoising score matching via Tweedie's formula to estimate higher order derivatives.
result Models trained with the proposed method can approximate second order derivatives more efficiently and accurately than via automatic differentiation.
Method preserves order in hierarchical clustering of ordered data.
problem Order preserving hierarchical clustering of directed acyclic graphs.
method Combination of classical hierarchical clustering and ultrametric fitting.
result Optimal clustering preserves both cluster quality and order.
A new method clusters categorical data by learning their optimal order and distance.
problem Clustering categorical data lacks a well-defined metric space.
method Order distance metric learning for categorical data.
result Superior clustering accuracy on categorical and mixed datasets.
The paper examines the reliability of limit order book representations in the face of data perturbation.
problem The reliability of limit order book representations under data perturbation.
method Experimental analysis of existing representations and guidelines for future research.
result Existing representations of limit order book data are vulnerable to data perturbation.
Bayesian method learns causal orderings from heterogeneous data.
problem Learning causal structure from heterogeneous data.
method Order-based Bayesian framework for Gaussian DAG models.
result Causal ordering is identifiable up to two permutations.
Improves zeroth-order optimization for private machine learning with public data.
problem High computation and memory cost of first-order DP methods.
method PAZO (Public Data Assisted Zeroth-order Optimization) framework.
result Achieves superior privacy/utility tradeoffs across tasks.
Bayesian method detects Markov order in network paths more reliably.
problem Detecting Markov order in constrained categorical sequences.
method Multi-order Bayesian modelling framework.
result Bayesian method detects correct Markov order more reliably than competing methods.
SOR-Mamba improves Mamba for robust time series forecasting by minimizing channel order bias.
problem Robust time series forecasting with Mamba's sequential order bias.
method SOR-Mamba incorporates regularization to minimize channel order discrepancy and introduces CCM for channel correlation preservation.
result SOR-Mamba enhances robustness to channel order and improves forecasting accuracy.
The existing literature provides evidence that limit order book data can be used to predict short-term price movements in stock markets. This paper proposes a new neural network architecture for predicting return jump arrivals in equity markets with high-frequency limit order book data. This new architecture, based on …
Differentiable relaxation for inferring partial orders from noisy linear data.
problem Inference of partial orders from linear data with noisy observations.
method Introducing a differentiable relaxation to model noisy linear extensions, replacing discontinuous precedence and feasibility with smooth surrogates.
result Smooth posterior that preserves partial-order semantics, supports gradient-based inference, and converges to hard likelihood.
Bayesian method reconstructs hidden higher-order interactions from network data.
problem Lack of explicit higher-order interactions in pairwise network data.
method Bayesian approach based on parsimony, infers higher-order structures when statistically supported.
result Demonstrated applicability to various datasets, synthetic and empirical.
Optimized variable orderings improve autoregressive model performance.
problem Challenges in variable ordering affect autoregressive model efficiency.
method Learn graphical model structure to inform optimal variable orderings.
result Graph-informed orderings yield higher-fidelity samples.
We show that multivariate Hawkes processes coupled with the nonparametric estimation procedure first proposed in Bacry and Muzy (2015) can be successfully used to study complex interactions between the time of arrival of orders and their size, observed in a limit order book market. We apply this methodology to high-fre…
Predicts node sequences in graphs using multi-order network models.
problem Predicting sequences of node traversals in graphs.
method Combines multiple higher-order network models into a multi-order model, fitting and selecting the optimal maximum order.
result Outperforms state-of-the-art algorithms for next-element and full sequence prediction.
Users form information trails as they browse the web, checkin with a geolocation, rate items, or consume media. A common problem is to predict what a user might do next for the purposes of guidance, recommendation, or prefetching. First-order and higher-order Markov chains have been widely used methods to study such se…
Hierarchical Partial-Order Models for Ranking
problem Rank aggregation combining ordered lists
method Hierarchical partial-order models
result Bayesian inference for latent poset hierarchy
This method infers models from data with physical insights, minimizing model order.
problem Learning models from data while preserving physical insights.
method Structure preservation and rank minimization via Sylvester equations.
result Models of low order are obtained with fewer degrees of freedom.
A novel transformer model improves classification of partially ordered sequences.
problem Classification of partially ordered sequences with uncertainty in timestamps.
method Developed a transformer-based model for partially ordered sequences, benchmarked against set models.
result Transformer-based model outperforms set models on three datasets.
Second-order optimizers retain residual information after data deletion, affecting machine unlearning.
problem Residual information in second-order optimizers after data deletion.
method Comparison of first-order and second-order learners, eigendecomposition analysis.
result Second-order optimizers retain residual information, not detectable by first-order analysis.
Modeling financial markets with a novel order flow model.
problem Inconsistent parameter values from long-range memory estimators.
method Tsallis q-exponential distribution for limit order cancellation times.
result Improved accuracy in predicting financial market dynamics.
New measures and tests for high-order interactions in complex data.
problem Challenges in capturing high-order interactions in multivariate data.
method Hierarchy of d-order interaction measures and kernel-based tests. result Established statistical significance of high-order interactions.
Paper uses MBO data for high-frequency price forecasting.
problem Lack of predictive analysis on granular MBO data.
method Introduced normalisation scheme for MBO data, trained deep neural networks.
result Ensemble of MBO and LOB models improves forecasting accuracy.
MarketGPT models financial time series with realistic order flow data.
problem Creating accurate financial market simulations.
method Generative pre-trained transformer (GPT) for long sequence generation.
result Model reproduces key features of real financial markets and stylized facts.
We introduce features for massive data streams. These stream features can be thought of as "ordered moments" and generalize stream sketches from "moments of order one" to "ordered moments of arbitrary order". In analogy to classic moments, they have theoretical guarantees such as universality that are important for lea…
Convolutional Neural Networks (CNN) have been pivotal to the success of many state-of-the-art classification problems, in a wide variety of domains (for e.g. vision, speech, graphs and medical imaging). A commonality within those domains is the presence of hierarchical, spatially agglomerative local-to-global interacti…
LOB-Bench benchmarks generative AI for financial data, outperforming traditional models.
problem Lack of consensus on evaluating generative AI models for financial data.
method Python-based benchmark with LOB statistics and market impact metrics.
result Generative autoregressive models outperform traditional models in LOB data.
Through the analysis of a dataset of ultra high frequency order book updates, we introduce a model which accommodates the empirical properties of the full order book together with the stylized facts of lower frequency financial data. To do so, we split the time interval of interest into periods in which a well chosen r…
Measures price impact in order-driven markets without relying on averages.
problem Measuring price impact in order-driven markets without relying on averages.
method Modeling the limit order book using state-dependent Hawkes processes and defining price impact profile as a function of the compensator of a stochastic process.
result The clustering of sell child orders has a bigger impact on price than their sizes.
A new autoregressive model learns the order of graph generation tasks.
problem Generating graphs in a meaningful order when the canonical order is not obvious.
method Introduces a variant of autoregressive models that dynamically decides the autoregressive order based on data.
result Achieves state-of-the-art results on molecular graph generation benchmarks.
Latent order book models have allowed for significant progress in our understanding of price formation in financial markets. In particular they are able to reproduce a number of stylized facts, such as the square-root impact law. An important question that is raised -- if one is to bring such models closer to real mark…
We present a novel factor analysis method that can be applied to the discovery of common factors shared among trajectories in multivariate time series data. These factors satisfy a precedence-ordering property: certain factors are recruited only after some other factors are activated. Precedence-ordering arise in appli…
Adaptive learning model forecasts financial prices using order book data.
problem Forecasting high-frequency financial time series with non-stationary data.
method Adaptive learning model based on order book data, with stationarity and non-stationarity considerations.
result The model outperforms top fixed models and improves forecasting accuracy.
Bayesian BIC for multi-trial data improves VAR model order selection.
problem Optimal VAR model order selection for multi-trial event-based data.
method Derive and apply Bayesian Information Criterion (BIC) for multi-trial ensemble data.
result Multi-trial BIC successfully recovers real model order and estimates small model order.
Nowadays, machine learning methods have been widely used in stock prediction. Traditional approaches assume an identical data distribution, under which a learned model on the training data is fixed and applied directly in the test data. Although such assumption has made traditional machine learning techniques succeed i…
Enhances stock movement prediction using Higher Order Transformers for multimodal time-series data.
problem Predicting stock movements in financial markets with complex dynamics.
method Introduced Higher Order Transformers, extending self-attention and transformer architecture to capture complex market dynamics. Employed low-rank tensor decomposition and kernel attention to manage computational complexity. Integrated technical and fundamental analysis from historical prices and tweets.
result Demonstrated effectiveness of the method on the Stocknet dataset, improving stock movement prediction.
Regularization is a popular technique in machine learning for model estimation and avoiding overfitting. Prior studies have found that modern ordered regularization can be more effective in handling highly correlated, high-dimensional data than traditional regularization. The reason stems from the fact that the ordered…
Explicit high-order feature interactions efficiently capture essential structural knowledge about the data of interest and have been used for constructing generative models. We present a supervised discriminative High-Order Parametric Embedding (HOPE) approach to data visualization and compression. Compared to deep emb…
Optimizes large stock order execution with LSTM neural networks.
problem Minimizing transaction costs in large stock order execution.
method Trained LSTM neural network to minimize transaction costs.
result LSTM strategy outperforms TWAP and VWAP strategies.
This paper employs machine learning algorithms to forecast German electricity spot market prices. The forecasts utilize in particular bid and ask order book data from the spot market but also fundamental market data like renewable infeed and expected demand. Appropriate feature extraction for the order book data is dev…
Mixes higher-order simplicial complexes for data augmentation.
problem Lack of labeled data for complex systems with multiway interactions.
method Proposes mixup mechanisms for simplicial complexes, including linear and nonlinear mixup, and a convex clustering mixup.
result Synthetic simplicial complexes interpolate between existing data based on homomorphism densities.
Paper proposes a new method for sparse covariance Cholesky factor estimation.
problem Estimating sparse covariance matrices for ordered data.
method Matrix loss penalization approach for sparse Cholesky factor estimation.
result The proposed method outperforms existing regression-based approaches in simulations and real data.
The study examines order flow in financial markets using fractional Lévy stable motion.
problem Challenges in selecting the best models for financial time series data.
method Investigates order disbalance time series from the perspective of fractional Lévy stable motion.
result Orders exhibit stable anti-correlation for 18 randomly selected stocks.
A model predicts influential nodes in complex networks by considering indirect interactions.
problem Identifying influential nodes in complex networks using indirect interactions.
method Proposes MOGen, a multi-order generative model that considers all indirect influences up to a maximum distance.
result MOGen consistently outperforms network models and path-based approaches in predicting influential nodes.
Proposes a method to enhance multi-view learning by maximizing higher order correlations.
problem Losing intrinsic interconnections among multiple views in pairwise correlation maximization.
method Formulates multi-view data as a low rank approximation problem using higher order correlation tensor and solves it with the generating polynomial method.
result Consistently outperforms prior methods on real multi-view data.
Order patterns and permutation entropy have become useful tools for studying biomedical, geophysical or climate time series. Here we study day-to-day market data, and Brownian motion which is a good model for their order patterns. A crucial point is that for small lags (1 up to 6 days), pattern frequencies in financial…