The study sets criteria for efficient communication in distributed online learning.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Algorithm beats best constant rebalancing portfolio in long-term investment.
New rules control false discoveries in online anomaly detection for time series data.
A new algorithm for competing agents in a two-sided market setting.
Serial problems can't be efficiently parallelized, affecting machine learning models.
Computational efficiency is an important consideration for deploying machine learning models for time series prediction in an online setting. Machine learning algorithms adjust model parameters automatically based on the data, but often require users to set additional parameters, known as hyperparameters. Hyperparamete…
Complex structures are typical in machine learning. Tailoring learning algorithms for every structure requires an effort that may be saved by defining a generic learning procedure adaptive to any complex structure. In this paper, we propose to map any complex structure onto a generic form, called serialization, over wh…
Stochastic convex optimization algorithms are the most popular way to train machine learning models on large-scale data. Scaling up the training process of these models is crucial, but the most popular algorithm, Stochastic Gradient Descent (SGD), is a serial method that is surprisingly hard to parallelize. In this pap…
Online VSMC efficiently learns SSM parameters in streaming data.
Transformers approximate Bayesian posteriors but not exactly.
We study the relation between serial correlation of financial returns and volatility at intraday level for the S&P500 stock index. At daily and weekly level, serial correlation and volatility are known to be negatively correlated (LeBaron effect). While confirming that the LeBaron effect holds also at intraday level, w…
Stochastic momentum methods trade compute efficiency for serial runtime.
The paper proposes estimators for bid-ask spreads with and without serial dependence.
The book chapter discusses tail risk analysis for financial data using extreme value statistics.
We investigate serial correlation, periodic, aperiodic and scaling behaviour of eigenmodes, i.e. daily price fluctuation time-series derived from eigenvectors, of correlation matrices of shares listed on the Johannesburg Stock Exchange (JSE) from January 1993 to December 2002. Periodic, or calendar, components are dete…
This paper takes a deep learning approach to understand consumer credit risk when e-commerce platforms issue unsecured credit to finance customers' purchase. The "NeuCredit" model can capture both serial dependences in multi-dimensional time series data when event frequencies in each dimension differ. It also captures …
Diffusion models explained via cognitive science.
We study the impact of volatility on intraday serial correlation, at time scales of less than 20 minutes, exploiting a data set with all transaction on SPX500 futures from 1993 to 2001. We show that, while realized volatility and intraday serial correlation are linked, this relation is driven by unexpected volatility o…
In practice daily volatility of portfolio returns is transformed to longer holding periods by multiplying by the square-root of time which assumes that returns are not serially correlated. Under this assumption this procedure of scaling can also be applied to contributions to volatility of the assets in the portfolio. …
ParaMonte::Python streamlines Bayesian data analysis with fast Monte Carlo and MCMC routines.
Recent years have witnessed exciting progress in the study of stochastic variance reduced gradient methods (e.g., SVRG, SAGA), their accelerated variants (e.g, Katyusha) and their extensions in many different settings (e.g., online, sparse, asynchronous, distributed). Among them, accelerated methods enjoy improved conv…
CoT enhances transformer accuracy on serial tasks by enabling serial computation.
W-RNN improves text classification by extracting serialized text semantics.
Proposes rCV to preserve serial correlations in time-series models.
In this paper we investigate the adaptive market efficiency of the agricultural commodity futures market, using a sample of eight futures contracts. Using a battery of nonlinear tests, we uncover the nonlinear serial dependence in the returns series. We run the Hinich portmanteau bicorrelation test to uncover the momen…
We consider the problem of fast time-series data clustering. Building on previous work modeling the correlation-based Hamiltonian of spin variables we present an updated fast non-expensive Agglomerative Likelihood Clustering algorithm (ALC). The method replaces the optimized genetic algorithm based approach (f-SPC) wit…
Algorithm achieves comparable performance to fully dynamic data with only a few batches.
MER algorithm speeds up VI solving with Markovian data.
We implement a master-slave parallel genetic algorithm (PGA) with a bespoke log-likelihood fitness function to identify emergent clusters within price evolutions. We use graphics processing units (GPUs) to implement a PGA and visualise the results using disjoint minimal spanning trees (MSTs). We demonstrate that our GP…
Paper introduces Decentralized Non-stationary Competing Bandits ( exttt{DNCB}) for dynamic matching markets.
Practitioners of Bayesian statistics have long depended on Markov chain Monte Carlo (MCMC) to obtain samples from intractable posterior distributions. Unfortunately, MCMC algorithms are typically serial, and do not scale to the large datasets typical of modern machine learning. The recently proposed consensus Monte Car…
ParaMonte simplifies Monte Carlo simulations for various scientific fields.
We present a general framework for accelerating a large class of widely used Markov chain Monte Carlo (MCMC) algorithms. Our approach exploits fast, iterative approximations to the target density to speculatively evaluate many potential future steps of the chain in parallel. The approach can accelerate computation of t…
Serial crystallography is the field of science that studies the structure and properties of crystals via diffraction patterns. In this paper, we introduce a new serial crystallography dataset comprised of real and synthetic images; the synthetic images are generated through the use of a simulator that is both scalable …
The Sharpe ratio, which is defined as the ratio of the excess expected return of an investment to its standard deviation, has been widely cited in the financial literature by researchers and practitioners. However, very little attention has been paid to the statistical properties of the estimation of the ratio. Lo (200…
Estimates Hurst exponent of log-volatility using KS statistic, addressing serial correlation in financial data.
In this paper we address the problem of discovering a small set of frequent serial episodes from sequential data so as to adequately characterize or summarize the data. We discuss an algorithm based on the Minimum Description Length (MDL) principle and the algorithm is a slight modification of an earlier method, called…
mGRN improves multivariate time series prediction by managing marginal and joint memories.
Method predicts LFSM increments from past observations using codifference.
New algorithm for decentralized matching markets without prior preference rankings.
A new algorithm reduces bias in estimating model parameters.
A new training method speeds up ResNet training by 3x with minimal accuracy loss.
We introduce two Python frameworks to train neural networks on large datasets: Blocks and Fuel. Blocks is based on Theano, a linear algebra compiler with CUDA-support. It facilitates the training of complex neural network models by providing parametrized Theano operations, attaching metadata to Theano's symbolic comput…
A method for noise reduction in functional time series using FPCA.
The FSRM uses a multifractional process to capture price multifractality, revealing serial information for forecasting.
Cluster jackknife improves inference for staggered DID methods.
We extend conformal inference to general settings that allow for time series data. Our proposal is developed as a randomization method and accounts for potential serial dependence by including block structures in the permutation scheme. As a result, the proposed method retains the exact, model-free validity when the da…
New method for PKM inverse dynamics second derivatives efficiently.