Addressing the ongoing examination of high-frequency trading practices in financial markets, we report the results of an extensive empirical study estimating the maximum possible profitability of the most aggressive such practices, and arrive at figures that are surprisingly modest. By "aggressive" we mean any trading …
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Stochastic Gradient Descent (SGD) is widely used in machine learning problems to efficiently perform empirical risk minimization, yet, in practice, SGD is known to stall before reaching the actual minimizer of the empirical risk. SGD stalling has often been attributed to its sensitivity to the conditioning of the probl…
The Normal Means problem plays a fundamental role in many areas of modern high-dimensional statistics, both in theory and practice. And the Empirical Bayes (EB) approach to solving this problem has been shown to be highly effective, again both in theory and practice. However, almost all EB treatments of the Normal Mean…
Aggregation defenses improve deep learning models' robustness against data poisoning attacks.
Practical algorithm for contextual bandits with large action spaces.
NTK theory fails to predict practical behavior of large-width neural networks.
Empirical study shows consistent meta-RL algorithms adapt to OOD tasks.
Develops methods to improve demand counterfactuals from imperfect proxies.
The Heston model is validated for option pricing using theoretical derivations and empirical market data.
We develop a family of accelerated stochastic algorithms that minimize sums of convex functions. Our algorithms improve upon the fastest running time for empirical risk minimization (ERM), and in particular linear least-squares regression, across a wide range of problem settings. To achieve this, we establish a framewo…
Invertibility conditions for observation-driven time series models often fail to be guaranteed in empirical applications. As a result, the asymptotic theory of maximum likelihood and quasi-maximum likelihood estimators may be compromised. We derive considerably weaker conditions that can be used in practice to ensure t…
EQO uses a simple bonus term for efficient exploration in tabular RL.
Empirical Gaussian Processes learn flexible priors from data.
Corrects sample selection bias in empirical risk minimization using importance sampling.
A major challenge in contextual bandits is to design general-purpose algorithms that are both practically useful and theoretically well-founded. We present a new technique that has the empirical and computational advantages of realizability-based approaches combined with the flexibility of agnostic methods. Our algorit…
Certified training improves robustness against adversarial attacks.
Despite widespread interest and practical use, the theoretical properties of random forests are still not well understood. In this paper we contribute to this understanding in two ways. We present a new theoretically tractable variant of random regression forests and prove that our algorithm is consistent. We also prov…
DAL enhances black-box bandit algorithms for non-stationary environments.
Natural gradient descent, which preconditions a gradient descent update with the Fisher information matrix of the underlying statistical model, is a way to capture partial second-order information. Several highly visible works have advocated an approximation known as the empirical Fisher, drawing connections between ap…
New method uses SURE to denoise signals, outperforming NPMLE.
This thesis improves practical reinforcement learning methods with robustness, scalability, and efficiency.
Causal inference is central to many areas of artificial intelligence, including complex reasoning, planning, knowledge-base construction, robotics, explanation, and fairness. An active community of researchers develops and enhances algorithms that learn causal models from data, and this work has produced a series of im…
Paper analyzes high-dimensional portfolio risks and finds empirical out-of-sample relative loss is more reliable.
Signature kernel handles sequential data with theoretical and practical advantages.
Synthetic experiments are crucial for assessing causal machine learning methods.
Paper presents ERM with -divergence regularization and its properties.
Introduces BPEL for EL, enhancing flexibility and using MCMC for inference.
We propose a novel approach for sampling realistic financial correlation matrices. This approach is based on generative adversarial networks. Experiments demonstrate that generative adversarial networks are able to recover most of the known stylized facts about empirical correlation matrices estimated on asset returns.…
The data processing inequality doesn't always hold in practice, showing benefits in low-level tasks.
Selecting an optimizer is a central step in the contemporary deep learning pipeline. In this paper, we demonstrate the sensitivity of optimizer comparisons to the hyperparameter tuning protocol. Our findings suggest that the hyperparameter search space may be the single most important factor explaining the rankings obt…
The interest rates (or nominal yields) can be negative, this is an unavoidable fact which has already been visible during the Great Depression (1929-39). Nowadays we can find negative rates easily by e.g. auditing. Several theoretical and practical ideas how to model and eventually overcome empirical negative rates can…
The paper provides theoretical guarantees for neural network-based anomaly detection.
Empirical median performs well in estimating location with varying scales.
New method tackles MDPs by learning normalized representations efficiently.
An important class of distance metrics proposed for training generative adversarial networks (GANs) is the integral probability metric (IPM), in which the neural net distance captures the practical GAN training via two neural networks. This paper investigates the minimax estimation problem of the neural net distance ba…
New algorithm closes empirical gap in PFSGD performance.
We study PCA as a stochastic optimization problem and propose a novel stochastic approximation algorithm which we refer to as "Matrix Stochastic Gradient" (MSG), as well as a practical variant, Capped MSG. We study the method both theoretically and empirically.
Estimating properties of discrete distributions is a fundamental problem in statistical learning. We design the first unified, linear-time, competitive, property estimator that for a wide class of properties and for all underlying distributions uses just samples to achieve the performance attained by the empirical…
Penalized likelihood approaches are widely used for high-dimensional regression. Although many methods have been proposed and the associated theory is now well-developed, the relative efficacy of different approaches in finite-sample settings, as encountered in practice, remains incompletely understood. There is theref…
Recently, over-parameterized neural networks have been extensively analyzed in the literature. However, the previous studies cannot satisfactorily explain why fully trained neural networks are successful in practice. In this paper, we present a new theoretical framework for analyzing over-parameterized neural networks …
Paper establishes no-regret property for practical EGO optimization.
The article presents calculations that prove practical importance of the earlier derived theoretical relationship between the interest rate on the interbank credit market, volume of investment and the quantity of securities tradable on the stock exchange.
New adaptive first-order methods improve on quasi-Newton variants.
New method combines Bloom filters and belief propagation for efficient group testing.
BLAE solves batched linear bandits with optimal regret and practical performance.
New empirical PAC-Bayes bound for Markov chains with finite state space.
Given a set of empirical observations, conditional density estimation aims to capture the statistical relationship between a conditional variable and a dependent variable by modeling their conditional probability . The paper develops best practices for conditional den…
We aim to analyze the relation between two random vectors that may potentially have both different number of attributes as well as realizations, and which may even not have a joint distribution. This problem arises in many practical domains, including biology and architecture. Existing techniques assume the vectors to …