Enhances financial time series forecasting with a multi-period learning framework.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper analyzes SGD with increasingly weighted averaging for optimization and generalization.
Optimal energy trading strategy for intraday markets using Hawkes processes.
New class of heavy-tailed distributions shows weighted averages dominate individual variables.
The study introduces anytime learning schedules for large language models without fixed horizons.
Paper explores weighted averaging schemes for SGD, achieving asymptotic normality and optimality.
Unified analysis of finite weight averaging methods in deep learning.
GACTGAN synthesizes tabular data better with less computational overhead.
The paper extends a prediction method to curved spaces.
New algorithm samples Bayesian neural networks for improved calibration.
MPC framework reduces execution costs and schedule deviations in trading.
Simplified analysis of SGD for linear regression with weight averaging.
This paper sets out to provide a general framework for the pricing of average-type options via lower and upper bounds. This class of options includes Asian, basket and options on the volume-weighted average price. We demonstrate that in cases under discussion lower bounds allow for the dimensionality of the problem to …
We propose Stochastic Weight Averaging in Parallel (SWAP), an algorithm to accelerate DNN training. Our algorithm uses large mini-batches to compute an approximate solution quickly and then refines it by averaging the weights of multiple models computed independently and in parallel. The resulting models generalize equ…
TT-DAC-PS: A deterministic actor-critic approach for optimal trade execution
New bounds for SGD show improved performance in various settings.
Extends unbiased simulation method to Asian options.
Distributed statistical learning problems arise commonly when dealing with large datasets. In this setup, datasets are partitioned over machines, which compute locally, and communicate short messages. Communication is often the bottleneck. In this paper, we study one-step and iterative weighted parameter averaging in s…
The Ricci tensor (Ric) is fundamental to Einstein's geometric theory of gravitation. The 3-dimensional Ric of a spacelike surface vanishes at the moment of time symmetry for vacuum spacetimes. The 4-dimensional Ric is the Einstein tensor for such spacetimes. More recently the Ric was used by Hamilton to define a non-li…
Proposes SWA for adversarial training to improve model robustness.
Consider a family of portfolio strategies with the aim of achieving the asymptotic growth rate of the best one. The idea behind Cover's universal portfolio is to build a wealth-weighted average which can be viewed as a buy-and-hold portfolio of portfolios. When an optimal portfolio exists, the wealth-weighted average c…
New covariance estimator for financial portfolios.
This paper investigates robust and efficient DR/RDR estimators for WATEs.
WASH trains ensembles with shuffled weights to improve accuracy and reduce communication.
New method SF-AdamW trains large models without decay phases or memory overhead.
We study the problem of optimal execution of a trading order under Volume Weighted Average Price (VWAP) benchmark, from the point of view of a risk-averse broker. The problem consists in minimizing mean-variance of the slippage, with quadratic transaction costs. We devise multiple ways to solve it, in particular we stu…
This note justifies approximations of arithmetic forwards using weighted averages of overnight forwards.
In this paper, we focus on quantifying model stability as a function of random seed by investigating the effects of the induced randomness on model performance and the robustness of the model in general. We specifically perform a controlled study on the effect of random seeds on the behaviour of attention, gradient-bas…
We analyze the generalization and robustness of the batched weighted average algorithm for V-geometrically ergodic Markov data. This algorithm is a good alternative to the empirical risk minimization algorithm when the latter suffers from overfitting or when optimizing the empirical risk is hard. For the generalization…
ACOWA improves distributed sparse classification with extra communication round.
In mixture model-based clustering applications, it is common to fit several models from a family and report clustering results from only the `best' one. In such circumstances, selection of this best model is achieved using a model selection criterion, most often the Bayesian information criterion. Rather than throw awa…
Recurrent Neural Networks (RNN) are a type of statistical model designed to handle sequential data. The model reads a sequence one symbol at a time. Each symbol is processed based on information collected from the previous symbols. With existing RNN architectures, each symbol is processed using only information from th…
Bayesian inference for inverse problems using mean-shift interacting particles
RATE metrics evaluate treatment prioritization rules, subsuming existing methods.
We propose methods for distributed graph-based multi-task learning that are based on weighted averaging of messages from other machines. Uniform averaging or diminishing stepsize in these methods would yield consensus (single task) learning. We show how simply skewing the averaging weights or controlling the stepsize a…
In this short note, we study an optimization problem of expected implementation shortfall (IS) cost under general shaped market impact functions. In particular, we find that an optimal strategy is a VWAP (volume weighted average price) execution strategy when the market model is a Black-Scholes type with stochastic clo…
Paper uses DRL to optimize trade execution, outperforming VWAP and TWAP.
New algorithms bound graph structure sampling and learning high-dimensional graphical models.
The market impact (MI) of Volume Weighted Average Price (VWAP) orders is a convex function of a trading rate, but most empirical estimates of transaction cost are concave functions. How is this possible? We show that isochronic (constant trading time) MI is slightly convex, and isochoric (constant trading volume) MI is…
Data augmentation improves robustness in adversarial training.
Aggregating multiple learners through an ensemble of models aim to make better predictions by capturing the underlying distribution of the data more accurately. Different ensembling methods, such as bagging, boosting, and stacking/blending, have been studied and adopted extensively in research and practice. While baggi…
Introduces a new quantile regression method for financial and wage data analysis.
Unified approach to aggregating models and preferences.
In causal inference, a variety of causal effect estimands have been studied, including the sample, uncensored, target, conditional, optimal subpopulation, and optimal weighted average treatment effects. Ad-hoc methods have been developed for each estimand based on inverse probability weighting (IPW) and on outcome regr…
We present in this paper a new premium computation principle based on the use of prior information from multiple sources for computing the premium charged to a policyholder. Under this framework, based on the use of Ordered Weighted Averaging (OWA) operators, we propose alternative collective and Bayes premiums and des…
Presently the most successful approaches to semi-supervised learning are based on consistency regularization, whereby a model is trained to be robust to small perturbations of its inputs and parameters. To understand consistency regularization, we conceptually explore how loss geometry interacts with training procedure…
We discuss some consequences of the existence of the holomorphic quadratic Hopf differential on a conformally immersed constant mean curvature topological disc with analytic boundary. In particular, we derive a formula for the mean curvature as a weighted average of the normal curvature of the boundary curve, and a con…
Active learning seeks to build the best possible model with a budget of labelled data by sequentially selecting the next point to label. However the training set is no longer \textit{iid}, violating the conditions required by existing consistency results. Inspired by the success of Stone's Theorem we aim to regain cons…