Derives a family of hyperparameter scaling strategies for neural networks.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper proposes MSSDDPG for better financial trading strategies.
Deep RL strategies outperform classical models in trading.
New method trains shallow neural networks with subquadratic width scaling.
This paper introduces a new market making approach using scaled beta distributions.
In this note, we study a class of stochastic control problems where the optimal strategies are described by two parameters. These include a subset of singular control, impulse control, and two-player stochastic games. The parameters are first chosen by the two continuous/smooth fit conditions, and then the optimality o…
The paper revisits the investment simulation based on strategies exhibited by Generalized (m,2)-Zipf law to present an interesting characterization of the wildness in financial time series. The investigations of dominant strategies on each specific time series shows that longer words dominant in larger time scale exhib…
The optimal dividend problem by De Finetti (1957) has been recently generalized to the spectrally negative Lévy model where the implementation of optimal strategies draws upon the computation of scale functions and their derivatives. This paper proposes a phase-type fitting approximation of the optimal strategy. We con…
Study optimal periodic dividend strategies for risky businesses with transaction costs.
Different investment strategies are adopted in short-term and long-term depending on the time scales, even though time scales are adhoc in nature. Empirical mode decomposition based Hurst exponent analysis and variance technique have been applied to identify the time scales for short-term and long-term investment from …
New voting strategies show committee-based consensus can scale efficiently.
We explore the use of Evolution Strategies (ES), a class of black box optimization algorithms, as an alternative to popular MDP-based RL techniques such as Q-learning and Policy Gradients. Experiments on MuJoCo and Atari show that ES is a viable solution strategy that scales extremely well with the number of CPUs avail…
This work explores test-time scaling strategies for LLMs, improving sample efficiency and expressiveness.
Generalized algorithm for translation and scale-invariant prediction.
The omnipresence of deep learning architectures such as deep convolutional neural networks (CNN)s is fueled by the synergistic combination of ever-increasing labeled datasets and specialized hardware. Despite the indisputable success, the reliance on huge amounts of labeled data and specialized hardware can be a limiti…
We determine the critical batch size for large language models and find it scales with data size, not model size.
Empirical studies indicate the presence of multi-scales in the volatility of underlying assets: a fast-scale on the order of days and a slow-scale on the order of months. In our previous works, we have studied the portfolio optimization problem in a Markovian setting under each single scale, the slow one in [Fouque and…
We revisit the stochastic limited-memory BFGS (L-BFGS) algorithm. By proposing a new framework for the convergence analysis, we prove improved convergence rates and computational complexities of the stochastic L-BFGS algorithms compared to previous works. In addition, we propose several practical acceleration strategie…
This paper tackles hyperparameter tuning for large-scale kernel ridge regression.
Paper proposes efficient GCN learning method for limited data.
Using daily returns of the S&P 500 stocks from 2001 to 2011, we perform a backtesting study of the portfolio optimization strategy based on the extreme risk index (ERI). This method uses multivariate extreme value theory to minimize the probability of large portfolio losses. With more than 400 stocks to choose from, ou…
Initializing the weights and the biases is a key part of the training process of a neural network. Unlike the subsequent optimization phase, however, the initialization phase has gained only limited attention in the literature. In this paper we discuss some consequences of commonly used initialization strategies for va…
A key problem in location-based modeling and forecasting lies in identifying suitable spatial and temporal resolutions. In particular, judicious spatial partitioning can play a significant role in enhancing the performance of location-based forecasting models. In this work, we investigate two widely used tessellation s…
Investment strategies for rank-dependent utility agents are derived in a continuous-time market.
This paper studies stability of the exponential utility maximization when there are small variations on agent's utility function. Two settings are considered. First, in a general semimartingale model where random endowments are present, a sequence of utilities defined on R converges to the exponential utility. Under a …
We introduce various quantitative and mathematical definitions for price momentum of financial instruments. The price momentum is quantified with velocity and mass concepts originated from the momentum in physics. By using the physical momentum of price as a selection criterion, the weekly contrarian strategies are imp…
Paper develops AGLD for MCMC with bounds for various data access strategies.
Supervised dimensionality reduction strategies have been of great interest. However, current supervised dimensionality reduction approaches are difficult to scale for situations characterized by large datasets given the high computational complexities associated with such methods. While stochastic approximation strateg…
Research provides explicit NPV expressions for double barrier strategies.
This study reveals the critical role of scale vectors in large language models, improving optimization and expressivity.
In this paper, we revisit the optimal periodic dividend problem, in which dividend payments can only be made at the jump times of an independent Poisson process. In the dual (spectrally positive Lévy) model, recent results have shown the optimality of a periodic barrier strategy, which pays dividends at Poissonian divi…
Deploying deep learning (DL) models across multiple compute devices to train large and complex models continues to grow in importance because of the demand for faster and more frequent training. Data parallelism (DP) is the most widely used parallelization strategy, but as the number of devices in data parallel trainin…
The paper analyzes LETF option markets using moneyness scaling to find statistical arbitrage opportunities.
Large-scale Gaussian process inference has long faced practical challenges due to time and space complexity that is superlinear in dataset size. While sparse variational Gaussian process models are capable of learning from large-scale data, standard strategies for sparsifying the model can prevent the approximation of …
AdAdaGrad optimizes batch sizes for deep learning models, reducing the generalization gap.
The paper studies scaling limits of hedging prices in financial models.
Enhances out-of-domain calibration of neural networks.
For large scale on-line inference problems the update strategy is critical for performance. We derive an adaptive scan Gibbs sampler that optimizes the update frequency by selecting an optimum mini-batch size. We demonstrate performance of our adaptive batch-size Gibbs sampler by comparing it against the collapsed Gibb…
Paper proposes a new framework for combining investment strategies without market-specific assumptions.
Optimizes trading strategies with price impact, predictable returns, and stochastic volatility.
State-of-the-art methods for Convolutional Sparse Coding usually employ Fourier-domain solvers in order to speed up the convolution operators. However, this approach is not without shortcomings. For example, Fourier-domain representations implicitly assume circular boundary conditions and make it hard to fully exploit …
Best-of-Majority improves inference performance in Pass@ settings.
Benefitting from large-scale training datasets and the complex training network, Convolutional Neural Networks (CNNs) are widely applied in various fields with high accuracy. However, the training process of CNNs is very time-consuming, where large amounts of training samples and iterative operations are required to ob…
SHAKE-GNN scales GNNs for large graphs with multi-scale representations.
Avanzi et al. (2016) recently studied an optimal dividend problem where dividends are paid both periodically and continuously with different transaction costs. In the Brownian model with Poissonian periodic dividend payment opportunities, they showed that the optimal strategy is either of the pure-continuous, pure-peri…
This paper studies the optimal dividend problem with capital injection under the constraint that the cumulative dividend strategy is absolutely continuous. We consider an open problem of the general spectrally negative case and derive the optimal solution explicitly using the fluctuation identities of the refracted-ref…
Local GP approach improves simulation efficiency for large datasets.
Adapts to estimate functions from noisy ERT data.