Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

5.1%10.2%15.4%20.5% · May 202619922001200920172026
48 results for scaling strategies

Paper proposes MSSDDPG for better financial trading strategies.

problem Extracting accurate features from noisy, non-stationary financial time series.
method Multi-scale stroke deep deterministic policy gradient reinforcement learning model (MSSDDPG).
result MSSDDPG outperforms other strategies in China's CSI 300 and SSE Composite.

New method trains shallow neural networks with subquadratic width scaling.

problem Training shallow neural networks with optimal width scaling.
method Polyak-Lojasiewicz condition, smoothness, standard data assumptions, random matrix theory.
result Subquadratic scaling on network width with standard initialization strategies.

In this note, we study a class of stochastic control problems where the optimal strategies are described by two parameters. These include a subset of singular control, impulse control, and two-player stochastic games. The parameters are first chosen by the two continuous/smooth fit conditions, and then the optimality o…

2016-05-17abs ↗pdf ↗

Study optimal periodic dividend strategies for risky businesses with transaction costs.

problem Optimal periodic dividend strategies for spectrally positive Lévy risk processes with fixed transaction costs.
method Investigates periodic (bu,bl)(b_u,b_l) strategies for a Poisson arrival process of decision times.
result A periodic (bu,bl)(b_u,b_l) strategy is optimal with lump sum dividends net of transaction costs.

Different investment strategies are adopted in short-term and long-term depending on the time scales, even though time scales are adhoc in nature. Empirical mode decomposition based Hurst exponent analysis and variance technique have been applied to identify the time scales for short-term and long-term investment from …

2019-06-13abs ↗pdf ↗

New voting strategies show committee-based consensus can scale efficiently.

problem Ensuring honest committees in committee-based consensus protocols.
method Empirical analysis of simpler voting strategies and their convergence to optimality.
result Simpler voting strategies converge to optimality exponentially quickly, ensuring robustness and efficiency.

This work explores test-time scaling strategies for LLMs, improving sample efficiency and expressiveness.

problem Understanding the sample efficiency and expressiveness of test-time scaling strategies for LLMs.
method Established separation and expressiveness results for self-consistency, best-of-nn, and self-correction strategies.
result Self-correction enables Transformers to simulate online learning over multiple tasks without prior knowledge.

Generalized algorithm for translation and scale-invariant prediction.

problem Sequential prediction with expert advice, focusing on translation and scale invariance.
method Designing a generalized online algorithm using the universal prediction perspective to compete against a generic class of expert selection strategies.
result No preliminary knowledge of loss sequences is required; performance bounds are stable under arbitrary scalings and translations.

We determine the critical batch size for large language models and find it scales with data size, not model size.

problem Determining the optimal batch size for large-scale model training.
method We propose a measure of critical batch size, pre-trained models, and systematic hyper-parameter sweeps.
result The critical batch size scales primarily with data size, not model size.

This paper tackles hyperparameter tuning for large-scale kernel ridge regression.

problem Hyperparameter tuning is crucial but often left to users, hindering efficiency and usability.
method Proposes a complexity regularization criterion based on a data-dependent penalty for efficient optimization.
result Demonstrates the benefit of the proposed approach through extensive empirical evaluation.

Paper proposes efficient GCN learning method for limited data.

problem Learning GCNs from data with extremely limited annotations.
method Adaptive sampling strategy and model compression.
result Cut down annotation requirement by 90% and compress parameters 6x.

Initializing the weights and the biases is a key part of the training process of a neural network. Unlike the subsequent optimization phase, however, the initialization phase has gained only limited attention in the literature. In this paper we discuss some consequences of commonly used initialization strategies for va…

2019-03-27abs ↗pdf ↗

Investment strategies for rank-dependent utility agents are derived in a continuous-time market.

problem Time inconsistency in rank-dependent utility models.
method Study of consistent planners seeking intra-personal equilibrium strategies.
result Explicit final wealth profile replicating equilibrium strategies, with scaling function derived.

Research provides explicit NPV expressions for double barrier strategies.

problem Calculating expected NPVs of double barrier strategies for regular diffusions.
method Explicit expression using bivariate q-scale function with perturbation technique.
result Explicit expressions for expected NPVs are derived for certain cases.

This study reveals the critical role of scale vectors in large language models, improving optimization and expressivity.

problem Understanding and optimizing the scale vectors in large language models.
method Systematic study of scale vectors from expressivity, optimization, and architectural perspectives; theoretical and empirical analysis of weight decay; proposing and evaluating improvements.
result Scale vectors improve optimization through a self-amplifying preconditioning effect and are beneficial for expressivity in certain architectures.

In this paper, we revisit the optimal periodic dividend problem, in which dividend payments can only be made at the jump times of an independent Poisson process. In the dual (spectrally positive Lévy) model, recent results have shown the optimality of a periodic barrier strategy, which pays dividends at Poissonian divi…

2017-08-04abs ↗pdf ↗

The paper analyzes LETF option markets using moneyness scaling to find statistical arbitrage opportunities.

problem Statistical discrepancies between levered and unlevered ETF option implied volatility smiles.
method Bootstrap uniform confidence bands, dynamic semiparametric factor model, moneyness scaling, Heston stochastic volatility.
result Trading opportunities exist on LETF market, and a statistical arbitrage strategy generates positive returns.

AdAdaGrad optimizes batch sizes for deep learning models, reducing the generalization gap.

problem The generalization gap between large-batch and small-batch training in deep learning.
method AdAdaGrad introduces adaptive batch size strategies derived from adaptive sampling methods.
result AdAdaGradNorm converges to a first-order stationary point with a rate of O(1/K) in K iterations.

The paper studies scaling limits of hedging prices in financial models.

problem Scaling limits of exponential utility indifference prices in financial models.
method Formulated dual problem as stochastic control, solved HJB equation for upper bound, used duality result for lower bound.
result Represented scaling limit in terms of specific relative entropy and constructed asymptotic optimal hedging strategies.

Enhances out-of-domain calibration of neural networks.

problem Improving calibration performance of deep neural networks in out-of-domain settings.
method Consistency-guided temperature scaling (CTS) that considers style and content consistency.
result Significantly enhances out-of-domain calibration performance.

For large scale on-line inference problems the update strategy is critical for performance. We derive an adaptive scan Gibbs sampler that optimizes the update frequency by selecting an optimum mini-batch size. We demonstrate performance of our adaptive batch-size Gibbs sampler by comparing it against the collapsed Gibb…

2018-01-27abs ↗pdf ↗

Paper proposes a new framework for combining investment strategies without market-specific assumptions.

problem Lack of a distribution-free and consistent preference framework for decision-making in combining investment strategies.
method Introduces a novel framework for decision-making in combining strategies, free from market conditions and statistical assumptions.
result Proposed strategies outperform individual component strategies in long-term wealth accumulation, with small tradeoffs in Sharpe ratios.

Optimizes trading strategies with price impact, predictable returns, and stochastic volatility.

problem Dynamic portfolio optimization under complex market conditions.
method Multi-scale volatility expansion, singular and regular perturbations, asymptotic approximations.
result Improved portfolio strategy with reduced profit and loss (PnL) through corrections for small price impact.

State-of-the-art methods for Convolutional Sparse Coding usually employ Fourier-domain solvers in order to speed up the convolution operators. However, this approach is not without shortcomings. For example, Fourier-domain representations implicitly assume circular boundary conditions and make it hard to fully exploit …

2019-08-31abs ↗pdf ↗

Best-of-Majority improves inference performance in Pass@kk settings.

problem Inference in difficult tasks often underperforms with single-shot selection methods.
method Combining majority voting and Best-of-N, Best-of-Majority restricts candidates to high-frequency responses.
result Best-of-Majority achieves minimax optimal regret and outperforms other methods.

SHAKE-GNN scales GNNs for large graphs with multi-scale representations.

problem Scaling Graph Neural Networks (GNNs) to large graphs.
method SHAKE-GNN uses a hierarchy of Kirchhoff Forests for stochastic multi-resolution graph decompositions.
result SHAKE-GNN achieves competitive performance on large-scale graph classification benchmarks.

This paper studies the optimal dividend problem with capital injection under the constraint that the cumulative dividend strategy is absolutely continuous. We consider an open problem of the general spectrally negative case and derive the optimal solution explicitly using the fluctuation identities of the refracted-ref…

2017-09-19abs ↗pdf ↗

Local GP approach improves simulation efficiency for large datasets.

problem High computational cost of traditional Gaussian processes for large-scale simulations.
method Hybridizes global and local GP approximations with strategic placement of inducing points.
result Local inducing points enhance accuracy and computational efficiency.