Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

83166249332 · Jun 202019922001200920172026
48 results for Length Scale

New algorithm reduces regret bounds for Bayesian optimization with unknown hyperparameters.

problem Optimizing black-box functions with unknown hyperparameters, especially length scale.
method Length Scale Balancing (LB) - aggregating multiple surrogate models with varying length scales.
result LB achieves a regret bound only logaritically away from the oracle algorithm.

For curves of prescribed length embedded into the unit disc in two dimensions, we obtain scaling results for the minimal elastic energy as the length just exceeds 2π and in the large length limit. In the small excess length case, we prove convergence to a fourth order obstacle type problem with integral constraint on…

2018-12-12abs ↗pdf ↗

Much recent work has concerned sparse approximations to speed up the Gaussian process regression from the unfavorable O(n3) scaling in computational time to O(nm2). Thus far, work has concentrated on models with one covariance function. However, in many practical situations additive models with multiple covariance func…

2012-06-13abs ↗pdf ↗

The scaling properties of the time series of asset prices and trading volumes of stock markets are analysed. It is shown that similarly to the asset prices, the trading volume data obey multi-scaling length-distribution of low-variability periods. In the case of asset prices, such scaling behaviour can be used for risk…

2005-01-13abs ↗pdf ↗

New insights into how depth and width affect in-context learning in deep models.

problem Understanding how various resources impact in-context learning in deep models.
method Analyzed linear regression in a deep linear self-attention model, varying resources like depth, width, context length, and training steps.
result Increasing depth improves in-context learning even at infinite context length, contrary to previous findings.

We discuss algorithms for estimating the Shannon entropy h of finite symbol sequences with long range correlations. In particular, we consider algorithms which estimate h from the code lengths produced by some compression algorithm. Our interest is in describing their convergence with sequence length, assuming no limit…

2002-03-21abs ↗pdf ↗

Moduli spaces of hyperbolic surfaces with geodesic boundary components of fixed lengths may be endowed with a symplectic structure via the Weil-Petersson form. We show that, as the boundary lengths are sent to infinity, the Weil-Petersson form converges to a piecewise linear form first defined by Kontsevich. The proof …

2010-10-20abs ↗pdf ↗

Randomized positional encodings boost transformer performance on longer sequences.

problem Transformers struggle with generalizing to sequences of arbitrary length.
method Introduced randomized positional encodings that simulate longer sequences and randomly select positions.
result Randomized positional encodings increase test accuracy by 12.0% on average for sequences of unseen length.

Deep architecture such as hierarchical semi-Markov models is an important class of models for nested sequential data. Current exact inference schemes either cost cubic time in sequence length, or exponential time in model depth. These costs are prohibitive for large-scale problems with arbitrary length and depth. In th…

2014-08-06abs ↗pdf ↗

The Efficient Global Optimization (EGO) algorithm uses a conditional Gaus-sian Process (GP) to approximate an objective function known at a finite number of observation points and sequentially adds new points which maximize the Expected Improvement criterion according to the GP. The important factor that controls the e…

2016-03-08abs ↗pdf ↗

Based on empirical financial time-series, we show that the "silence-breaking" probability follows a super-universal power law: the probability of observing a large movement is inversely proportional to the length of the on-going low-variability period. Such a scaling law has been previously predicted theoretically [R. …

2008-12-24abs ↗pdf ↗

New language model shows context length impacts generation quality and reasoning ability.

problem Analyzing the impact of context length and reasoning on autoregressive generation.
method Introduced synthetic hierarchical languages, used an exact k-gram ansatz, derived asymptotic predictions, and validated empirically.
result Reasoning models with limited context can generate sequences from the true language, improving exponentially over standard models.

NUTS mixing time scales as d^(1/4) for Gaussian distributions.

problem Improving the efficiency of the No-U-Turn Sampler (NUTS) for Gaussian distributions.
method Coupling argument leveraging geometric structure of Gaussian concentration, uniformity analysis of NUTS transitions.
result The mixing time of NUTS scales as d^(1/4) for Gaussian distributions, up to logarithmic factors.

Recurrent Neural Networks (RNNs) are among the most popular models in sequential data analysis. Yet, in the foundational PAC learning language, what concept class can it learn? Moreover, how can the same recurrent unit simultaneously learn functions from different input tokens to different output tokens, without affect…

2019-02-04abs ↗pdf ↗

Preformer improves Transformer for long-term time series forecasting.

problem Transformer's quadratic complexity and lack of context-awareness for long-term forecasting.
method Introduces Multi-Scale Segment-Correlation mechanism for efficient time series segmentation and context-aware attention.
result Preformer outperforms other Transformer-based methods in long-term time series forecasting.

We prove a rigidity theorem for the geometry of the unit ball in random subspaces of the scl norm in B_1^H of a free group. In a free group F of rank k, a random word w of length n (conditioned to lie in [F,F]) has scl(w)=log(2k-1)n/6log(n) + o(n/log(n)) with high probability, and the unit ball in a subspace spanned by…

2011-04-10abs ↗pdf ↗

The paper defines and studies discrete p-density and compression-radius profiles of lattice knots.

problem Understanding geometric properties of lattice knots.
method Develops a framework for discrete p-density and compression-radius profiles of lattice knots, studying them on length-filtered sets and finite move-graph exploration.
result Density and compression-radius values are not monotone, illustrating distinct optimization problems.

Sine activation functions enable two-layer neural networks to learn modular addition more efficiently.

problem Learning modular addition with two-layer neural networks.
method Introduced and analyzed sine activation functions, providing theoretical and empirical evidence.
result Sine activation functions allow for constant-width network realizations of modular addition, whereas ReLU networks require linear width scaling.

The study examines how extra compute during testing affects the performance of large language models.

problem Understanding the conditions under which test-time scaling improves model performance.
method An in-context weight prediction task for linear regression was used to train transformers. The performance was analyzed under varying levels of test-time compute.
result Training transformers on diverse, relevant, and hard tasks leads to the best performance for test-time scaling.

Improves Gaussian process factor models for multi-population recordings.

problem Cubic runtime scaling with trial length and group number limits application to large-scale recordings.
method Two approximate approaches: inducing variables and frequency domain.
result Achieved orders of magnitude speed-up with minimal statistical performance impact.

New insights show embedding lengths correlate with semantic properties.

problem Contrastive embedding norms ignore embedding magnitudes but correlate with semantic properties.
method Formal theoretical framework and analysis of optimization dynamics.
result Embedding lengths encode semantic information as a byproduct of training.

Algorithm learns linear systems from partial observations with near-optimal rate.

problem Identifying linear dynamical systems from partial observations, especially those with long-term memory.
method Multi-scale low-rank approximation using SVD on Hankel matrices of increasing sizes, combined with Fourier domain concentration bounds.
result Near-optimal rate of $\widetilde O\left(\sqrt\frac{d}{T} ight)$ in H2\mathcal{H}_2 error, with logarithmic dependence on memory length.

Proposes a new tail risk measure based on the most probable maximum risk event size.

problem Current risk measures like VaR and ES are limited in their applicability and require specifying a confidence level.
method Develops a new risk measure called MPMR that does not require a confidence level and scales with the length of the time interval.
result The new risk measure, MPMR, scales with the number of observations by a power law, allowing for reliable estimations of long-term risks based on short-term estimations.

The scaling properties of oil price fluctuations are described as a non-stationary stochastic process realized by a time series of finite length. An original model is used to extract the scaling exponent of the fluctuation functions within a non-stationary process formulation. It is shown that, when returns are measure…

2008-09-06abs ↗pdf ↗

Analyzing and interpreting time-dependent stochastic data requires accurate and robust density estimation. In this paper we extend the concept of normalizing flows to so-called temporal Normalizing Flows (tNFs) to estimate time dependent distributions, leveraging the full spatio-temporal information present in the data…

2019-12-19abs ↗pdf ↗

The study examines correlations of logarithms of integers at different scalings.

problem Analyzing pair correlations of logarithms of integers at various scalings.
method Examined correlations of logarithms of positive integers at different scalings, proving the existence of pair correlation functions.
result Level repulsion at linear scaling, total loss of mass at superlinear scalings, and Poissonian behavior at sublinear scalings.

New kernel interprets 3D anisotropic data with rotations and improved predictions.

problem Capturing rotated anisotropy in 3D spatial fields.
method Introduces a Lie-algebraic kernel with three principal length-scales and an explicit rotation.
result Posterior recovers rotated anisotropy and improves prediction over axis-aligned kernels.

Liouville's theorem says that in dimension greater than two, all conformal maps are Möbius transformations. We prove an analogous statement about simplicial complexes, where two simplicial complexes are considered discretely conformally equivalent if they are combinatorially equivalent and the lengths of corresponding …

2019-11-03abs ↗pdf ↗

Research examines correlations of complex logarithms of lattice points, showing level repulsion and Poissonian behavior.

problem Analyzing correlations of complex logarithms of lattice points.
method Proving existence of pair correlation functions and examining behavior at various scalings.
result Level repulsion observed at linear scaling, Poissonian behavior at sublinear scalings.