New algorithm reduces regret bounds for Bayesian optimization with unknown hyperparameters.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
For curves of prescribed length embedded into the unit disc in two dimensions, we obtain scaling results for the minimal elastic energy as the length just exceeds and in the large length limit. In the small excess length case, we prove convergence to a fourth order obstacle type problem with integral constraint on…
Much recent work has concerned sparse approximations to speed up the Gaussian process regression from the unfavorable O(n3) scaling in computational time to O(nm2). Thus far, work has concentrated on models with one covariance function. However, in many practical situations additive models with multiple covariance func…
Unified framework for critical scaling of inverse temperature in self-attention.
The scaling properties of the time series of asset prices and trading volumes of stock markets are analysed. It is shown that similarly to the asset prices, the trading volume data obey multi-scaling length-distribution of low-variability periods. In the case of asset prices, such scaling behaviour can be used for risk…
Mamba struggles with long context lengths, but spectrum scaling improves performance.
Paper reconciles two methods of describing Riemannian spaces.
New methods for Markov Blanket discovery using MML outperform existing approaches.
New insights into how depth and width affect in-context learning in deep models.
We discuss algorithms for estimating the Shannon entropy h of finite symbol sequences with long range correlations. In particular, we consider algorithms which estimate h from the code lengths produced by some compression algorithm. Our interest is in describing their convergence with sequence length, assuming no limit…
Moduli spaces of hyperbolic surfaces with geodesic boundary components of fixed lengths may be endowed with a symplectic structure via the Weil-Petersson form. We show that, as the boundary lengths are sent to infinity, the Weil-Petersson form converges to a piecewise linear form first defined by Kontsevich. The proof …
The question of how best to estimate a continuous probability density from finite data is an intriguing open problem at the interface of statistics and physics. Previous work has argued that this problem can be addressed in a natural way using methods from statistical field theory. Here I describe new results that allo…
Recurrent auto-encoder model summarises sequential data through an encoder structure into a fixed-length vector and then reconstructs the original sequence through the decoder structure. The summarised vector can be used to represent time series features. In this paper, we propose relaxing the dimensionality of the dec…
Randomized positional encodings boost transformer performance on longer sequences.
Deep architecture such as hierarchical semi-Markov models is an important class of models for nested sequential data. Current exact inference schemes either cost cubic time in sequence length, or exponential time in model depth. These costs are prohibitive for large-scale problems with arbitrary length and depth. In th…
We provide a numerically robust and fast method capable of exploiting the local geometry when solving large-scale stochastic optimisation problems. Our key innovation is an auxiliary variable construction coupled with an inverse Hessian approximation computed using a receding history of iterates and gradients. It is th…
The Efficient Global Optimization (EGO) algorithm uses a conditional Gaus-sian Process (GP) to approximate an objective function known at a finite number of observation points and sequentially adds new points which maximize the Expected Improvement criterion according to the GP. The important factor that controls the e…
Based on empirical financial time-series, we show that the "silence-breaking" probability follows a super-universal power law: the probability of observing a large movement is inversely proportional to the length of the on-going low-variability period. Such a scaling law has been previously predicted theoretically [R. …
New language model shows context length impacts generation quality and reasoning ability.
Bayesian Optimization (BO) has become a core method for solving expensive black-box optimization problems. While much research focussed on the choice of the acquisition function, we focus on online length-scale adaption and the choice of kernel function. Instead of choosing hyperparameters in view of maximum likelihood…
NUTS mixing time scales as d^(1/4) for Gaussian distributions.
Recurrent Neural Networks (RNNs) are among the most popular models in sequential data analysis. Yet, in the foundational PAC learning language, what concept class can it learn? Moreover, how can the same recurrent unit simultaneously learn functions from different input tokens to different output tokens, without affect…
Accelerates signature kernel computation for sequences.
Preformer improves Transformer for long-term time series forecasting.
We prove a rigidity theorem for the geometry of the unit ball in random subspaces of the scl norm in B_1^H of a free group. In a free group F of rank k, a random word w of length n (conditioned to lie in [F,F]) has scl(w)=log(2k-1)n/6log(n) + o(n/log(n)) with high probability, and the unit ball in a subspace spanned by…
The paper defines and studies discrete p-density and compression-radius profiles of lattice knots.
We investigate quotation and transaction activities in the foreign exchange market for every week during the period of June 2007 to December 2010. A scaling relationship between the mean values of number of quotations (or number of transactions) for various currency pairs and the corresponding standard deviations holds…
Fast, fully-automated histograms for large data sets.
Sine activation functions enable two-layer neural networks to learn modular addition more efficiently.
The study examines how extra compute during testing affects the performance of large language models.
Defines magnitude for length spaces with measures, agreeing with finite spaces' magnitude.
We have discovered 12 independent new empirical scaling laws in foreign exchange data-series that hold for close to three orders of magnitude and across 13 currency exchange rates. Our statistical analysis crucially depends on an event-based approach that measures the relationship between different types of events. The…
Improves Gaussian process factor models for multi-population recordings.
New insights show embedding lengths correlate with semantic properties.
Algorithm learns linear systems from partial observations with near-optimal rate.
Proposes a new tail risk measure based on the most probable maximum risk event size.
The scaling properties of oil price fluctuations are described as a non-stationary stochastic process realized by a time series of finite length. An original model is used to extract the scaling exponent of the fluctuation functions within a non-stationary process formulation. It is shown that, when returns are measure…
New method constructs potential functions for Kähler-Einstein metrics.
Analyzing and interpreting time-dependent stochastic data requires accurate and robust density estimation. In this paper we extend the concept of normalizing flows to so-called temporal Normalizing Flows (tNFs) to estimate time dependent distributions, leveraging the full spatio-temporal information present in the data…
The study examines correlations of logarithms of integers at different scalings.
In current clinical practices, electroencephalograms (EEG) are reviewed and analyzed by trained neurologists to provide supports for therapeutic decisions. Manual reviews can be laborious and error prone. Automatic and accurate seizure/non-seizure classification methods are desirable. A critical challenge is that seizu…
Study on geodesics on high genus expander surfaces, proving filling and non-simple properties.
Slice Sampling has emerged as a powerful Markov Chain Monte Carlo algorithm that adapts to the characteristics of the target distribution with minimal hand-tuning. However, Slice Sampling's performance is highly sensitive to the user-specified initial length scale hyperparameter and the method generally struggles with …
We highlight a pitfall when applying stochastic variational inference to general Bayesian networks. For global random variables approximated by an exponential family distribution, natural gradient steps, commonly starting from a unit length step size, are averaged to convergence. This useful insight into the scaling of…
Theory predicts neural scaling exponents from language statistics.
New kernel interprets 3D anisotropic data with rotations and improved predictions.
Liouville's theorem says that in dimension greater than two, all conformal maps are Möbius transformations. We prove an analogous statement about simplicial complexes, where two simplicial complexes are considered discretely conformally equivalent if they are combinatorially equivalent and the lengths of corresponding …
Research examines correlations of complex logarithms of lattice points, showing level repulsion and Poissonian behavior.