Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

4589134178 · Jun 202019922001200920172026
48 results for Long Length

In this article, we prove that every arithmetic locally symmetric orbifold of classical type without Euclidean or compact factors has arbitrarily long arithmetic progressions in its primitive length spectrum. Moreover, we show the stronger property that every primitive length occurs in arbitrarily long arithmetic progr…

2016-02-04abs ↗pdf ↗

In this article, we investigate when the set of primitive geodesic lengths on a Riemannian manifold have arbitrarily long arithmetic progressions. We prove that in the space of negatively curved metrics, a metric having such arithmetic progressions is quite rare. We introduce almost arithmetic progressions, a coarsific…

2014-01-29abs ↗pdf ↗

Shorter adversarial prompts help protect LLMs from jailbreak attacks.

problem Protecting large language models from jailbreak attacks with long adversarial suffixes.
method Adversarial training on shorter adversarial suffixes to defend against longer adversarial suffixes.
result Aligning LLMs on shorter adversarial suffixes can effectively defend against jailbreak attacks with longer suffixes.

Much recent work has concerned sparse approximations to speed up the Gaussian process regression from the unfavorable O(n3) scaling in computational time to O(nm2). Thus far, work has concentrated on models with one covariance function. However, in many practical situations additive models with multiple covariance func…

2012-06-13abs ↗pdf ↗

The paper tackles long-context linear system identification with improved sample complexity bounds.

problem Identifying dynamical systems with long dependencies over fixed context windows.
method Established sample complexity bounds for systems with linear dependencies over a context window of length p.
result The learning process is not hindered by slow mixing properties in extended context windows.

Proving that next-token prediction makes language models generate coherent long documents.

problem Understanding why language models generate coherent documents despite focusing on next-token prediction.
method Proving the power of next-token prediction in learning longer-range structure using Recurrent Neural Networks (RNN).
result Optimizing next-token prediction in RNNs yields a model that closely approximates the training distribution, even for long-range coherence.

HGConv uses HRR to efficiently detect malware, outperforming existing methods.

problem Efficiently detecting malware with long sequences.
method Holographic Global Convolutional Networks (HGConv) utilizing Holographic Reduced Representations (HRR).
result Achieved state-of-the-art results on malware benchmarks.

Study curves evolving by gradient flow of elastic energy, proving existence, smoothing, and convergence.

problem Evolution of curves with fixed length and clamped boundary conditions.
method Negative L2L^2-gradient flow of elastic energy, existence, parabolic smoothing, constrained Lojasiewicz-Simon gradient inequality.
result Convergence to a critical point as time tends to infinity.

Transformer model predicts stock trends using technical data and sentiment analysis.

problem Lack of accurate long-term stock trend prediction using traditional models.
method Developed a Transformer-based model integrating technical stock data and sentiment analysis.
result Transformer model shows significant improvement in directional accuracy over RNNs, especially for longer sequence lengths.

We show there exist tunnel number one hyperbolic 3-manifolds with arbitrarily long unknotting tunnel. This provides a negative answer to an old question of Colin Adams.

2008-12-04abs ↗pdf ↗

We discuss algorithms for estimating the Shannon entropy h of finite symbol sequences with long range correlations. In particular, we consider algorithms which estimate h from the code lengths produced by some compression algorithm. Our interest is in describing their convergence with sequence length, assuming no limit…

2002-03-21abs ↗pdf ↗

The paper introduces combinatorial Calabi flows to find hyperbolic metrics on surfaces with boundary.

problem Finding hyperbolic metrics on surfaces with totally geodesic boundaries of given lengths.
method Introducing combinatorial Calabi flows and proving their long time existence and global convergence.
result Proves the long time existence and global convergence of combinatorial Calabi flow on surfaces with boundary.

Paper finds optimal shapes for minimizing average lengths of billiard trajectories in specific polygons.

problem Finding optimal shapes to minimize the average length of billiard trajectories.
method Used techniques from Teichmüller theory.
result Optimal shapes minimize average lengths of billiard trajectories in specific polygons.

Music relies heavily on repetition to build structure and meaning. Self-reference occurs on multiple timescales, from motifs to phrases to reusing of entire sections of music, such as in pieces with ABA structure. The Transformer (Vaswani et al., 2017), a sequence model based on self-attention, has achieved compelling …

2018-09-12abs ↗pdf ↗

Research shows a significant increase in stay lengths for digital nomads in the U.S. during and after the pandemic.

problem Shifts in stay lengths for digital nomads during and after the pandemic.
method Analysis of Airbnb reservations data from 2019-2024, using statistical models to quantify changes.
result Mean stay lengths increased from 3.68 to 4.36 nights, stabilizing near 4.07 after 2021, indicating a 10% increase from pre-pandemic levels.

Preformer improves Transformer for long-term time series forecasting.

problem Transformer's quadratic complexity and lack of context-awareness for long-term forecasting.
method Introduces Multi-Scale Segment-Correlation mechanism for efficient time series segmentation and context-aware attention.
result Preformer outperforms other Transformer-based methods in long-term time series forecasting.

ESNs with transfer learning predict long-term chaotic patterns in spatiotemporal dynamical systems.

problem Predicting long-term statistical patterns of spatiotemporally chaotic dynamical systems.
method Echo state networks (ESNs) with transfer learning.
result ESNs with transfer learning accurately predict long-term statistical properties of spatiotemporally chaotic PDEs.

Meta-learning framework improves short utterance speaker recognition.

problem Poor performance of existing models with short utterances.
method Prototypical Networks with support and query sets, enforcing classification against entire training set.
result Significant performance gains on VoxCeleb datasets.

Hrrformer uses HRR to speed up self-attention for long sequences.

problem Infeasibility of using transformers with very long sequence lengths.
method Re-cast self-attention using Holographic Reduced Representations (HRR).
result Achieves near state-of-the-art accuracy with O(THlogH)\mathcal{O}(T H \log H) time complexity and O(TH)\mathcal{O}(T H) space complexity.

Study finds phase transition in context-sensitive language model with short-range interactions.

problem Understanding phase transitions in language models with short-range interactions.
method Constructed a random language model with short-range interactions and investigated its statistical properties.
result Phase transition occurs in context-sensitive language models with constant context length.

Compact Recurrent Transformer (CRT) improves Transformer efficiency for long sequences.

problem Efficiently scaling Transformer architecture to long sequences with limited compute resources.
method Combines shallow Transformer models with recurrent neural networks and persistent memory.
result CRT achieves comparable or superior performance to full-length Transformers with shorter segments and reduced FLOPs.

Rough Transformers improve efficiency for medical time-series data.

problem Efficiently modeling irregularly sampled, long-range time-series data.
method Introducing Rough Transformers, a Transformer variant with continuous-time representations and multi-view signature attention.
result Rough Transformers outperform vanilla Transformers while using less computational resources.

Transformers become faster by linearizing self-attention.

problem Quadratic complexity of transformers makes them slow for long sequences.
method Expressed self-attention as a linear dot-product and used matrix product associativity to reduce complexity.
result Linear transformers are up to 4000x faster on long sequences.

The article is devoted to investigating the application of aggregating algorithms to the problem of the long-term forecasting. We examine the classic aggregating algorithms based on the exponential reweighing. For the general Vovk's aggregating algorithm we provide its generalization for the long-term forecasting. For …

2018-03-18abs ↗pdf ↗

New model predicts ICU patients' stay duration efficiently.

problem Efficient ICU bed allocation under resource constraints.
method Temporal Pointwise Convolution (TPC) model combining temporal and pointwise convolutions.
result Significant performance improvements over LSTM and Transformer models.

New model predicts ICU patient stays more accurately.

problem Efficient ICU bed allocation under resource constraints.
method Temporal Pointwise Convolutional Networks (TPC) combining temporal and pointwise convolutions.
result Significant performance improvements over LSTM and Transformer models.

Transformers are powerful sequence models, but require time and memory that grows quadratically with the sequence length. In this paper we introduce sparse factorizations of the attention matrix which reduce this to O(nn)O(n \sqrt{n}). We also introduce a) a variation on architecture and initialization to train deeper net…

2019-04-23abs ↗pdf ↗

Study critical exponents on hyperbolic surfaces with long boundaries using Weil-Petersson measures.

problem Analyzing critical exponents on hyperbolic surfaces with long boundaries.
method Using spine graph construction and comparing normalized Weil-Petersson and Kontsevich measures.
result Asymptotic convergence-in-mean result of normalized Weil-Petersson measures to normalized Kontsevich measures.

We prove the absence of a universal diameter bound on lengths of curves in a sweep-out of a Riemannian 2-sphere. If such bound existed it would yield a simple proof of existence of short geodesic segments and closed geodesics on a sphere of small diameter.

2011-05-31abs ↗pdf ↗

D-LinOSS models learn to dissipate energy, improving performance on long-range tasks.

problem Representational limitations of LinOSS models in long-range reasoning.
method Introducing Damped Linear Oscillatory State-Space models (D-LinOSS) that learn to dissipate latent state energy on arbitrary time scales.
result D-LinOSS consistently outperforms previous LinOSS methods on long-range learning tasks, achieving faster convergence and reducing hyperparameter search space.

Algorithm learns linear systems from partial observations with near-optimal rate.

problem Identifying linear dynamical systems from partial observations, especially those with long-term memory.
method Multi-scale low-rank approximation using SVD on Hankel matrices of increasing sizes, combined with Fourier domain concentration bounds.
result Near-optimal rate of $\widetilde O\left(\sqrt\frac{d}{T} ight)$ in H2\mathcal{H}_2 error, with logarithmic dependence on memory length.

Study validates Lillo-Mike-Farmer model predicting financial market long-range correlations.

problem Quantifying long-range correlations in financial markets.
method Analyzed nine years of market data to classify traders as order-splitting or random, measured metaorder-length distributions, and compared to LMF model predictions.
result Agreement between LMF model predictions and actual data, validating the model.

The FitzHugh-Nagumo equation provides a simple mathematical model of cardiac tissue as an excitable medium hosting spiral wave vortices. Here we present extensive numerical simulations studying long-term dynamics of knotted vortex string solutions for all torus knots up to crossing number 11. We demonstrate that FitzHu…

2017-06-20abs ↗pdf ↗

Global Memory Augmentation (GMAT) improves Transformer performance on long documents.

problem Large memory requirements of Transformer pairwise dot-product attention for long sequences.
method Integrates a dense global memory of length M into sparse Transformer blocks.
result Significant improvement on various tasks, including synthetic tasks, masked language modeling, and reading comprehension.

Adaptive beamforming collapses in highly non-stationary environments, but the Universal Switching Beamformer resolves this by dynamically adjusting memory length.

problem Adaptive beamforming performance degrades in highly non-stationary environments.
method Integrating sequential prediction into the beamforming architecture.
result The USB achieves agility and precision in tracking highly non-stationary scenes.