Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

98195293390 · Jun 202019922001200920182026
48 results for random sequences

Develops methods to construct exchangeable sequences of random multisets.

problem Creating models for random multisets with unknown base measures.
method Uses exchangeable sequences of point processes and conditional-i.i.d. negative binomial processes.
result Provides constructions for negative binomial processes with random base measures.

SVM and N-best algorithm classify microbial marker clades from genome sequences.

problem Classifying microbial clades from genome sequences, especially new species.
method Support vector machine (SVM) with N-best algorithm, time series feature extraction, random fragment generation, k-mer size selection.
result Recognition accuracy rates above 28% in top-1 candidate, above 91% in top-10 candidate.

Study Bergman kernels and zero distributions of random sections on Kähler manifolds.

problem Asymptotic distribution of common zeros of random sections on Kähler manifolds.
method Analysis of Bergman kernels and equidistribution for sequences of line bundles.
result Established asymptotic expansion of Bergman kernels and equidistribution of zeros.

RS-Del provides robustness for sequence classifiers against edit distance attacks.

problem Certifying robustness of discrete sequence classifiers against edit distance attacks.
method Randomized deletion (RS-Del) for discrete sequence classifiers, focusing on edit distance-bounded adversaries.
result Achieved a certified accuracy of 91% at an edit distance radius of 128 bytes on malware detection.

Randomized positional encodings boost transformer performance on longer sequences.

problem Transformers struggle with generalizing to sequences of arbitrary length.
method Introduced randomized positional encodings that simulate longer sequences and randomly select positions.
result Randomized positional encodings increase test accuracy by 12.0% on average for sequences of unseen length.

LMC improves sampling from complex distributions using quasi-random sequences.

problem Sampling from complex high-dimensional distributions with high accuracy.
method Using completely uniformly distributed (CUD) sequences in Langevin Monte Carlo (LMC) to generate Gaussian perturbations.
result LMC with low-discrepancy CUD sequences achieves smaller estimation error than standard LMC.

We obtain an index of the complexity of a random sequence by allowing the role of the measure in classical probability theory to be played by a function we call the generating mechanism. Typically, this generating mechanism will be a finite automata. We generate a set of biased sequences by applying a finite state auto…

2008-12-10abs ↗pdf ↗

Study shows mass distribution of random holomorphic sections follows a central limit theorem.

problem Understanding mass distribution of random holomorphic sections.
method Proved a central limit theorem for mass distribution of random holomorphic sections associated with positive line bundles.
result Almost every sequence of random holomorphic sections exhibits quantum ergodicity.

A new method optimizes neural sequence models for better task performance.

problem Training neural sequence models with maximum likelihood estimation ignores task losses.
method Maximum likelihood guided parameter search (MGS) in the parameter space.
result MGS optimizes sequence-level losses, reducing repetition and non-termination.

Strong stability of ergodic iterations proven without ergodic driving sequence.

problem Ensuring strong stability of ergodic iterations under non-ergodic driving sequences.
method Revisiting processes driven by stationary ergodic sequences, proving strong stability under mild conditions on recursive maps.
result Strong stability of iterations proven without ergodic driving sequence.

This research explores various sampling methods and probability distributions for hard alignment in sequence-to-sequence TTS synthesis.

problem Improving alignment accuracy in sequence-to-sequence text-to-speech synthesis.
method Investigated various sampling methods (greedy, beam, random) and probability distributions (Bernoulli, Concrete) for hard alignment.
result Deterministic search is more preferable than stochastic search for natural alignment transition.

This study examines how financial tick data becomes more random with time aggregation.

problem Investigating the randomness of financial tick data over time.
method Applied statistical randomness tests from NIST and TestU01 batteries to ultra-high frequency financial data.
result Financial tick data becomes increasingly random as the aggregation level of transaction time increases.

Transformers can predict pseudo-random sequences from LCGs with unseen parameters and moduli.

problem Learning pseudo-random number sequences from linear congruential generators with unknown parameters and moduli.
method Investigated the ability of Transformers to learn LCG sequences with varying complexity and moduli. Analyzed embedding layers and attention patterns.
result Transformers can predict pseudo-random sequences from LCGs with unseen parameters and moduli, up to mexttest=216m_{ ext{test}} = 2^{16}, using a two-step strategy.

Open problem seeks an online learning algorithm for binary classification.

problem Existence of an online learning algorithm for binary classification with sublinear mistakes.
method Assumption of sequence allowing learning algorithm's existence.
result Specific condition determines sequence's learnability.

Study explores strategies for randomized allocation in delayed rewards bandits.

problem Understanding the exploration-exploitation tradeoff in randomized strategies with delayed rewards.
method Examines two strategies: updating exploration sequence at every time point vs. updating only when a new reward is observed.
result The strategy updating only when a new reward is observed leads to strong consistency in allocation for a wider scope of situations.

The paper estimates variance of random sections on complex manifolds.

problem Estimating variance of random holomorphic sections on compact Kahler manifolds.
method Analyzes a sequence of smooth Hermitian holomorphic line bundles on a compact Kahler manifold X, considering specific probability measures.
result Provides variance estimates for various measures including Gaussian and Fubini-Study measures.

We investigate the statistics of records in a random sequence {xB(0)=0,xB(1),,xB(n)=xB(0)=0}\{x_B(0)=0,x_B(1),\cdots, x_B(n)=x_B(0)=0\} of nn time steps. The sequence xB(k)x_B(k)'s represents the position at step kk of a random walk `bridge' of nn steps that starts and ends at the origin. At each step, the increment of the position is a random ju…

2015-05-22abs ↗pdf ↗

Generative model simulates financial market price variations from order flow.

problem Simulating intra-day price variations driven by order flow.
method Sequence Generative Adversarial Networks framework applied to model order flow.
result Generated price sequences from generative model better match real price variations.

This study examines a single attention layer's capabilities using random features.

problem Understanding the learning and generalization of a single multi-head attention layer.
method Random feature setting with large number of heads, frozen query and key matrices, and trainable value matrices.
result Random-feature attention layer can express a broad class of permutation-invariant target functions.

ECL improves sequence modeling by learning evolutionary structure.

problem Ignoring evolutionary structure in sequence modeling leads to suboptimal performance.
method ECL progressively exposes models to sequences of increasing evolutionary distance.
result ECL improves performance across multiple biological domains and tasks.

FLDCRF improves sequence labeling performance with latent dynamics interactions.

problem Sequence labeling with improved performance and latent dynamics interactions.
method Factored Latent-Dynamic Conditional Random Fields (FLDCRF) with multiple latent dynamics interactions.
result FLDCRF outperforms state-of-the-art models across multiple datasets.

Modeling latent dynamics in high-dimensional event sequences without prior knowledge.

problem Modeling latent dynamics in high-dimensional event sequences with unknown marker relations.
method Adversarial imitation learning framework decomposed into latent structural intensity model, efficient random walk model, and seq2seq discriminator.
result Effective detection of hidden network among markers and decent prediction for future events.

A new model uses normalizing flows for discrete sequences, improving generation speed.

problem Modeling discrete sequences like text using normalizing flows poses challenges.
method Proposes a VAE-based model with autoregressive and non-autoregressive flow architectures.
result Flow-based models can match or improve on autoregressive baselines for discrete sequence tasks.

Given an edge-independent random graph G(n,p), we determine various facts about the cohomology of graph products of groups for the graph G(n,p). In particular, the random graph product of a sequence of finite groups is a rational duality group with probability tending to 1 as n goes to infinity. This includes random ri…

2012-10-16abs ↗pdf ↗

Sum-product networks enhance sequence modeling with higher-order factors.

problem Modeling complex relations in sequence data with first-order models.
method Combining sum-product networks with higher-order linear-chain conditional random fields.
result Improved performance in sequence labeling tasks compared to state-of-the-art methods.

Paper optimizes statistical estimation for randomized smoothing to reduce adversarial robustness certification time.

problem Efficiently estimating robustness of points against adversarial attacks.
method Developed estimation procedures using confidence sequences and randomized Clopper-Pearson intervals.
result Achieved optimal sample complexities and stronger certificates with reduced computational burden.

GGP models multivariate time series with latent sub-sequences for diverse behaviors.

problem Modeling multivariate time series with diverse behaviors and patterns.
method Graph Gamma Process (GGP) linear dynamical systems with latent sub-sequences.
result GGP models exhibit good predictive performance and reveal interpretable latent patterns.

The study shows that random surfaces built from polygons converge to a Poisson-Dirichlet partition.

problem Understanding geometric properties of random surfaces constructed from polygons.
method Uniformly pairing polygon sides to form surfaces, analyzing degree sequences and geometric properties using probabilistic techniques.
result Several geometric properties of the graph are universal, converging to a Poisson-Dirichlet partition as non o \infty.

A new approach models credit card transactions using HMMs to detect fraud.

problem Detecting credit card fraud using isolated event analysis.
method Model sequences from three perspectives using HMMs and combine likelihoods as features.
result Improved fraud detection effectiveness compared to state-of-the-art methods.

Paper shows comparing single performance scores is insufficient for non-deterministic systems, proposing to compare score distributions.

problem Insufficient comparison of non-deterministic sequence tagging systems.
method Compare score distributions based on multiple executions of LSTM-networks.
result LSTM-networks produce superior and more stable performance when compared using score distributions.

This paper examines how different decoding algorithms for LLMs align with various goals.

problem Consistency of decoding algorithms with different goals in LLMs.
method Analysis of greedy, lookahead, random sampling, and temperature-scaled random sampling algorithms.
result Random sampling is consistent with the true probability distribution, but other goals require optimal algorithms for specific probability distributions.

Efficiently improves non-autoregressive sequence models for better translation performance.

problem Heavy inference latency and inconsistent output sentences in non-autoregressive models.
method Incorporates a structured inference module with an efficient CRF approximation and dynamic transition technique.
result Significantly better translation performance (BLEU score 26.80) compared to previous non-autoregressive models.

We consider the problem of improving the efficiency of randomized Fourier feature maps to accelerate training and testing speed of kernel methods on large datasets. These approximate feature maps arise as Monte Carlo approximations to integral representations of shift-invariant kernel functions (e.g., Gaussian kernel).…

2014-12-29abs ↗pdf ↗

A new network embedding method using diffusion to overcome limitations of random walks.

problem Limitations of random walk based network embedding methods in fragile sampling and disequilibrium networks.
method Proposes a network diffusion based embedding method that captures both depth and breadth information and uses cascades for global network information.
result The diffusion based models are more robust in fragile sampling and highly imbalanced networks.

The study proves a central limit theorem for Gaussian holomorphic sections on Kähler manifolds.

problem Understanding statistical properties of zeros of random holomorphic sections.
method Proves a central limit theorem for smooth linear statistics of zero divisors of Gaussian sections in line bundles over Kähler manifolds.
result Derives first-order asymptotics and upper decay estimates for Bergman kernels.

The study analyzes how stochastic recursive algorithms converge to Markov chains.

problem Understanding convergence of stochastic recursive algorithms to Markov chains.
method Analyzes iterated random operators and contraction operators over Polish spaces.
result The distribution of random sequences converges to the invariant distribution of the Markov chain.