Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

9192837 · Jun 201919922001200920172026
48 results for alphabet frequencies

Study explains Zipf's law using geometric mechanisms from a finite alphabet.

problem Explains Zipf's law in language without relying on linguistic elements.
method Uses the Full Combinatorial Word Model (FCWM) to generate geometric distributions of word lengths.
result Supports predictions of power-law rank-frequency curves, matching various languages.

Estimating a large alphabet probability distribution from a limited number of samples is a fundamental problem in machine learning and statistics. A variety of estimation schemes have been proposed over the years, mostly inspired by the early work of Laplace and the seminal contribution of Good and Turing. One of the b…

2018-07-06abs ↗pdf ↗

In this paper we consider the problem of transmitting a continuous alphabet discrete-time source over an AWGN channel. The design of good curves for this purpose relies on geometrical properties of spherical codes and projections of NN-dimensional lattices. We propose a constructive scheme based on a set of curves on …

2012-02-09abs ↗pdf ↗

We consider the problem of predicting the next observation given a sequence of past observations, and consider the extent to which accurate prediction requires complex algorithms that explicitly leverage long-range dependencies. Perhaps surprisingly, our positive results show that for a broad class of sequences, there …

2016-12-08abs ↗pdf ↗

Knots and links are interpreted as homotopy classes of nanowords and nanophrases in an alphabet consisting of 4 letters. Similar results hold for curves on surfaces. We also discuss versions of the Jones link polynomial and the link quandles for nanophrases.

2005-06-20abs ↗pdf ↗

We study the problem of learning overcomplete HMMs---those that have many hidden states but a small output alphabet. Despite having significant practical importance, such HMMs are poorly understood with no known positive or negative results for efficient learning. In this paper, we present several new results---both po…

2017-11-07abs ↗pdf ↗

New algorithms reduce rejection sampling complexity for shape-constrained distributions.

problem Generating exact samples from shape-constrained distributions efficiently.
method Sublinear query complexity algorithms for rejection sampling.
result Sublinear complexity algorithms for sampling from shape-constrained distributions.

We consider the problem of enumerating relevant features hidden in other irrelevant information for multi-labeled data, which is formalized as learning juntas. A kk-junta function is a function which depends on only kk coordinates of the input. For relatively small kk w.r.t. the input size nn, learning kk-junta fu…

2019-03-15abs ↗pdf ↗

We discuss a topological approach to words introduced by the author. Words on an arbitrary alphabet are approximated by Gauss words and then studied up to natural modifications inspired by the Reidemeister moves on knot diagrams. This leads us to a notion of homotopy for words. We introduce several homotopy invariants …

2006-09-19abs ↗pdf ↗

The theoretical basis for a candidate variational principle for the information bottleneck (IB) method is formulated within the ambit of the generalized nonadditive statistics of Tsallis. Given a nonadditivity parameter q q , the role of the \textit{additive duality} of nonadditive statistics (q=2q q^*=2-q ) in relating…

2008-11-19abs ↗pdf ↗

The study optimizes distribution estimation from samples with relative entropy error, adapting to sparse distributions.

problem Estimating discrete distributions with high-probability accuracy in relative entropy.
method Analysis of Laplace estimator and confidence-dependent smoothing techniques, including data-dependent smoothing.
result Optimal high-probability risk bounds for various estimators, including a new data-dependent smoothing method.

We generalize presentations of the fundamental group of discriminant complements and arrive at a class of presentations associated naturally with words in the free monoid of the alphabet σ1,,σn1σ_1,\dots,σ_{n-1}. Our study addresses invariance properties of these presentations and the presented groups under various operatio…

2020-01-24abs ↗pdf ↗

BestChanID identifies the channel with maximal capacity using training sequences.

problem Identifying the channel with maximal capacity among several discrete memoryless channels.
method Formulated as a multi-armed bandit problem, proposed a capacity estimator, and developed gap-elimination algorithms.
result Guaranteed to output the DMC with the largest capacity with a desired confidence.

In this article we present a finite generating set G2G_2 of H2\mathcal{H}_2, the genus-2 Goeritz group of S3S^3, in terms of Dehn twists about certain simple closed curves on the standard Heegaard surface. We present an algorithm that describes an element ψH2ψ\in\mathcal{H}_2 as a word in the alphabet of G2G_2 in a cert…

2019-12-18abs ↗pdf ↗

Paper proposes a fast stochastic algorithm for neural network quantization with error bounds.

problem Error analysis for quantized neural networks with non-convex loss functions and nonlinear activations.
method Greedy path-following mechanism combined with stochastic quantizer.
result Established full-network error bounds for quantized neural networks.

Viewing Dehn's algorithm as a rewriting system, we generalise to allow an alphabet containing letters which do not necessarily represent group elements. This extends the class of groups for which the algorithm solves the word problem to include nilpotent groups, many relatively hyperbolic groups including geometrically…

2007-06-20abs ↗pdf ↗

Robust hypothesis testing designs a test for worst-case distributions using kernel methods.

problem Design a robust test for hypothesis testing under uncertainty sets.
method Data-driven uncertainty sets constructed using kernel mean embeddings and maximum mean discrepancy (MMD). Bayesian and Neyman-Pearson settings investigated.
result Proposed robust kernel tests are exponentially consistent and asymptotically optimal.

This paper is concerned with jointly recovering nn node-variables {xi}1in\left\{ x_{i}\right\}_{1\leq i\leq n} from a collection of pairwise difference measurements. Imagine we acquire a few observations taking the form of xixjx_{i}-x_{j}; the observation pattern is represented by a measurement graph G\mathcal{G} with an ed…

2015-04-06abs ↗pdf ↗

HyFAD improves time series imputation by combining time and frequency diffusion.

problem Improve time series imputation by handling frequency-sensitive denoising and balancing global and local dynamics.
method HyFAD is a hybrid time-frequency diffusion model with frequency-aware embedding, built on DDPM paradigm.
result HyFAD achieves state-of-the-art performance in time series imputation.

SSMs have a built-in bias towards low-frequency components, which can be adjusted.

problem Frequency bias in SSMs affects their performance on long-range sequences.
method Proposed two mechanisms to tune frequency bias: scaling initialization or applying a Sobolev-norm-based filter.
result Tuning frequency bias improves SSMs' performance on long-range sequence learning tasks.

Visual spoofing bypasses spam filters and plagiarism detection.

problem Vulnerability in spam filters that can be exploited by visually similar but differently encoded characters.
method Replaces characters with visually similar but differently encoded characters from a different alphabet.
result Spammers can create messages that bypass existing spam filters.

We present some nonparametric methods for graphical modeling. In the discrete case, where the data are binary or drawn from a finite alphabet, Markov random fields are already essentially nonparametric, since the cliques can take only a finite number of values. Continuous data are different. The Gaussian graphical mode…

2012-01-04abs ↗pdf ↗

A Gauss paragraph is a combinatorial formulation of a generic closed curve with multiple components on some surface. A virtual string is a collection of circles with arrows that represent the crossings of such a curve. Every closed curve has an underlying virtual string and every virtual string has an underlying Gauss …

2004-12-01abs ↗pdf ↗

The paper examines how parabolic frequency behaves under Ricci flow and Ricci-harmonic flow on manifolds.

problem Understanding the behavior of parabolic frequency under Ricci flow and Ricci-harmonic flow.
method Investigates the monotonicity of parabolic frequency for solutions of linear and heat equations with bounded curvatures.
result Establishes monotonicity results for parabolic frequency under specific curvature conditions.

New method constrains CNN filter frequencies to improve robustness.

problem CNN bias towards low frequency components, leading to poor performance in scenario transformations.
method Frequency domain regularization by constraining filter spectra, training valid frequency range end-to-end.
result Demonstrated effectiveness in defending adversarial perturbations, reducing generalization gap, and improving transfer learning.

We advocate the use of a notion of entropy that reflects the relative abundances of the symbols in an alphabet, as well as the similarities between them. This concept was originally introduced in theoretical ecology to study the diversity of ecosystems. Based on this notion of entropy, we introduce geometry-aware count…

2019-06-19abs ↗pdf ↗

Study uses multi-kernel Hawkes models to analyze high-frequency price dynamics.

problem Understanding responsive speeds of market participants in high-frequency trading.
method Multi-kernel Hawkes models with conditional Hessian analysis for optimization.
result Existence of multi-kernels (UHF, VHF, HF) in high-frequency price dynamics.

Paper extends SI method for detecting CPs in complex systems' frequency domain.

problem Identifying change points in complex systems' frequency domain.
method Extends SI framework to frequency domain using DFT properties and develops valid p-values.
result Reliable detection of genuine CPs with strong statistical guarantees.

Proves monotonicity of parabolic frequency on all manifolds without curvature assumptions.

problem Monotonicity of parabolic frequency on manifolds.
method Analyzes parabolic frequency function on manifolds, proving monotonicity without curvature assumptions.
result Monotonicity of parabolic frequency on all manifolds, no curvature assumption needed.