Deep learning uses alphabet frequencies to accurately classify fake news.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study explains Zipf's law using geometric mechanisms from a finite alphabet.
Paper explores using hand gestures to type English letters.
Estimating a large alphabet probability distribution from a limited number of samples is a fundamental problem in machine learning and statistics. A variety of estimation schemes have been proposed over the years, mostly inspired by the early work of Laplace and the seminal contribution of Good and Turing. One of the b…
A simple text model shows word lengths follow Zipf's law.
In this paper we consider the problem of transmitting a continuous alphabet discrete-time source over an AWGN channel. The design of good curves for this purpose relies on geometrical properties of spherical codes and projections of -dimensional lattices. We propose a constructive scheme based on a set of curves on …
New Sauer inequality improves multiclass hypothesis class bounds.
We consider the problem of predicting the next observation given a sequence of past observations, and consider the extent to which accurate prediction requires complex algorithms that explicitly leverage long-range dependencies. Perhaps surprisingly, our positive results show that for a broad class of sequences, there …
We propose a novel receiver for orthogonal frequency division multiplexing (OFDM) transmissions in impulsive noise environments. Impulsive noise arises in many modern wireless and wireline communication systems, such as Wi-Fi and powerline communications, due to uncoordinated interference that is much stronger than the…
Knots and links are interpreted as homotopy classes of nanowords and nanophrases in an alphabet consisting of 4 letters. Similar results hold for curves on surfaces. We also discuss versions of the Jones link polynomial and the link quandles for nanophrases.
We study the problem of learning overcomplete HMMs---those that have many hidden states but a small output alphabet. Despite having significant practical importance, such HMMs are poorly understood with no known positive or negative results for efficient learning. In this paper, we present several new results---both po…
New algorithms reduce rejection sampling complexity for shape-constrained distributions.
Independent component analysis (ICA) is a statistical method for transforming an observable multi-dimensional random vector into components that are as statistically independent as possible from each other. Usually the ICA framework assumes a model according to which the observations are generated (such as a linear tra…
This abstract explores an RNN-based approach to online handwritten recognition problem. Our method uses data from an accelerometer and a gyroscope mounted on a handheld pen-like device to train and run a character pre-diction model. We have built a dataset of timestamped gyroscope and accelerometer data gathered during…
We consider the problem of enumerating relevant features hidden in other irrelevant information for multi-labeled data, which is formalized as learning juntas. A -junta function is a function which depends on only coordinates of the input. For relatively small w.r.t. the input size , learning -junta fu…
We discuss a topological approach to words introduced by the author. Words on an arbitrary alphabet are approximated by Gauss words and then studied up to natural modifications inspired by the Reidemeister moves on knot diagrams. This leads us to a notion of homotopy for words. We introduce several homotopy invariants …
The theoretical basis for a candidate variational principle for the information bottleneck (IB) method is formulated within the ambit of the generalized nonadditive statistics of Tsallis. Given a nonadditivity parameter , the role of the \textit{additive duality} of nonadditive statistics () in relating…
Motivation: Proteins are known to undergo conformational changes in the course of their functions. The changes in conformation are often attributable to a small fraction of residues within the protein. Therefore identification of these variable regions is important for an understanding of protein function. Results: We …
We characterize the effectiveness of a classical algorithm for recovering the Markov graph of a general discrete pairwise graphical model from i.i.d. samples. The algorithm is (appropriately regularized) maximum conditional log-likelihood, which involves solving a convex program for each node; for Ising models this is …
The study optimizes distribution estimation from samples with relative entropy error, adapting to sparse distributions.
This paper studies the prediction of chord progressions for jazz music by relying on machine learning models. The motivation of our study comes from the recent success of neural networks for performing automatic music composition. Although high accuracies are obtained in single-step prediction scenarios, most models fa…
We generalize presentations of the fundamental group of discriminant complements and arrive at a class of presentations associated naturally with words in the free monoid of the alphabet . Our study addresses invariance properties of these presentations and the presented groups under various operatio…
BestChanID identifies the channel with maximal capacity using training sequences.
In this article we present a finite generating set of , the genus-2 Goeritz group of , in terms of Dehn twists about certain simple closed curves on the standard Heegaard surface. We present an algorithm that describes an element as a word in the alphabet of in a cert…
Paper proposes a fast stochastic algorithm for neural network quantization with error bounds.
Web-based framework detects Parkinson's disease from speech recordings.
This study calculates the maximum error of a famous estimation method.
Viewing Dehn's algorithm as a rewriting system, we generalise to allow an alphabet containing letters which do not necessarily represent group elements. This extends the class of groups for which the algorithm solves the word problem to include nilpotent groups, many relatively hyperbolic groups including geometrically…
Robust hypothesis testing designs a test for worst-case distributions using kernel methods.
This paper is concerned with jointly recovering node-variables from a collection of pairwise difference measurements. Imagine we acquire a few observations taking the form of ; the observation pattern is represented by a measurement graph with an ed…
HyFAD improves time series imputation by combining time and frequency diffusion.
SSMs have a built-in bias towards low-frequency components, which can be adjusted.
Visual spoofing bypasses spam filters and plagiarism detection.
Geometrically interprets frequency in electric circuits.
Trading affects grid frequency fluctuations, making them more extreme.
We present some nonparametric methods for graphical modeling. In the discrete case, where the data are binary or drawn from a finite alphabet, Markov random fields are already essentially nonparametric, since the cliques can take only a finite number of values. Continuous data are different. The Gaussian graphical mode…
Study on frequencies of non-simple curves in surfaces of large genus.
CNNs show sensitivity to low-frequency signals due to image frequency distribution.
A Gauss paragraph is a combinatorial formulation of a generic closed curve with multiple components on some surface. A virtual string is a collection of circles with arrows that represent the crossings of such a curve. Every closed curve has an underlying virtual string and every virtual string has an underlying Gauss …
The paper examines how parabolic frequency behaves under Ricci flow and Ricci-harmonic flow on manifolds.
New method constrains CNN filter frequencies to improve robustness.
We advocate the use of a notion of entropy that reflects the relative abundances of the symbols in an alphabet, as well as the similarities between them. This concept was originally introduced in theoretical ecology to study the diversity of ecosystems. Based on this notion of entropy, we introduce geometry-aware count…
Study uses multi-kernel Hawkes models to analyze high-frequency price dynamics.
Paper extends SI method for detecting CPs in complex systems' frequency domain.
The paper defines a frequency for mean curvature flow and proves its monotonicity.
Paper defines parabolic frequency for Ricci flow solutions, proving monotonicity and uniqueness.
Proves monotonicity of parabolic frequency on all manifolds without curvature assumptions.
New Fourier-based diffusion model improves high-frequency generation quality.