Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

336598130 · Jun 202019922001200920172026
48 results for Hamming distance

New STH distance finds patterns in event timeseries without resampling.

problem Lack of efficient analysis methods for event and state timeseries.
method Define STE-ts, propose STH, leveraging both time and state duration.
result Improved precision and computation time compared to resampled metrics.

ExDAG solves DAG learning problems with low structural Hamming distance.

problem Learning DAGs with low structural Hamming distance under identifiability assumptions.
method Mixed-integer quadratic programming (MIQP) with branch-and-bound-and-cut algorithm and lazy constraints.
result ExDAG guarantees global convergence and provides a real-time quality assessment.

Study shows effective resistance distance yields more accurate network barycenter than Hamming distance.

problem Identifying the best metric for computing the Fréchet mean network.
method Compared the effectiveness of Hamming distance and effective resistance distance in capturing network topology.
result Effective resistance distance produces a more accurate Fréchet mean network.

Differentially private data structures for estimating distances between strings.

problem Estimating distances between query strings and database strings while ensuring privacy.
method Proposes differentially private data structures for Hamming and edit distances using randomized response technique.
result Efficient data structures that provide accurate distance estimates with strong privacy guarantees.

New metric space for ReLU codes connects to network safety and robustness.

problem Lack of metrics capturing network safety and robustness beyond accuracy.
method Introduces a metric space of ReLU activation codes with a truncated Hamming distance.
result Establishes an isometry between ReLU codes and polyhedral bodies related to safety and robustness.

Asynchronous Gibbs sampling has been recently shown to be fast-mixing and an accurate method for estimating probabilities of events on a small number of variables of a graphical model satisfying Dobrushin's condition~\cite{DeSaOR16}. We investigate whether it can be used to accurately estimate expectations of functions…

2018-11-26abs ↗pdf ↗

Classifies homeomorphism groups of countable Stone spaces up to coarse equivalence.

problem Classifying non-locally compact topological groups using geometric group theory.
method Classification based on coarsely bounded sets and quasi-isometry.
result Groups in the second class are quasi-isometric to the Hamming cube.

A new, efficient kk-modes algorithm improves clustering of categorical data.

problem Clustering categorical data using existing methods like kk-means is inefficient.
method Developed a novel kk-modes algorithm called OTQT, which improves on existing methods.
result OTQT finds more accurate clusters per iteration and is faster overall.

We present a powerful new loss function and training scheme for learning binary hash codes with any differentiable model and similarity function. Our loss function improves over prior methods by using log likelihood loss on top of an accurate approximation for the probability that two inputs fall within a Hamming dista…

2018-10-01abs ↗pdf ↗

Our work proves robustness of embedding schemes to discrete changes in text.

problem Discrete changes in text, like replacing a word, affect model robustness.
method Formal proofs and quantitative bounds for embedding schemes (concatenation, TF-IDF, Paragraph Vector).
result Embedding schemes are robust to discrete changes in text with Hölder or Lipschitz properties.

Hashing, or learning binary embeddings of data, is frequently used in nearest neighbor retrieval. In this paper, we develop learning to rank formulations for hashing, aimed at directly optimizing ranking-based evaluation metrics such as Average Precision (AP) and Normalized Discounted Cumulative Gain (NDCG). We first o…

2017-05-23abs ↗pdf ↗

Locality-sensitive hashing converts high-dimensional feature vectors, such as image and speech, into bit arrays and allows high-speed similarity calculation with the Hamming distance. There is a hashing scheme that maps feature vectors to bit arrays depending on the signs of the inner products between feature vectors a…

2012-12-26abs ↗pdf ↗

A new metric compares true and learned causal graphs considering data and graph structure.

problem Comparing true and learned causal graphs accurately.
method Continuous Structural Intervention Distance (CSID) using conditional mean embeddings and maximum mean discrepancy.
result Validated the CSID with synthetic data, showing its effectiveness in comparing causal graphs.

We present a powerful new loss function and training scheme for learning binary hash functions. In particular, we demonstrate our method by creating for the first time a neural network that outperforms state-of-the-art Haar wavelets and color layout descriptors at the task of automated scene matching. By accurately rel…

2018-02-09abs ↗pdf ↗

Causal inference relies on the structure of a graph, often a directed acyclic graph (DAG). Different graphs may result in different causal inference statements and different intervention distributions. To quantify such differences, we propose a (pre-) distance between DAGs, the structural intervention distance (SID). T…

2013-06-05abs ↗pdf ↗

In this note we prove that for each positive integer mm there exists a bi-Lipschitz embedding ZmHam(S2)Z^m\to Ham(S^2), where Ham(S2)Ham(S^2) is equipped with the entropy metric. In particular, the same result holds when the entropy metric is substituted with the autonomous metric.

2019-09-12abs ↗pdf ↗

We prove that π1(Ham(M))π_1(\text{Ham}(M)) contains an infinite cyclic subgroup, where Ham(M)\text{Ham}(M) is the Hamiltonian group of the one point blow up of CP3{\Bbb C}P^3. We give a sufficient condition for the group π1(Ham(M))π_1(\text{Ham}(M)) to contain an infinite cyclic subgroup, when MM is a general toric manifold.

2005-06-09abs ↗pdf ↗

We verify here some variants of topological and dynamical flavor of the injectivity radius conjecture in Hofer geometry, Lalonde-Savelyev \cite{citeLalondeSavelyevOntheinjectivityradiusinHofergeometry} in the case of Ham(S2)Ham (S^2) and Ham(Σ,ω)Ham(Σ, ω), for ΣΣ a closed positive genus surface. In particular we show that any lo…

2015-01-12abs ↗pdf ↗

Let S(s,w)\mathfrak{S}(\underline{s},w) be the graph whose vertices are all subexpressions with target ww of a fixed expression s\underline{s} in generators of a Coxeter group and edges are the pairs of subexpressions with Hamming distance 2. We prove that S(s,w)\mathfrak{S}(\underline{s},w) is connected and its cycle space …

2025-06-12abs ↗pdf ↗

Let SS be a compact oriented surface. We construct homogeneous quasimorphisms on Diff(S,area)Diff(S, area), on Diff0(S,area)Diff_0(S, area) and on Ham(S)Ham(S) generalizing the constructions of Gambaudo-Ghys and Polterovich. We prove that there are infinitely many linearly independent homogeneous quasimorphisms on Diff(S,area)Diff(S, area), on $Diff_0(…

2017-07-19abs ↗pdf ↗

Multi-label classification is a type of supervised learning where an instance may belong to multiple labels simultaneously. Predicting each label independently has been criticized for not exploiting any correlation between labels. In this paper we propose a novel approach, Nearest Labelset using Double Distances (NLDD)…

2017-02-15abs ↗pdf ↗

We present a lower bound for a fragmentation norm and construct a bi-Lipschitz embedding I ⁣:RnHam(M)I\colon \mathbb{R}^n\to\mathrm{Ham}(M) with respect to the fragmentation norm on the group Ham(M)\mathrm{Ham}(M) of Hamiltonian diffeomorphisms of a symplectic manifold (M,ω)(M,ω). As an application, we provide an answer to Brandenbursk…

2019-01-07abs ↗pdf ↗

Proposes a new Bayesian score for learning network structure from related datasets.

problem Learning network structure from heterogeneous related data sets.
method Bayesian Hierarchical Dirichlet (BHD) score based on a hierarchical model.
result BHD outperforms BDeu in reconstruction accuracy and sparsity for related datasets.

PAIR-CI calibrates CI tests for causal discovery with incomplete data.

problem Miscalibration of CI tests when imputing incomplete data.
method Integrates multiple imputation directly into the inferential procedure via a paired permutation design.
result PAIR-CI reduces false positive rates to below 5% in simulations.

Consider a linear regression model where the design matrix X has n rows and p columns. We assume (a) p is much large than n, (b) the coefficient vector beta is sparse in the sense that only a small fraction of its coordinates is nonzero, and (c) the Gram matrix G = X'X is sparse in the sense that each row has relativel…

2012-04-29abs ↗pdf ↗

In this work we construct Calabi quasi-morphisms on the universal cover of the group Ham(M) of Hamiltonian diffeomorphisms for some non-monotone symplectic manifolds. This complements a result by Entov and Polterovich which applies in the monotone case. Moreover, in contrast to their work, we show that these quasi-morp…

2005-08-04abs ↗pdf ↗

Differentiable causal discovery methods perform robustly under model violations.

problem Causal discovery algorithms struggle with real-world data due to unverifiable causal assumptions.
method Benchmarked differentiable causal discovery methods under eight model assumption violations.
result Differentiable causal discovery methods exhibit robust performance under Structural Hamming Distance and Structural Intervention Distance metrics.

A fast binary embedding method preserves Euclidean distances in high-dimensional data.

problem Preserving Euclidean distances in high-dimensional datasets.
method Stable noise-shaping quantization of AxA x with AA a sparse Gaussian random matrix, followed by a linear transformation.
result Euclidean distances are approximated by the 1\ell_1 norm on binary sequences, leading to accurate binary codes.

Vectors of data are at the heart of machine learning and data mining. Recently, vector quantization methods have shown great promise in reducing both the time and space costs of operating on vectors. We introduce a vector quantization algorithm that can compress vectors over 12x faster than existing techniques while al…

2017-06-30abs ↗pdf ↗

The paper proposes a semi-parametric Bayesian network model using Gaussian Processes and Horseshoe priors.

problem Learning semi-parametric relationships in Expert Bayesian Networks with minimal nonlinear components.
method Uses Gaussian Processes and Horseshoe priors to model relationships, prioritizes modifying expert graphs, and generates diverse graphs.
result Models outperform state-of-the-art semi-parametric Bayesian Network models in synthetic and real-world datasets.

Probabilistic graphical models are graphical representations of probability distributions. Graphical models have applications in many fields including biology, social sciences, linguistic, neuroscience. In this paper, we propose directed acyclic graphs (DAGs) learning via bootstrap aggregating. The proposed procedure i…

2014-06-09abs ↗pdf ↗