ExDAG solves DAG learning problems with low structural Hamming distance.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Differentially private data structures for estimating distances between strings.
Study robust mean estimation under coordinate-level corruptions using Hamming distance.
New method identifies extreme risk propagation in financial networks.
Study shows effective resistance distance yields more accurate network barycenter than Hamming distance.
We revisit fuzzy neural network with a cornerstone notion of generalized hamming distance, which provides a novel and theoretically justified framework to re-interpret many useful neural network techniques in terms of fuzzy logic. In particular, we conjecture and empirically illustrate that, the celebrated batch normal…
New STH distance finds patterns in event timeseries without resampling.
A new metric compares true and learned causal graphs considering data and graph structure.
We extend the definition of Weinstein's Action homomorphism to Hamiltonian actions with equivariant moment maps of (possibly infinite-dimensional) Lie groups on symplectic manifolds, and show that under conditions including a uniform bound on the symplectic areas of geodesic triangles the resulting homomorphism extends…
Causal inference relies on the structure of a graph, often a directed acyclic graph (DAG). Different graphs may result in different causal inference statements and different intervention distributions. To quantify such differences, we propose a (pre-) distance between DAGs, the structural intervention distance (SID). T…
New metric space for ReLU codes connects to network safety and robustness.
Proposes a new Bayesian score for learning network structure from related datasets.
Score based learning (SBL) is a promising approach for learning Bayesian networks in the discrete domain. However, when employing SBL in the continuous domain, one is either forced to move the problem to the discrete domain or use metrics such as BIC/AIC, and these approaches are often lacking. Discretization can have …
The existence of adversarial examples in which an imperceptible change in the input can fool well trained neural networks was experimentally discovered by Szegedy et al in 2013, who called them "Intriguing properties of neural networks". Since then, this topic had become one of the hottest research areas within machine…
Asynchronous Gibbs sampling has been recently shown to be fast-mixing and an accurate method for estimating probabilities of events on a small number of variables of a graphical model satisfying Dobrushin's condition~\cite{DeSaOR16}. We investigate whether it can be used to accurately estimate expectations of functions…
Classifies homeomorphism groups of countable Stone spaces up to coarse equivalence.
A new, efficient -modes algorithm improves clustering of categorical data.
We present a powerful new loss function and training scheme for learning binary hash codes with any differentiable model and similarity function. Our loss function improves over prior methods by using log likelihood loss on top of an accurate approximation for the probability that two inputs fall within a Hamming dista…
Develops homotopies for Lagrangian field theory using advanced algebraic structures.
Large scale agglomerative clustering is hindered by computational burdens. We propose a novel scheme where exact inter-instance distance calculation is replaced by the Hamming distance between Kernelized Locality-Sensitive Hashing (KLSH) hashed values. This results in a method that drastically decreases computation tim…
Sharp threshold found for Frechet mean of inhomogeneous graphs.
New measures assess differences in causal graphs' separations.
Our work proves robustness of embedding schemes to discrete changes in text.
Hashing, or learning binary embeddings of data, is frequently used in nearest neighbor retrieval. In this paper, we develop learning to rank formulations for hashing, aimed at directly optimizing ranking-based evaluation metrics such as Average Precision (AP) and Normalized Discounted Cumulative Gain (NDCG). We first o…
PAIR-CI calibrates CI tests for causal discovery with incomplete data.
The Quadratic Assignment Problem (QAP) is a well-known permutation-based combinatorial optimization problem with real applications in industrial and logistics environments. Motivated by the challenge that this NP-hard problem represents, it has captured the attention of the optimization community for decades. As a resu…
Locality-sensitive hashing converts high-dimensional feature vectors, such as image and speech, into bit arrays and allows high-speed similarity calculation with the Hamming distance. There is a hashing scheme that maps feature vectors to bit arrays depending on the signs of the inner products between feature vectors a…
Differentiable causal discovery methods perform robustly under model violations.
We present a powerful new loss function and training scheme for learning binary hash functions. In particular, we demonstrate our method by creating for the first time a neural network that outperforms state-of-the-art Haar wavelets and color layout descriptors at the task of automated scene matching. By accurately rel…
The paper proposes a semi-parametric Bayesian network model using Gaussian Processes and Horseshoe priors.
We study the problem of recovering the latent ground truth labeling of a structured instance with categorical random variables in the presence of noisy observations. We present a new approximate algorithm for graphs with categorical variables that achieves low Hamming error in the presence of noisy vertex and edge obse…
In this note we prove that for each positive integer there exists a bi-Lipschitz embedding , where is equipped with the entropy metric. In particular, the same result holds when the entropy metric is substituted with the autonomous metric.
This paper introduces a novel real-time Fuzzy Supervised Learning with Binary Meta-Feature (FSL-BM) for big data classification task. The study of real-time algorithms addresses several major concerns, which are namely: accuracy, memory consumption, and ability to stretch assumptions and time complexity. Attaining a fa…
We prove that contains an infinite cyclic subgroup, where is the Hamiltonian group of the one point blow up of . We give a sufficient condition for the group to contain an infinite cyclic subgroup, when is a general toric manifold.
We verify here some variants of topological and dynamical flavor of the injectivity radius conjecture in Hofer geometry, Lalonde-Savelyev \cite{citeLalondeSavelyevOntheinjectivityradiusinHofergeometry} in the case of and , for a closed positive genus surface. In particular we show that any lo…
Let be the graph whose vertices are all subexpressions with target of a fixed expression in generators of a Coxeter group and edges are the pairs of subexpressions with Hamming distance 2. We prove that is connected and its cycle space …
In this paper, we develop four malware detection methods using Hamming distance to find similarity between samples which are first nearest neighbors (FNN), all nearest neighbors (ANN), weighted all nearest neighbors (WANN), and k-medoid based nearest neighbors (KMNN). In our proposed methods, we can trigger the alarm i…
Local approach learns causal structure of linear Gaussian polytree models from interventional data.
New framework estimates staged tree models using hierarchical clustering on the probability simplex.
We establish bounds on the KL divergence between two multivariate Gaussian distributions in terms of the Hamming distance between the edge sets of the corresponding graphical models. We show that the KL divergence is bounded below by a constant when the graphs differ by at least one edge; this is essentially the tighte…
Let be a compact oriented surface. We construct homogeneous quasimorphisms on , on and on generalizing the constructions of Gambaudo-Ghys and Polterovich. We prove that there are infinitely many linearly independent homogeneous quasimorphisms on , on $Diff_0(…
Multi-label classification is a type of supervised learning where an instance may belong to multiple labels simultaneously. Predicting each label independently has been criticized for not exploiting any correlation between labels. In this paper we propose a novel approach, Nearest Labelset using Double Distances (NLDD)…
We present a lower bound for a fragmentation norm and construct a bi-Lipschitz embedding with respect to the fragmentation norm on the group of Hamiltonian diffeomorphisms of a symplectic manifold . As an application, we provide an answer to Brandenbursk…
New method recovers transportable DAG structures from different datasets.
Probabilistic graphical models are graphical representations of probability distributions. Graphical models have applications in many fields including biology, social sciences, linguistic, neuroscience. In this paper, we propose directed acyclic graphs (DAGs) learning via bootstrap aggregating. The proposed procedure i…
Unified analysis of multilabel Fisher discriminants with improved dimensionality and robustness.
Unified analysis of multilabel Fisher discriminants with improved dimensionality and robustness.
We apply Gromov's ham sandwich method to get (1) domain monotonicity (up to a multiplicative constant factor); (2) reverse domain monotonicity (up to a multiplicative constant factor); and (3) universal inequalities for Neumann eigenvalues of the Laplacian on bounded convex domains in a Euclidean space.