Lipschitz equivalence of self-similar sets is an important area in the study of fractal geometry. It is known that two dust-like self-similar sets with the same contraction ratios are always Lipschitz equivalent. However, when self-similar sets have touching structures the problem of Lipschitz equivalence becomes much …
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Skeleton is a new notion designed for constructing space-filling curves of self-similar sets. It is shown in [Dai, Rao and Zhang, Space-filling curves of self-similar sets (II): Edge-to-trail substitution rule,https://doi.org/10.1088/1361-6544/ab1275] that for a connected self-similar set, space-filling curves can be c…
In this paper, we introduce the notion of asymptotic self-similar sets on general doubling metric spaces by extending the notion of self-similar sets, and determine their Hausdorff dimensions, which gives an extension of Balogh and Rohner 's result. This is carried out by introducing the notions of almost similarity ma…
Estimates set overlap and similarity using random samples.
New clustering method using point-set kernel measures similarity.
In [9] Kaimanovich introduced the concept of augmented tree on the symbolic space of a self-similar set. It is hyperbolic in the sense of Gromov, and it was shown in [13] that under the open set condition, a self-similar set can be identified with the hyperbolic boundary of the tree. In the paper, we investigate in det…
Paper explains contrastive learning using cosine similarity and proposes mitigations for batch size effects.
New stability measures for similar features improve feature selection accuracy.
Researchers create a Fredholm module on fractal shapes like the Cantor set.
Bayesian methods detect significant IIA violations in similarity choice data.
Given only information in the form of similarity triplets "Object A is more similar to object B than to object C" about a data set, we propose two ways of defining a kernel function on the data set. While previous approaches construct a low-dimensional Euclidean embedding of the data set that reflects the given similar…
The two main theorems of this paper provide a characterization of hyperbolic affine iterated function systems defined on Rm. Atsushi Kameyama (Distances on Topological Self-Similar Sets, Proceedings of Symposia in Pure Mathematics, Volume 72.1, 2004) asked the following fundamental question: given a topological self-si…
We introduce GSimCNN (Graph Similarity Computation via Convolutional Neural Networks) for predicting the similarity score between two graphs. As the core operation of graph similarity search, pairwise graph similarity computation is a challenging problem due to the NP-hard nature of computing many graph distance/simila…
In this paper, we propose a family of graph partition similarity measures that take the topology of the graph into account. These graph-aware measures are alternatives to using set partition similarity measures that are not specifically designed for graph partitions. The two types of measures, graph-aware and set parti…
Proposes a new stability measure for model fitting on similar feature data sets.
Graph similarity computation is one of the core operations in many graph-based applications, such as graph similarity search, graph database analysis, graph clustering, etc. Since computing the exact distance/similarity between two graphs is typically NP-hard, a series of approximate methods have been proposed with a t…
The study explores homeomorphism groups of self-similar 2-manifolds, including the 2-sphere and Cantor set.
Develops a fast method to analyze image similarities without pre-trained models.
In a context of document co-clustering, we define a new similarity measure which iteratively computes similarity while combining fuzzy sets in a three-partite graph. The fuzzy triadic similarity (FT-Sim) model can deal with uncertainty offers by the fuzzy sets. Moreover, with the development of the Web and the high ava…
ProtoBandit uses bandits to find prototypes efficiently.
CatSIM measures image similarity robustly to small changes.
We consider the problem of predicting several response variables using the same set of explanatory variables. This setting naturally induces a group structure over the coefficient matrix, in which every explanatory variable corresponds to a set of related coefficients. Most of the existing methods that utilize this gro…
Automates similarity measure construction from data.
SOAK assesses data subset similarity for better model training.
Twinning splits data into fast, statistically similar sets.
Paper develops active learning for clustering unknown pairwise similarities.
Study properties of self-similar continua with finite intersection property.
Fractal Lipschitz-Killing curvature measures C^f_k(F,.), k = 0, ..., d, are determined for a large class of self-similar sets F in R^d. They arise as weak limits of the appropriately rescaled classical Lipschitz-Killing curvature measures C_k(F_r,.) from geometric measure theory of parallel sets F_r for small distances…
Hierarchical clustering based on pairwise similarities is a common tool used in a broad range of scientific applications. However, in many problems it may be expensive to obtain or compute similarities between the items to be clustered. This paper investigates the hierarchical clustering of N items based on a small sub…
Given an iterated function system (IFS) of contractive similitudes, the theory of Gromov hyperbolic graph on the IFS has been established recently. In the paper, we introduce a notion of simple augmented tree which is a Gromov hyperbolic graph. By generalizing a combinatorial device of rearrangeable matrix, we show tha…
Study on self-similar sets on Riemannian manifolds with new separation conditions.
First constructed genus 2 Cantor set in 3D space.
Excessive reuse of test data has become commonplace in today's machine learning workflows. Popular benchmarks, competitions, industrial scale tuning, among other applications, all involve test data reuse beyond guidance by statistical confidence bounds. Nonetheless, recent replication studies give evidence that popular…
Minwise hashing is the standard technique in the context of search and databases for efficiently estimating set (e.g., high-dimensional 0/1 vector) similarities. Recently, b-bit minwise hashing was proposed which significantly improves upon the original minwise hashing in practice by storing only the lowest b bits of e…
The problem of hierarchical clustering items from pairwise similarities is found across various scientific disciplines, from biology to networking. Often, applications of clustering techniques are limited by the cost of obtaining similarities between pairs of items. While prior work has been developed to reconstruct cl…
Study on self-similar solutions of supercritical Fujita equation, proving entropy and energy gap.
In this article we verify an orbifold version of a conjecture of Nimershiem from 1998. Namely, for every flat -manifold , we show that the set of similarity classes of flat metrics on which occur as a cusp cross-section of a hyperbolic -orbifold is dense in the space of similarity classes of flat metri…
Study uses trajectory embedding to measure place function similarity at fine spatial granularity.
Algorithm aggregates rewards from multiple players to learn related tasks in online bandit learning.
There are plenty of problems where the data available is scarce and expensive. We propose a generator of semi-artificial data with similar properties to the original data which enables development and testing of different data mining algorithms and optimization of their parameters. The generated data allow a large scal…
The performance of many machine learning techniques depends on the choice of an appropriate similarity or distance measure on the input space. Similarity learning (or metric learning) aims at building such a measure from training data so that observations with the same (resp. different) label are as close (resp. far) a…
Multi-label classification is an important learning problem with many applications. In this work, we propose a principled similarity-based approach for multi-label learning called SML. We also introduce a similarity-based approach for predicting the label set size. The experimental results demonstrate the effectiveness…
Infinite fractal tree solves shortest connection problem.
Proposes a robust similarity measure for sparse time series data.
For many analytical problems the challenge is to handle huge amounts of available data. However, there are data science application areas where collecting information is difficult and costly, e.g., in the study of geological phenomena, rare diseases, faults in complex systems, insurance frauds, etc. In many such cases,…
BiLRP explains deep similarity models by decomposing scores into feature contributions.
We found a new simple family of Cantor sets whose projections are one-dimensional.
CLS measures dataset similarity through decision rule performance.