Automates similarity measure construction from data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Unified understanding of neural representation similarity measures.
A personalized learning system needs a large pool of items for learners to solve. When working with a large pool of items, it is useful to measure the similarity of items. We outline a general approach to measuring the similarity of items and discuss specific measures for items used in introductory programming. Evaluat…
New stability measures for similar features improve feature selection accuracy.
Quantum networks learn task-dependent asymmetric similarity measures.
Method measures weight similarity in neural networks using normalization and statistical inference.
Paper develops multivariate time series similarity and distance measures.
In this paper, we propose a family of graph partition similarity measures that take the topology of the graph into account. These graph-aware measures are alternatives to using set partition similarity measures that are not specifically designed for graph partitions. The two types of measures, graph-aware and set parti…
New measure quantifies function similarity for optimization.
Proposes a method to measure similarity between anomaly scores from different methods.
Clustering is an underspecified task: there are no universal criteria for what makes a good clustering. This is especially true for relational data, where similarity can be based on the features of individuals, the relationships between them, or a mix of both. Existing methods for relational clustering have strong and …
Proposes a robust similarity measure for sparse time series data.
A good measure of similarity between data points is crucial to many tasks in machine learning. Similarity and metric learning methods learn such measures automatically from data, but they do not scale well respect to the dimensionality of the data. In this paper, we propose a method that can learn efficiently similarit…
Defines a similarity measure for classification distributions.
Modified cosine distance improves similarity performance in data with variance and correlation.
CLS measures dataset similarity through decision rule performance.
Fractal Lipschitz-Killing curvature measures C^f_k(F,.), k = 0, ..., d, are determined for a large class of self-similar sets F in R^d. They arise as weak limits of the appropriately rescaled classical Lipschitz-Killing curvature measures C_k(F_r,.) from geometric measure theory of parallel sets F_r for small distances…
New clustering method using point-set kernel measures similarity.
Develops an ordinal-similarity framework for scalable and interpretable representation alignment.
Recent work has sought to understand the behavior of neural networks by comparing representations between layers and between different trained models. We examine methods for comparing neural network representations based on canonical correlation analysis (CCA). We show that CCA belongs to a family of statistics for mea…
Time series are ubiquitous, and a measure to assess their similarity is a core part of many computational systems. In particular, the similarity measure is the most essential ingredient of time series clustering and classification systems. Because of this importance, countless approaches to estimate time series similar…
Study uses trajectory embedding to measure place function similarity at fine spatial granularity.
Recently there has been an increase in the studies on time-series data mining specifically time-series clustering due to the vast existence of time-series in various domains. The large volume of data in the form of time-series makes it necessary to employ various techniques such as clustering to understand the data and…
Cross-domain visual data matching is one of the fundamental problems in many real-world vision tasks, e.g., matching persons across ID photos and surveillance videos. Conventional approaches to this problem usually involves two steps: i) projecting samples from different domains into a common space, and ii) computing (…
STRAPSim measures ETF portfolio similarity better than existing methods.
Producing overlapping schemes is a major issue in clustering. Recent proposed overlapping methods relies on the search of an optimal covering and are based on different metrics, such as Euclidean distance and I-Divergence, used to measure closeness between observations. In this paper, we propose the use of another meas…
Generalizes Black-Scholes model for option pricing under uncertainty.
The study explores how to infer the geometry of space forms from similarity comparisons.
In this study, we establish a basis for selecting similarity measures when applying machine learning techniques to solve materials science problems. This selection is considered with an emphasis on the distinctiveness between materials that reflect their nature well. We perform a case study with a dataset of rare-earth…
Despite the success of the popular kernelized support vector machines, they have two major limitations: they are restricted to Positive Semi-Definite (PSD) kernels, and their training complexity scales at least quadratically with the size of the data. Many natural measures of similarity between pairs of samples are not…
Proposes a new stability measure for model fitting on similar feature data sets.
New measures link neural representation geometry to decoding ability.
Spectral Clustering(SC) is a prominent data clustering technique of recent times which has attracted much attention from researchers. It is a highly data-driven method and makes no strict assumptions on the structure of the data to be clustered. One of the central pieces of spectral clustering is the construction of an…
kdiff measures distances for time series and structured data.
New framework to test neural network representation similarity measures.
Inner product-based convolution has been the founding stone of convolutional neural networks (CNNs), enabling end-to-end learning of visual representation. By generalizing inner product with a bilinear matrix, we propose the neural similarity which serves as a learnable parametric similarity measure for CNNs. Neural si…
From a sequence of similarity networks, with edges representing certain similarity measures between nodes, we are interested in detecting a change-point which changes the statistical property of the networks. After the change, a subset of anomalous nodes which compares dissimilarly with the normal nodes. We study a sim…
Paper proposes a supervised similarity framework for corporate bonds using RF proximities.
In many situations, the choice of an adequate similarity measure or metric on the feature space dramatically determines the performance of machine learning methods. Building automatically such measures is the specific purpose of metric/similarity learning. In Vogel et al. (2018), similarity learning is formulated as a …
Similarity between objects is multi-faceted and it can be easier for human annotators to measure it when the focus is on a specific aspect. We consider the problem of mapping objects into view-specific embeddings where the distance between them is consistent with the similarity comparisons of the form "from the t-th vi…
Supervised learning has been very successful for automatic segmentation of images from a single scanner. However, several papers report deteriorated performances when using classifiers trained on images from one scanner to segment images from other scanners. We propose a transfer learning classifier that adapts to diff…
A new stable similarity measure for time series using persistent homology.
We propose a family of near-metrics based on local graph diffusion to capture similarity for a wide class of data sets. These quasi-metametrics, as their names suggest, dispense with one or two standard axioms of metric spaces, specifically distinguishability and symmetry, so that similarity between data points of arbi…
International trade has been increasingly organized in the form of global value chains (GVCs) where different stages of production are located in different countries. This recent phenomenon has substantial consequences for both trade policy design at the national or regional level and business decision making at the fi…
In [Mor], we have introduced a notion of flat laminations on surfaces endowed with a flat structure, similar to geodesic laminations on hyperbolic surfaces. Here is a sequel to this article that aims at defining transversal measures on flat laminations similar to transversal measures on hyperbolic laminations, taking i…
Improves confidence calibration in neural networks by smoothing labels based on class similarity.
We propose a method to visualize class similarity in large-scale classifiers.
This article explores some properties of universal covers of compact Kahler manifolds, under the assumption of Caratheodory measure hyperbolicity. In particular, by comparing invariant volume forms, an inequality is established between the volume of canonical bundle of a compact Kahler manifolds and the Caratheodory me…