A novel weighted distance improves fuzzy c-means clustering accuracy.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New method calculates Ricci curvature from distances between weighted volumes.
Unified framework evaluates different nearest neighbor classification methods.
Consider a weighted or unweighted k-nearest neighbor graph that has been built on n data points drawn randomly according to some density p on R^d. We study the convergence of the shortest path distance in such graphs as the sample size tends to infinity. We prove that for unweighted kNN graphs, this distance converges …
Fundamental weight systems identified as quantum states.
WE constructs GP kernels for mixed inputs using weighted EDMs.
In a recently published paper [1], it is shown that deep neural networks (DNNs) with random Gaussian weights preserve the metric structure of the data, with the property that the distance shrinks more when the angle between the two data points is smaller. We agree that the random projection setup considered in [1] pres…
We study the use of power weighted shortest path distance functions for clustering high dimensional Euclidean data, under the assumption that the data is drawn from a collection of disjoint low dimensional manifolds. We argue, theoretically and experimentally, that this leads to higher clustering accuracy. We also pres…
DW-KNN improves KNN by integrating distance and neighbor reliability for better prediction accuracy.
We define a class of Euclidean distances on weighted graphs, enabling to perform thermodynamic soft graph clustering. The class can be constructed form the "raw coordinates" encountered in spectral clustering, and can be extended by means of higher-dimensional embeddings (Schoenberg transformations). Geographical flow …
Multi-label classification is a type of supervised learning where an instance may belong to multiple labels simultaneously. Predicting each label independently has been criticized for not exploiting any correlation between labels. In this paper we propose a novel approach, Nearest Labelset using Double Distances (NLDD)…
TAWT improves cross-task learning efficiency and guarantees.
Machine learning is often used in virtual screening to find compounds that are pharmacologically active on a target protein. The weave module is a type of graph convolutional deep neural network that uses not only features focusing on atoms alone (atom features) but also features focusing on atom pairs (pair features);…
Proof of existence and uniqueness of weighted Voronoi-Delaunay on polyhedral surfaces.
Distributional (or distribution-valued) data are a new type of data arising from several sources and are considered as realizations of distributional variables. A new set of fuzzy c-means algorithms for data described by distributional variables is proposed. The algorithms use the Wasserstein distance between dist…
Study shows challenges in converting RNNs to FSMs due to computational complexity.
Nonuniform tubular neighborhoods of curves in Euclidean n-space are studied by using weighted distance functions and generalizing the normal exponential map. Different notions of injectivity radii are introduced to investigate singular but injective exponential maps. A generalization of the thickness formula is obtaine…
Paper defines a new distance metric for comparing learning tasks.
A new method detects small holes in noisy data.
This paper approximates 1-Wasserstein distance using tree-based embedding.
This paper refines MMD for domain adaptation by balancing intra-class and inter-class distances.
Study of metrics on positive-definite matrices from power potential, linking to power means.
A new kNN imputation method improves classification performance on datasets with missing data.
Our work presents extensive empirical evidence that layer rotation, i.e. the evolution across training of the cosine distance between each layer's weight vector and its initialization, constitutes an impressively consistent indicator of generalization performance. In particular, larger cosine distances between final an…
SWRLDA improves LDA for multi-class classification with edge classes.
Quantifying similarity between data objects is an important part of modern data science. Deciding what similarity measure to use is very application dependent. In this paper, we combine insights from systems theory and machine learning, and investigate the weighted cepstral distance, which was previously defined for si…
New bounds for average graph distance using curvature and centrality.
Nearest Neighbors Algorithm is a Lazy Learning Algorithm, in which the algorithm tries to approximate the predictions with the help of similar existing vectors in the training dataset. The predictions made by the K-Nearest Neighbors algorithm is based on averaging the target values of the spatial neighbors. The selecti…
We propose (WIPS) for neural network-based graph embedding. In addition to the parameters of neural networks, we optimize the weights of the inner product by allowing positive and negative values. Despite its simplicity, WIPS can approximate arbitrary general similarities in…
New geometric analysis of PWSPDs balances density and geometry in high-dimensional data.
A new method estimates uncertainty without explicit prediction models.
A Discriminative Deep Forest (DisDF) as a metric learning algorithm is proposed in the paper. It is based on the Deep Forest or gcForest proposed by Zhou and Feng and can be viewed as a gcForest modification. The case of the fully supervised learning is studied when the class labels of individual training examples are …
Deep neural networks' infinite-width behavior approximated by Gaussian models.
A new classifier uses Fermat distance for semi-supervised learning in high dimensions.
A novel linear classification method that possesses the merits of both the Support Vector Machine (SVM) and the Distance-weighted Discrimination (DWD) is proposed in this article. The proposed Distance-weighted Support Vector Machine method can be viewed as a hybrid of SVM and DWD that finds the classification directio…
In this work we study the properties of deep neural networks (DNN) with random weights. We formally prove that these networks perform a distance-preserving embedding of the data. Based on this we then draw conclusions on the size of the training data and the networks' structure. A longer version of this paper with more…
BRDAD uses bagging and regularization to improve anomaly detection without labeled data.
Extends graph encoder embedding to weighted graphs and matrices.
Global optimization problems whose objective function is expensive to evaluate can be solved effectively by recursively fitting a surrogate function to function samples and minimizing an acquisition function to generate new samples. The acquisition step trades off between seeking for a new optimization vector where the…
Improved convergence rates for MLE in mixture models using penalized log-likelihood.
Proves rigidity of circle packings in the plane, generalizing previous work.
A Siamese Deep Forest (SDF) is proposed in the paper. It is based on the Deep Forest or gcForest proposed by Zhou and Feng and can be viewed as a gcForest modification. It can be also regarded as an alternative to the well-known Siamese neural networks. The SDF uses a modified training set consisting of concatenated pa…
Energy distance measures feature heterogeneity in federated learning.
In this paper we study sectional curvature bounds for Riemannian manifolds with density from the perspective of a weighted torsion free connection introduced recently by the last two authors. We develop two new tools for studying weighted sectional curvature bounds: a new weighted Rauch comparison theorem and a modifie…
Most random ReLU networks are vulnerable to small, Euclidean adversarial perturbations.
Clustering is a separation of data into groups of similar objects. Every group called cluster consists of objects that are similar to one another and dissimilar to objects of other groups. In this paper, the K-Means algorithm is implemented by three distance functions and to identify the optimal distance function for c…
In a spatially embedded network, that is a network where nodes can be uniquely determined in a system of coordinates, links' weights might be affected by metric distances coupling every pair of nodes (dyads). In order to assess to what extent metric distances affect relationships (link's weights) in a spatially embedde…
Estimates TV distance between autoregressive models under different access models.