FUJI scores similarity of ranked lists more robustly.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Complex classification performance metrics such as the F-measure and Jaccard index are often used, in order to handle class-imbalanced cases such as information retrieval and image segmentation. These performance metrics are not decomposable, that is, they cannot be expressed in a per-example manner, which hinder…
This short article aims at demonstrate that the Intersection over Union (or Jaccard index) is not a submodular function. This mistake has been made in an article which is cited and used as a foundation in another article. The Intersection of Union is widely used in machine learning as a cost function especially for imb…
Two private algorithms estimate Jaccard similarity efficiently.
Two simple proofs of the triangle inequality for the Jaccard distance in terms of nonnegative, monotone, submodular functions are given and discussed.
The probability Jaccard similarity was recently proposed as a natural generalization of the Jaccard similarity to measure the proximity of sets whose elements are associated with relative frequencies or probabilities. In combination with a hash algorithm that maps those weighted sets to compact signatures which allow f…
Stability in clinical prediction models is crucial for transferability between studies, yet has received little attention. The problem is paramount in high dimensional data which invites sparse models with feature selection capability. We introduce an effective method to stabilize sparse Cox model of time-to-events usi…
This project explores several Machine Learning methods to predict movie genres based on plot summaries. Naive Bayes, Word2Vec+XGBoost and Recurrent Neural Networks are used for text classification, while K-binary transformation, rank method and probabilistic classification with learned probability threshold are employe…
The minimization of loss functions is the heart and soul of Machine Learning. In this paper, we propose an off-the-shelf optimization approach that can minimize virtually any non-differentiable and non-decomposable loss function (e.g. Miss-classification Rate, AUC, F1, Jaccard Index, Mathew Correlation Coefficient, etc…
We propose an approach to reduce both computational complexity and data storage requirements for the online positioning stage of a fingerprinting-based indoor positioning system (FIPS) by introducing segmentation of the region of interest (RoI) into sub-regions, sub-region selection using a modified Jaccard index, and …
Proposes a new network to improve nuclei segmentation in histopathology images.
Extends Tanimoto kernel to real-valued functions.
Positive definite kernels are an important tool in machine learning that enable efficient solutions to otherwise difficult or intractable problems by implicitly linearizing the problem geometry. In this paper we develop a set-theoretic interpretation of the Earth Mover's Distance (EMD) and propose Earth Mover's Interse…
It has been noticed that some external CVIs exhibit a preferential bias towards a larger or smaller number of clusters which is monotonic (directly or inversely) in the number of clusters in candidate partitions. This type of bias is caused by the functional form of the CVI model. For example, the popular Rand index (R…
C-MinHash reduces the number of permutations needed for MinHash from thousands to just two.
Efficient methods for sparse random projections improve classification accuracy in very high-dimensional data.
JORC-UMAP improves UMAP by incorporating geometric and topological priors.
Conformal prediction fails under severe feature turnover in COVID-19 supply chain tasks.
CNN automates vitiligo lesion segmentation quickly and accurately.
New algorithms optimize metrics for binary classification with class imbalance.
STRAPSim measures ETF portfolio similarity better than existing methods.
Generative models' evaluation scores can be misleading, leading to inflated grades.
This study examined how the correlation and network structure of 30 global indices and 145 local Korean indices belonging to the KOSPI 200 have changed during the 13-year period, 2000-2012. The correlations among the indices were calculated. The results showed that although the average correlations of the global indice…
Confidence intervals improve evaluation of binary prediction rules in data mining.
ReliefE ranks features faster and better in high-dimensional data.
We propose using five data-driven community detection approaches from social networks to partition the label space for the task of multi-label classification as an alternative to random partitioning into equal subsets as performed by RAkELd: modularity-maximizing fastgreedy and leading eigenvector, infomap, walktrap an…
Develops algorithms for optimizing multi-label metrics with provable guarantees.
C-OPH improves One Permutation Hashing by using a shorter circulant permutation.
Estimates set overlap and similarity using random samples.
The automatic digitizing of paper maps is a significant and challenging task for both academia and industry. As an important procedure of map digitizing, the semantic segmentation section mainly relies on manual visual interpretation with low efficiency. In this study, we select urban planning maps as a representative …
Quantum Clustering is a powerful method to detect clusters in data with mixed density. However, it is very sensitive to a length parameter that is inherent to the Schrödinger equation. In addition, linking data points into clusters requires local estimates of covariance that are also controlled by length parameters. Th…
This study presents a multimodal machine learning model to predict ICD-10 diagnostic codes. We developed separate machine learning models that can handle data from different modalities, including unstructured text, semi-structured text and structured tabular data. We further employed an ensemble method to integrate all…
New STH distance finds patterns in event timeseries without resampling.
New metrics improve performance in imbalanced classification problems.
Paper proposes a feature-wise change detection method for improving indoor positioning accuracy.
We provide a general theoretical analysis of expected out-of-sample utility, also referred to as decision-theoretic classification, for non-decomposable binary classification metrics such as F-measure and Jaccard coefficient. Our key result is that the expected out-of-sample utility for many performance metrics is prov…
In this paper, we introduce the notion of motif closure and describe higher-order ranking and link prediction methods based on the notion of closing higher-order network motifs. The methods are fast and efficient for real-time ranking and link prediction-based applications such as web search, online advertising, and re…
We propose a new class of metrics on sets, vectors, and functions that can be used in various stages of data mining, including exploratory data analysis, learning, and result interpretation. These new distance functions unify and generalize some of the popular metrics, such as the Jaccard and bag distances on sets, Man…
A new clustering evaluation index based on density estimation.
Study on symmetric operators on non-compact manifolds, focusing on their index modulo 2.
New index formula connects numerical and -theoretic indices.
The paper explores global index formulas for one-dimensional holomorphic foliations.
Explain Arnold's proof of the Morse index theorem using Maslov index.
The p-index improves investment performance for NYSE stocks but not for SSE stocks.
Paper introduces danceability index as a new bridge index definition.
Study Whittle index learning algorithms for restless bandits with constant stepsizes.
A new index rebalancing strategy reduces large constituent weights without undesirable effects.
We study bounded pseudoconvex domains in complex Euclidean space. We define an index associated to the boundary and show this new index is equivalent to the Diederich-Fornæss index defined in 1977. This connects the Diederich-Fornæss index to boundary conditions and refines the Levi pseudoconvexity. We also prove the $…