The adjusted Rand index (ARI) is commonly used in cluster analysis to measure the degree of agreement between two data partitions. Since its introduction, exploring the situations of extreme agreement and disagreement under different circumstances has been a subject of interest, in order to achieve a better understandi…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Proposes new random models for fuzzy clustering similarity measures.
The misclassification error distance and the adjusted Rand index are two of the most commonly used criteria to evaluate the performance of clustering algorithms. This paper provides an in-depth comparison of the two criteria, aimed to better understand exactly what they measure, their properties and their differences. …
Meta-learning neural networks for better clustering representations.
In unsupervised machine learning, agreement between partitions is commonly assessed with so-called external validity indices. Researchers tend to use and report indices that quantify agreement between two partitions for all clusters simultaneously. Commonly used examples are the Rand index and the adjusted Rand index. …
The goal of lifetime clustering is to develop an inductive model that maps subjects into clusters according to their underlying (unobserved) lifetime distribution. We introduce a neural-network based lifetime clustering model that can find cluster assignments by directly maximizing the divergence between the empiri…
Paper addresses xVA models for market-implied skew and smile.
Adjusted for chance measures are widely used to compare partitions/clusterings of the same data set. In particular, the Adjusted Rand Index (ARI) based on pair-counting, and the Adjusted Mutual Information (AMI) based on Shannon information theory are very popular in the clustering community. Nonetheless it is an open …
The main goal of this study is to extract a set of brain networks in multiple time-resolutions to analyze the connectivity patterns among the anatomic regions for a given cognitive task. We suggest a deep architecture which learns the natural groupings of the connectivity patterns of human brain in multiple time-resolu…
A new measure DCSI quantifies separability for density-based clustering.
A system is presented that segments, clusters and predicts musical audio in an unsupervised manner, adjusting the number of (timbre) clusters instantaneously to the audio input. A sequence learning algorithm adapts its structure to a dynamically changing clustering tree. The flow of the system is as follows: 1) segment…
In mixture model-based clustering applications, it is common to fit several models from a family and report clustering results from only the `best' one. In such circumstances, selection of this best model is achieved using a model selection criterion, most often the Bayesian information criterion. Rather than throw awa…
A new measure normalizes clustering accuracy to evaluate algorithms better.
CDL index improves clustering validation for non-convex data.
Study EM and GD for clustering with penalties for misspecification and high dimensions.
A new clustering framework optimizes customer search data for personalized travel recommendations.
Unified framework for comparing clusterings from information-theoretic and pair-counting perspectives.
STICC clusters geographic objects considering both spatial contiguity and attributes.
Mixtures of Unigrams are one of the simplest and most efficient tools for clustering textual data, as they assume that documents related to the same topic have similar distributions of terms, naturally described by Multinomials. When the classification task is particularly challenging, such as when the document-term ma…
Study compares clustering methods for mixed-type data.
KLD token adjusts supply based on macroeconomic debt index, creating deflationary effect.
This paper identifies and analyzes biases in risk-adjusted index weighting methods, affecting social welfare and market fairness.
New star-shaped acceptability indexes generalize existing methods.
New method approximates M-estimator and predictions without solving fixed-point equations.
New metric improves clustering in persistent homology.
New k-means method handles random data better than traditional techniques.
Paper introduces Arte-Blue Chip Index for diversifying portfolios with art investments.
The study of genetic variants can help find correlating population groups to identify cohorts that are predisposed to common diseases and explain differences in disease susceptibility and how patients react to drugs. Machine learning algorithms are increasingly being applied to identify interacting GVs to understand th…
Randomizes AD models for better option pricing.
Visual summarization of clinical data collected on patients contained within the electronic health record (EHR) may enable precise and rapid triage at the time of patient presentation to an emergency department (ED). The triage process is critical in the appropriate allocation of resources and in anticipating eventual …
The paper investigates how irrelevant features affect clustering performance.
New method clusters matrix-valued data by latent variables.
We consider the problem of pricing derivatives written on some industrial loss index via utility indifference pricing. The industrial loss index is modelled by a compound Poisson process and the insurer can adjust her portfolio by choosing the risk loading, which in turn determines the demand. We compute the price of a…
It has been noticed that some external CVIs exhibit a preferential bias towards a larger or smaller number of clusters which is monotonic (directly or inversely) in the number of clusters in candidate partitions. This type of bias is caused by the functional form of the CVI model. For example, the popular Rand index (R…
Exchange Traded Funds (ETFs) have been gaining increasing popularity in the investment community as is evidenced by the high growth both in the number of ETFs and their net assets since 2000. As ETFs are in nature similar to index mutual funds, in this paper we examined if this growing demand for ETFs can be explained …
The Hype Index measures media attention to equities using NLP.
Unsupervised image segmentation aims at clustering the set of pixels of an image into spatially homogeneous regions. We introduce here a class of Bayesian nonparametric models to address this problem. These models are based on a combination of a Potts-like spatial smoothness component and a prior on partitions which is…
Study examines volatility-based strategy for Chinese ETF options, improving returns in volatile markets.
This paper proposes a Lasso-type estimator for a high-dimensional sparse parameter identified by a single index conditional moment restriction (CMR). In addition to this parameter, the moment function can also depend on a nuisance function, such as the propensity score or the conditional choice probability, which we es…
Sentiment analysis of DAX40 stocks improves performance by 5.38% annually.
The MAXFLAT low-pass filter improves factor adjustment for better portfolio performance in China's stock market.
Enhanced indexation uses equity and index options for better performance.
Study improves MACD trading strategy with volume and price adjustments.
The paper introduces mortgage-rate-adjusted home prices to help buyers and adjust housing indices.
We note a simple mechanism that may at least partially resolve several outstanding economic puzzles, including why the cyclically adjusted price to earnings ratio of the S&P 500 index has been oddly high for the past two decades, why gains to capital have outpaced gains to wages, and the persistence of the equity premi…
DOPE efficiently estimates ATE with complex covariates.
Two types of zeroth-order stochastic algorithms have recently been designed for nonconvex optimization respectively based on the first-order techniques SVRG and SARAH/SPIDER. This paper addresses several important issues that are still open in these methods. First, all existing SVRG-type zeroth-order algorithms suffer …
Paper defines and proves a new analytic index for Fredholm operators.