Selective inference controls Type I error in k-means clustering tests.
problem Inflated Type I error in classical hypothesis tests for k-means clusters.
method Selective inference approach to control Type I error.
result Proposes a computable finite-sample p-value for selective inference.
Randomized hierarchical clustering tests for stability and detects clusters.
problem Greedy hierarchical clustering's sensitivity to data perturbations.
method Randomization scheme and p-values at each node.
result Valid hypothesis testing procedures for clustering results.
ElbowSig assesses clustering structure at multiple scales.
problem Selecting optimal number of clusters in unsupervised learning.
method Formalizes elbow heuristic with a normalized discrete curvature statistic.
result Validates multiscale clustering structure over various resolutions.
Community detection in networks is a key exploratory tool with applications in a diverse set of areas, ranging from finding communities in social and biological networks to identifying link farms in the World Wide Web. The problem of finding communities or clusters in a network has received much attention from statisti…
Study finds price-based clustering outperforms AI and human methods in stock market analysis.
problem Investigates if AI can improve stock clustering compared to traditional methods.
method Compares price-based, human-informed, and AI-driven clustering methods using synthetic factor models.
result Price-based clustering reduces RMSE by 15.9% relative to GICS and 14.7% relative to LLM embeddings.
Unsupervised clustering is one of the most fundamental challenges in machine learning. A popular hypothesis is that data are generated from a union of low-dimensional nonlinear manifolds; thus an approach to clustering is identifying and separating these manifolds. In this paper, we present a novel approach to solve th…
We describe our language-independent unsupervised word sense induction system. This system only uses topic features to cluster different word senses in their global context topic space. Using unlabeled data, this system trains a latent Dirichlet allocation (LDA) topic model then uses it to infer the topics distribution…
This paper introduces a novel clustering algorithm for heteroscedastic Gaussian data without needing to know the number of clusters.
problem Clustering heteroscedastic Gaussian data without prior knowledge of the number of clusters.
method Introduces a novel cost function and fixed-point analysis to estimate centroids, introduces Wald kernel for measurement plausibility, and derives CENTRE-X algorithm.
result CENTRE-X algorithm can estimate centroids without prior knowledge of the number of clusters and performs comparably to standard algorithms K-means and Mean-Shift.
Cryptocurrencies show varying levels of efficiency over time, forming clusters with younger ones mimicking older ones.
problem Determining the efficiency of cryptocurrencies over time.
method Permutation entropy and statistical complexity over sliding time-windows of price log returns.
result 37% of cryptocurrencies are efficient over 80% of the time, while 20% are efficient in less than 20% of the time.
A new method, InfoGuide, improves automatic clustering analysis.
problem Lack of automatic clustering analysis frameworks.
method Capturing traces of information gain between clustering retrievals.
result InfoGuide can enable more automatic clustering analysis.
Network Lasso improves semi-supervised regression on network data.
problem Improving regression accuracy on network data with limited labeled examples.
method Applying network Lasso to semi-supervised regression problems, leveraging message passing over an empirical graph.
result Network Lasso's accuracy is linked to the existence of large network flows over the empirical graph.
A Kyle-inspired model with adaptive agents explains excess volatility and volatility clustering.
problem Reconciling asymmetrically informed traders with adaptive market hypothesis.
method Proposes a model with adaptive agents using inductive reasoning, reconciling Kyle model with Adaptive Market Hypothesis.
result Microfoundations for GARCH models and volatility clustering explained.
Proposes selective inference for testing differences in means between clusters.
problem Inflated type I error rate when testing differences in means between clusters.
method Selective inference approach to control selective type I error rate.
result Controls selective type I error rate by accounting for data-driven cluster definition.
In sensor networks, it is not always practical to set up a fusion center. Therefore, there is need for fully decentralized clustering algorithms. Decentralized clustering algorithms should minimize the amount of data exchanged between sensors in order to reduce sensor energy consumption. In this respect, we propose one…
A neuro-inspired architecture learns without supervision using clustering and predictive coding.
problem Achieving continual learning without supervision.
method Neuro-inspired architecture based on online clustering and hierarchical predictive coding.
result The architecture achieves continual learning without supervision.
The paper reviews methods for determining the number of communities in network data.
problem Determining the number of communities in network data.
method Statistical methods for hypothesis testing and clustering in network models.
result SCORE and NCV methods evaluated for clustering in Degree-Corrected Block Models, with NCV facing challenges.
Efficiently clusters survival curves without computationally intensive resampling.
problem Identifying clusters of survival curves efficiently and scalably.
method Log-rank test combined with k-means clustering.
result Achieves comparable results to bootstrap-based methods but with improved efficiency.
Active learning (AL) repeatedly trains the classifier with the minimum labeling budget to improve the current classification model. The training process is usually supervised by an uncertainty evaluation strategy. However, the uncertainty evaluation always suffers from performance degeneration when the initial labeled …
The paper confirms two groups of gamma-ray bursts using a new nonparametric metric.
problem Determining the number of inherent groups in gamma-ray bursts.
method A new nonparametric interpoint distance-based measure, combined with clustering methods.
result Confirms two groups of short and long gamma-ray bursts.
This paper uses UOT metrics for better dimensionality reduction and classification/clustering.
problem Improving dimensionality reduction and classification/clustering methods.
method Uses Hellinger--Kantorovich metric from unbalanced optimal transport (UOT).
result UOT outperforms Euclidean and OT-based methods in classification and clustering tasks.
This study links blockchain design to cryptos' distributional characteristics.
problem Understanding the relationship between blockchain design and cryptos' distributional characteristics.
method Used spectral clustering to cluster cryptos based on their blockchain mechanisms and operational features.
result Clusters of cryptos share similar blockchain mechanisms, supporting the hypothesis.
Geometric QHD tests improve hub detection in correlated data.
problem Detecting hubs in correlated data with evolving correlations.
method Geometric QHD tests combining QCD and QHD, clustering.
result Improved hub detection in correlated data.
Machine learning systems increasingly depend on pipelines of multiple algorithms to provide high quality and well structured predictions. This paper argues interaction effects between clustering and prediction (e.g. classification, regression) algorithms can cause subtle adverse behaviors during cross-validation that m…
A novel density-based approach QC detects outliers in data with high precision.
problem Detecting outliers in data with high precision and sensitivity.
method Quantum Clustering (QC) approach based on the density of data points.
result QC effectively finds hidden outliers and subtle outliers with parameter adjustment.
Three approaches for personalized models in machine learning.
problem Training a single model for all users is not optimal in many scenarios.
method User clustering, data interpolation, and model interpolation.
result Learning-theoretic guarantees and efficient algorithms for personalized models.
The formation of price in a financial market is modelled as a chain of Ising spin with three fundamental figures of trading. We investigate the time behaviour of the model, and we compare the results with the real EURO/USD change rate. By using the test of local Poisson hypothesis, we show that this minimal model leads…
Consider a two-class clustering problem where we observe Xi=ℓiμ+Zi, Zi∼iidN(0,Ip), 1≤i≤n. The feature vector μ∈Rp is unknown but is presumably sparse. The class labels ℓi∈{−1,1} are also unknown and the main interest is to estimate them. We are interested …
In this paper, we consider a generic probabilistic discriminative learner from the functional viewpoint and argue that, to make it learn well, it is necessary to constrain its hypothesis space to a set of non-trivial piecewise constant functions. To achieve this goal, we present a scalable unsupervised regularization f…
This article investigates the correlation structure of the global crude oil market using the daily returns of 71 oil price time series across the world from 1992 to 2012. We identify from the correlation matrix six clusters of time series exhibiting evident geographical traits, which supports Weiner's (1991) regionaliz…
Paper proposes clustering model for ICC based on histologic patterns.
problem Challenges in grading rare cancers like ICC due to small sample sizes and difficulty in extracting patterns.
method Unsupervised deep convolutional autoencoder clustering model trained on 246 ICC digitized slides.
result Three clusters significantly associated with recurrence-free survival in Cox-proportional hazard models.
The paper develops methods for causal function estimation and inference with multiway clustered data.
problem Estimation and inference for causal functions under multiway clustering.
method Two-step procedure using machine learning for nuisance parameters and projection onto basis functions.
result Rejects the null hypothesis of uniformly zero effects and reveals heterogeneous treatment effects.
A new learning method uses data to learn from large model sets.
problem Learning with large sets of candidate models where uniform convergence is hard.
method Data-dependent learning that incorporates empirical data less reliant on prior assumptions.
result Demonstrates improved generalization in various learning assumptions.
New logic approach to machine learning prediction.
problem Predicting based on finite samples.
method Formalized measure of belief violations in modal Logic of Observations and Hypotheses (LOH).
result Machine learning algorithms minimize their version of incongruity.
Efficient algorithm for self-directed learning of convex clusters on graphs.
problem Self-directed classification of nodes on graphs with convex clusters.
method Developed efficient algorithms for (geodesically) convex clusters on graphs.
result Polynomial runtime algorithm with 3(h(G)+1)4lnn mistakes for graphs with two convex clusters. The Morse-Smale complex of a function f decomposes the sample space into cells where f is increasing or decreasing. When applied to nonparametric density estimation and regression, it provides a way to represent, visualize, and compare multivariate functions. In this paper, we present some statistical results on es…
We present five methods to the problem of network anomaly detection. These methods cover most of the common techniques in the anomaly detection field, including Statistical Hypothesis Tests (SHT), Support Vector Machines (SVM) and clustering analysis. We evaluate all methods in a simulated network that consists of nomi…
Paper detects and estimates breaks in high-dimensional functional time series.
problem Detecting and estimating structural breaks in heterogeneous mean functions of high-dimensional functional time series.
method Proposes a new test statistic combining functional CUSUM and power enhancement components, with a clustering algorithm for group structure estimation.
result The proposed techniques have satisfactory performance in finite samples, detecting and estimating breaks effectively.
Despite our extensive knowledge of biophysical properties of neurons, there is no commonly accepted algorithmic theory of neuronal function. Here we explore the hypothesis that single-layer neuronal networks perform online symmetric nonnegative matrix factorization (SNMF) of the similarity matrix of the streamed data. …
Local elasticity in neural networks makes predictions resilient to dissimilar updates.
problem Understanding resilience of neural network predictions to updates from dissimilar data.
method Simulation and geometric interpretation using neural tangent kernel.
result Local elasticity persists in neural networks with nonlinear activation functions, not in linear ones.
A tick size is the smallest increment of a security price. It is clear that at the shortest time scale on which individual orders are placed the tick size has a major role which affects where limit orders can be placed, the bid-ask spread, etc. This is the realm of market microstructure and there is a vast literature o…
It is widely believed that fluctuations in transaction volume, as reflected in the number of transactions and to a lesser extent their size, are the main cause of clustered volatility. Under this view bursts of rapid or slow price diffusion reflect bursts of frequent or less frequent trading, which cause both clustered…
Method uses TV minimization for semi-supervised learning on network data.
problem Semi-supervised learning from partially-labeled network data.
method Graph signal recovery interpretation, total variation minimization, primal-dual method for non-smooth convex optimization.
result TV minimization recovers clusters in the empirical graph of the data under certain network conditions.
Leveraged ETFs can outperform their targets in certain market conditions, contrary to the volatility drag hypothesis.
problem The long-term performance decay of leveraged ETFs due to volatility drag.
method Unified framework incorporating AR(1) and AR-GARCH models, continuous-time regime switching, and flexible rebalancing frequencies.
result Return dynamics, including return autocorrelation, volatility clustering, and regime persistence, determine LETF performance.
Probabilistic embeddings improve speaker diarization accuracy.
problem Improving speaker diarization accuracy using embeddings.
method Extracting x-vectors and precision matrices from speech segments, interfacing with PLDA model, applying agglomerative clustering, joint training of PLDA and extractor.
result Joint training of PLDA and probabilistic x-vector extractor yields accuracy gains.
A new method detects concept drift in streaming data using k-means space partitioning.
problem Detecting distribution changes in streaming data.
method Equal intensity k-means space partitioning (EI-kMeans) and heuristic sensitivity improvement.
result EI-kMeans improves drift detection accuracy and sensitivity.
Motivated by community detection, we characterise the spectrum of the non-backtracking matrix B in the Degree-Corrected Stochastic Block Model. Specifically, we consider a random graph on n vertices partitioned into two equal-sized clusters. The vertices have i.i.d. weights {φu}u=1n with second moment $Φ…
Based on the Aristotelian concept of potentiality vs. actuality allowing for the study of energy and dynamics in language, we propose a field approach to lexical analysis. Falling back on the distributional hypothesis to statistically model word meaning, we used evolving fields as a metaphor to express time-dependent c…
A framework for hypothesis testing on attributed graphs using sampling.
problem Statistical testing on graph data, especially large attributed graphs.
method Sampling-based framework with PHASE and PHASEopt for accurate and efficient hypothesis testing.
result PHASE and PHASEopt improve accuracy and efficiency of hypothesis testing in attributed graphs.