GBOC detects anomalies in time series data using granular-ball vectors.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The hierarchical Dirichlet process (HDP) has become an important Bayesian nonparametric model for grouped data, such as document collections. The HDP is used to construct a flexible mixed-membership model where the number of components is determined by the data. As for most Bayesian nonparametric models, exact posterio…
Paper proposes a new method for predicting DER adoption with hierarchical guarantees.
New algorithms for hierarchical classification using conformal prediction.
The study improves prediction regions for hierarchical data using a projection step.
Logarithmic separation profile in hyperbolic groups shows hierarchical structure.
This paper improves hierarchical community detection efficiency using local structural properties.
Clustering is a fundamental problem in many scientific applications. Standard methods such as -means, Gaussian mixture models, and hierarchical clustering, however, are beset by local minima, which are sometimes drastically suboptimal. Recently introduced convex relaxations of -means and hierarchical clustering s…
Paper develops a new objective for hierarchical clustering in Euclidean space.
End-to-end deep model for coherent probabilistic forecasts in hierarchical time series.
CCMM efficiently solves large-scale convex clustering problems.
Improved sampling for network community detection.
We propose a restricted collapsed draw (RCD) sampler, a general Markov chain Monte Carlo sampler of simultaneous draws from a hierarchical Chinese restaurant process (HCRP) with restriction. Models that require simultaneous draws from a hierarchical Dirichlet process with restriction, such as infinite Hidden markov mod…
Bottom-up algorithms outperform top-down in hierarchical community detection at intermediate levels.
Modern vehicles are equipped with increasingly complex sensors. These sensors generate large volumes of data that provide opportunities for modeling and analysis. Here, we are interested in exploiting this data to learn aspects of behaviors and the road network associated with individual drivers. Our dataset is collect…
Generative adversarial networks (GANs) are deep neural networks that allow us to sample from an arbitrary probability distribution without explicitly estimating the distribution. There is a generator that takes a latent vector as input and transforms it into a valid sample from the distribution. There is also a discrim…
Hierarchical clustering is a widely used approach for clustering datasets at multiple levels of granularity. Despite its popularity, existing algorithms such as hierarchical agglomerative clustering (HAC) are limited to the offline setting, and thus require the entire dataset to be available. This prohibits their use o…
SHVC improves image compression with fewer parameters.
New insights into optimizing latent representations in hierarchical VAEs.
Proposes a new cross-validation method to estimate model performance.
We propose a new splitting criterion for a meta-learning approach to multiclass classifier design that adaptively merges the classes into a tree-structured hierarchy of increasingly difficult binary classification problems. The classification tree is constructed from empirical estimates of the Henze-Penrose bounds on t…
Generators found for automorphisms of special groups.
New methods improve prediction intervals across multiple environments.
Due to the escalating growth of big data sets in recent years, new Bayesian Markov chain Monte Carlo (MCMC) parallel computing methods have been developed. These methods partition large data sets by observations into subsets. However, for Bayesian nested hierarchical models, typically only a few parameters are common f…
The Hierarchical Mixture of Experts (HME) is a well-known tree-based model for regression and classification, based on soft probabilistic splits. In its original formulation it was trained by maximum likelihood, and is therefore prone to over-fitting. Furthermore the maximum likelihood framework offers no natural metri…
Study proves consistency of spectral clustering on hierarchical networks.
Hierarchical clustering is a popular method for analyzing data which associates a tree to a dataset. Hartigan consistency has been used extensively as a framework to analyze such clustering algorithms from a statistical point of view. Still, as we show in the paper, a tree which is Hartigan consistent with a given dens…
Annotating large unlabeled datasets can be a major bottleneck for machine learning applications. We introduce a scheme for inferring labels of unlabeled data at a fraction of the cost of labeling the entire dataset. Our scheme, bounded expectation of label assignment (BELA), greedily queries an oracle (or human labeler…
Proposes a method to partition univariate data into unimodal subsets.
As data collections become larger, exploratory regression analysis becomes more important but more challenging. When observations are hierarchically clustered the problem is even more challenging because model selection with mixed effect models can produce misleading results when nonlinear effects are not included into…
The increase in data collection has made data annotation an interesting and valuable task in the contemporary world. This paper presents a new methodology for quickly annotating data using click-supervision and hierarchical object detection. The proposed work is semi-automatic in nature where the task of annotations is…
New tree and forest methods use oblique splits for better risk bounds.
We consider clustering based on significance tests for Gaussian Mixture Models (GMMs). Our starting point is the SigClust method developed by Liu et al. (2008), which introduces a test based on the k-means objective (with k = 2) to decide whether the data should be split into two clusters. When applied recursively, thi…
Paper learns dictionaries for sparse signal recovery using automatic differentiation.
Proposes oblique predictive clustering trees for faster, more efficient learning.
Unified framework for semi-supervised learning reduces annotation needs.
We discuss an autoencoder model in which the encoding and decoding functions are implemented by decision trees. We use the soft decision tree where internal nodes realize soft multivariate splits given by a gating function and the overall output is the average of all leaves weighted by the gating values on their path. …
Paper introduces hierarchical softmax for global hierarchical classification tasks.
Recently proposed budding tree is a decision tree algorithm in which every node is part internal node and part leaf. This allows representing every decision tree in a continuous parameter space, and therefore a budding tree can be jointly trained with backpropagation, like a neural network. Even though this continuity …
This study compares hierarchical and non-hierarchical models for open-domain multi-turn dialog generation.
We survey agglomerative hierarchical clustering algorithms and discuss efficient implementations that are available in R and other software environments. We look at hierarchical self-organizing maps, and mixture models. We review grid-based clustering, focusing on hierarchical density-based approaches. Finally we descr…
A nonparametric method for time series analysis extracts envelopes, detects peaks, and clusters data.
Paper proposes a Renyi entropy-based method for tuning hierarchical topic models.
Neural NMF discovers hierarchical topics in multilayer data.
The paper develops a decision support system for hierarchical text classification of conference proceedings.
We present a comprehensive framework for structured sparse coding and modeling extending the recent ideas of using learnable fast regressors to approximate exact sparse codes. For this purpose, we develop a novel block-coordinate proximal splitting method for the iterative solution of hierarchical sparse coding problem…
Bayesian Hierarchical Invariant Prediction refines ICP for better scalability and prior integration.
Hierarchical causal models help understand cause and effect in nested data.