High density clusters can be characterized by the connected components of a level set of the underlying probability density function generating the data, at some appropriate level . The complete hierarchical clustering can be characterized by a cluster tree ${\cal T}= \bigcup_λ L(λ)…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The clusters of a distribution are often defined by the connected components of a density level set. However, this definition depends on the user-specified level. We address this issue by proposing a simple, generic algorithm, which uses an almost arbitrary level set estimator to estimate the smallest level at which th…
The level set tree approach of Hartigan (1975) provides a probabilistically based and highly interpretable encoding of the clustering behavior of a dataset. By representing the hierarchy of data modes as a dendrogram of the level sets of a density estimator, this approach offers many advantages for exploratory analysis…
We study the connections between spectral clustering and the problems of maximum margin clustering, and estimation of the components of level sets of a density function. Specifically, we obtain bounds on the eigenvectors of graph Laplacian matrices in terms of the between cluster separation, and within cluster connecti…
We show that DBSCAN can estimate the connected components of the -density level set given i.i.d. samples from an unknown density . We characterize the regularity of the level set boundaries using parameter and analyze the estimation error under the Hausdorff metric. When the data …
BDMBC clusters data with varying densities using a new PLLS measure.
Machine learning predicts nuclear physics parameters with high accuracy.
Following Hartigan, a cluster is defined as a connected component of the t-level set of the underlying density, i.e., the set of points for which the density is greater than t. A clustering algorithm which combines a density estimate with spectral clustering techniques is proposed. Our algorithm is composed of two step…
Paper proves stronger Penrose inequality with matter density.
Study recovers Riemannian quantities from noisy data densities.
Machine learning predicts electronic density of states for condensed matter.
Paper uses machine learning to estimate IRI from pavement distress types, densities, and severities.
New bounds on generalization error using information density moments.
SLS optimizes minimum-volume regions for conditional quantiles, bypassing density estimation.
New insights into binary perceptron reveal phase transitions and algorithmic thresholds.
Generative models improved with smoothed score functions for better sample quality.
In this article we use the mean curvature flow with surgery to derive regularity estimates for the level set flow going past Brakke regularity in certain special conditions allowing for 2-convex regions of high density. We also show a stability result for the plane under the level set flow.
Single-level density-based approach has long been widely acknowledged to be a conceptually and mathematically convincing clustering method. In this paper, we propose an algorithm called "best-scored clustering forest" that can obtain the optimal level and determine corresponding clusters. The terminology "best-scored" …
New scoring rules for multivariate distributions and level sets.
MRCNet tackles crowd counting and density mapping in aerial imagery.
GCAO improves clustering of high-dimensional data by grouping low-density boundary points.
Hierarchical VAEs detect out-of-distribution data by identifying low-level in-distribution features.
We derive and analyze a generic, recursive algorithm for estimating all splits in a finite cluster tree as well as the corresponding clusters. We further investigate statistical properties of this generic clustering algorithm when it receives level set estimates from a kernel density estimator. In particular, we derive…
The L1 loss landscape of neural nets near local minima behaves differently, revealing exponential decay and increased vertex density.
New framework quantifies uncertainty in flexible density-based clustering.
This work explores efficient reinforcement learning with density features in low-rank MDPs.
LGKDE learns graph density using neural networks and perturbations.
New model predicts grain boundary migration in metals.
We compute, using a formula of Dittmann, the Bures metric tensor (g) for the eight-dimensional convex set of three-level quantum systems, employing a newly-developed Euler angle-based parameterization of the 3 x 3 density matrices. Most of the individual metric elements (g_{ij}) are found to be expressible in relativel…
New method uses random projections to estimate densities and modes efficiently.
We study in this paper the rate of convergence for learning densities under the Generative Adversarial Networks (GAN) framework, borrowing insights from nonparametric statistics. We introduce an improved GAN estimator that achieves a faster rate, through simultaneously leveraging the level of smoothness in the target d…
A clustering method for multivariate populations with similar dependence structures.
An efficient way to learn deep density models that have many layers of latent variables is to learn one layer at a time using a model that has only one layer of latent variables. After learning each layer, samples from the posterior distributions for that layer are used as training data for learning the next layer. Thi…
DBSCAN is a classical density-based clustering procedure with tremendous practical relevance. However, DBSCAN implicitly needs to compute the empirical density for each sample point, leading to a quadratic worst-case time complexity, which is too slow on large datasets. We propose DBSCAN++, a simple modification of DBS…
Study area and coarea formulas for graphs and submanifolds in Carnot groups.
Let be a compact, connected Riemannian manifold whose Riemannian volume measure is denoted by . Let be a non-constant eigenfunction of the Laplacian. The random wave conjecture suggests that in certain situations, the value distribution of under is approximately Gaussian. Wr…
We develop a new Low-level, First-order Probabilistic Programming Language (LF-PPL) suited for models containing a mix of continuous, discrete, and/or piecewise-continuous variables. The key success of this language and its compilation scheme is in its ability to automatically distinguish parameters the density functio…
We study time reversal, last passage time, and -transform of linear diffusions. For general diffusions with killing, we obtain the probability density of the last passage time to an arbitrary level and analyze the distribution of the time left until killing after the last passage time. With these tools, we develop a…
Personalized medicine seeks to identify the causal effect of treatment for a particular patient as opposed to a clinical population at large. Most investigators estimate such personalized treatment effects by regressing the outcome of a randomized clinical trial (RCT) on patient covariates. The realized value of the ou…
Study shows zero level sets of solutions to Allen-Cahn equation are minimal surfaces with zero mean curvature.
We introduce multiplicative LSTM (mLSTM), a recurrent neural network architecture for sequence modelling that combines the long short-term memory (LSTM) and multiplicative recurrent neural network architectures. mLSTM is characterised by its ability to have different recurrent transition functions for each possible inp…
Study examines financial contagion at community level, finding increased contagion density and widespread transmission.
The paper proposes a new method for probabilistic load forecasting using Bernstein-Polynomial Normalizing Flows.
We introduce a balloon estimator in a generalized expectation-maximization method for estimating all parameters of a Gaussian mixture model given one data sample per mixture component. Instead of limiting explicitly the model size, this regularization strategy yields low-complexity sparse models where the number of eff…
QNA uses quantum-inspired density operators to diagnose market dependence and structural risk.
New findings suggest deep generative models can misclassify outliers, requiring new evaluation methods.
SPQR package uses neural networks for flexible quantile regression.
A new method for density estimation using mixture discrepancy and moments.