A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
We consider the problem of clustering a set of high-dimensional data points into sets of low-dimensional linear subspaces. The number of subspaces, their dimensions, and their orientations are unknown. We propose a simple and low-complexity clustering algorithm based on thresholding the correlations between the data po…
The paper characterizes spin initial data sets saturating the BPS bound in asymptotically AdS spacetimes.
problem Characterizing spin initial data sets saturating the BPS bound in asymptotically AdS spacetimes.
method The paper introduces a theorem for replacing imaginary Killing spinors with strictly timelike or null ones and uses spinors to construct a codimension-2 slicing.
result The paper establishes a sharp dimension threshold for saturating the BPS bound in gravitational waves and rotating black holes in higher dimensions.
We show that for a K-unstable Fano variety, any divisorial valuation computing its stability threshold induces a non-trivial special test configuration preserving the stability threshold. When such a divisorial valuation exists, we show that the Fano variety degenerates to a uniquely determined twisted K-polystable Fan…
Let M be a compact, unit volume, Riemannian manifold with boundary. In this paper we study the homology of a random Čech-complex generated by a homogeneous Poisson process in M. Our main results are two asymptotic threshold formulas, an upper threshold above which the Čech complex recovers the k-th homology of $M…
Recently, a novel family of biologically plausible online algorithms for reducing the dimensionality of streaming data has been derived from the similarity matching principle. In these algorithms, the number of output dimensions can be determined adaptively by thresholding the singular values of the input data matrix. …
We show that every approximately differentially private learning algorithm (possibly improper) for a class H with Littlestone dimension~d requires Ω(log∗(d)) examples. As a corollary it follows that the class of thresholds over N can not be learned in a private manner; this resolves open qu…
We consider the problem of clustering noisy high-dimensional data points into a union of low-dimensional subspaces and a set of outliers. The number of subspaces, their dimensions, and their orientations are unknown. A probabilistic performance analysis of the thresholding-based subspace clustering (TSC) algorithm intr…
In this article, we consider the sparse tensor singular value decomposition, which aims for dimension reduction on high-dimensional high-order data with certain sparsity structure. A method named Sparse Tensor Alternating Thresholding for Singular Value Decomposition (STAT-SVD) is proposed. The proposed procedure featu…
In this paper we show that the computational complexity of the Iterative Thresholding and K-residual-Means (ITKrM) algorithm for dictionary learning can be significantly reduced by using dimensionality-reduction techniques based on the Johnson-Lindenstrauss lemma. The dimensionality reduction is efficiently carried out…
We improve deep threshold networks' memorization capacity exponentially.
problem Memorizing datasets with randomized labels using deep neural networks.
method Using Gaussian random weights in the first layer and binary or integer weights in subsequent layers, we prove a new dependence on minimum distance.
result We show that O(δ1+n) neurons and O(δd+n) weights are sufficient.
High-dimensional sparse modeling via regularization provides a powerful tool for analyzing large-scale data sets and obtaining meaningful, interpretable models. The use of nonconvex penalty functions shows advantage in selecting important features in high dimensions, but the global optimality of such methods still dema…
Deep Gaussian processes can have non-degenerate and non-Gaussian limits.
problem Understanding the behavior of deep Gaussian processes as depth grows.
method Studying the limit of compositional Gaussian processes where each layer is a Gaussian process.
result Identified a sharp bandwidth threshold above which the limit is degenerate, and proved that for bandwidths below this threshold, the limit is a non-degenerate and non-Gaussian distribution.
The problem of clustering noisy and incompletely observed high-dimensional data points into a union of low-dimensional subspaces and a set of outliers is considered. The number of subspaces, their dimensions, and their orientations are assumed unknown. We propose a simple low-complexity subspace clustering algorithm, w…
We consider the problem of Gaussian mixture clustering in the high-dimensional limit where the data consists of m points in n dimensions, n,m→∞ and α=m/n stays finite. Using exact but non-rigorous methods from statistical physics, we determine the critical value of α and the distance between…
Iterative thresholding algorithms seek to optimize a differentiable objective function over a sparsity or rank constraint by alternating between gradient steps that reduce the objective, and thresholding steps that enforce the constraint. This work examines the choice of the thresholding operator, and asks whether it i…