A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
For a certain class of distributions, we prove that the linear programming relaxation of k-medoids clustering---a variant of k-means clustering where means are replaced by exemplars from within the dataset---distinguishes points drawn from nonoverlapping balls with high probability once the number of points drawn a…
We improve deep threshold networks' memorization capacity exponentially.
problem Memorizing datasets with randomized labels using deep neural networks.
method Using Gaussian random weights in the first layer and binary or integer weights in subsequent layers, we prove a new dependence on minimum distance.
result We show that O(δ1+n) neurons and O(δd+n) weights are sufficient.
In this paper we continue our study of the Laplacian on manifolds with axial analytic asymptotically cylindrical ends initiated in~arXiv:1003.2538. By using the complex scaling method and the Phragmén-Lindelöf principle we prove exponential decay of the eigenfunctions corresponding to the non-threshold eigenvalues of t…
Support Vector Data Description (SVDD) is a machine learning technique used for single class classification and outlier detection. SVDD based K-chart was first introduced by Sun and Tsung for monitoring multivariate processes when underlying distribution of process parameters or quality characteristics depart from Norm…
Using data from 92 indices of stock exchanges worldwide, I analize the cluster formation and evolution from 2007 to 2010, which includes the Subprime Mortgage Crisis of 2008, using asset graphs based on distance thresholds. I also study the survivability of connections and of clusters through time and the influence of …
We consider the problem of Gaussian mixture clustering in the high-dimensional limit where the data consists of m points in n dimensions, n,m→∞ and α=m/n stays finite. Using exact but non-rigorous methods from statistical physics, we determine the critical value of α and the distance between…
Adaptive sampling theory has shown that, with proper assumptions on the signal class, algorithms exist to reconstruct a signal in Rd with an optimal number of samples. We generalize this problem to the case of spatial signals, where the sampling cost is a function of both the number of samples taken and t…
Subspace clustering refers to the problem of clustering high-dimensional data points into a union of low-dimensional linear subspaces, where the number of subspaces, their dimensions and orientations are all unknown. In this paper, we propose a variation of the recently introduced thresholding-based subspace clustering…
This work bounds the run-time of nonconvex optimization with early stopping.
problem Bounding the expected run-time of nonconvex optimization with early stopping.
method Derives conditions for well-defined early stopping based on validation function norms and bounds the expected number of iterations and gradient evaluations.
result Guarantees the validity of early stopping and provides bounds on the expected run-time for various optimization algorithms.
The problem of clustering noisy and incompletely observed high-dimensional data points into a union of low-dimensional subspaces and a set of outliers is considered. The number of subspaces, their dimensions, and their orientations are assumed unknown. We propose a simple low-complexity subspace clustering algorithm, w…
A new algorithm of the analysis of correlation among economy time series is proposed. The algorithm is based on the power law classification scheme (PLCS) followed by the analysis of the network on the percolation threshold (NPT). The algorithm was applied to the analysis of correlations among GDP per capita time serie…
This paper deals with prediction of anopheles number, the main vector of malaria risk, using environmental and climate variables. The variables selection is based on an automatic machine learning method using regression trees, and random forests combined with stratified two levels cross validation. The minimum threshol…