Develops an online nonparametric classifier for massive data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study builds a classifier for diffusions with unknown diffusion but known drifts.
Flexible classifier using Mahalanobis distances for non-elliptical distributions.
Active learning can't improve over passive in certain settings.
A new method detects changes in multivariate data using random forests.
Combines coarse learners for nonparametric probabilistic regression.
InClass nets use neural networks to estimate CIMMs without assuming fixed parameters.
Unsupervised classification methods learn a discriminative classifier from unlabeled data, which has been proven to be an effective way of simultaneously clustering the data and training a classifier from the data. Various unsupervised classification methods obtain appealing results by the classifiers learned in an uns…
Skew Gaussian Processes improve classification performance by allowing asymmetry.
When modeling a probability distribution with a Bayesian network, we are faced with the problem of how to handle continuous variables. Most previous work has either solved the problem by discretizing, or assumed that the data are generated by a single Gaussian. In this paper we abandon the normality assumption and inst…
Bayesian nonparametric method segments multi-sequence time series data.
Stream mining poses unique challenges to machine learning: predictive models are required to be scalable, incrementally trainable, must remain bounded in size (even when the data stream is arbitrarily long), and be nonparametric in order to achieve high accuracy even in complex and dynamic environments. Moreover, the l…
Paper tackles nonparametric classification with privacy constraints, achieving optimal accuracy.
Bayesian learning is built on an assumption that the model space contains a true reflection of the data generating mechanism. This assumption is problematic, particularly in complex data environments. Here we present a Bayesian nonparametric approach to learning that makes use of statistical models, but does not assume…
Human learners have the natural ability to use knowledge gained in one setting for learning in a different but related setting. This ability to transfer knowledge from one task to another is essential for effective learning. In this paper, we study transfer learning in the context of nonparametric classification based …
A new procedure, called DDa-procedure, is developed to solve the problem of classifying d-dimensional objects into q >= 2 classes. The procedure is completely nonparametric; it uses q-dimensional depth plots and a very efficient algorithm for discrimination analysis in the depth space [0,1]^q. Specifically, the depth i…
Proposes a robust method for counterfactual classification.
We study the sample complexity of semi-supervised learning (SSL) and introduce new assumptions based on the mismatch between a mixture model learned from unlabeled data and the true mixture model induced by the (unknown) class conditional distributions. Under these assumptions, we establish an labeled samp…
The paper introduces a method to quantify uncertainty in neural networks without parametric assumptions.
The nearest neighbor classifier fails in high dimensions, leading to this study.
The Neyman-Pearson (NP) paradigm in binary classification seeks classifiers that achieve a minimal type II error while enforcing the prioritized type I error controlled under some user-specified level . This paradigm serves naturally in applications such as severe disease diagnosis and spam detection, where people h…
We consider the two-group classification problem and propose a kernel classifier based on the optimal scoring framework. Unlike previous approaches, we provide theoretical guarantees on the expected risk consistency of the method. We also allow for feature selection by imposing structured sparsity using weighted kernel…
DIVA clusters dynamic data without needing cluster count, outperforming baselines.
In the framework of supervised classification (discrimination) for functional data, it is shown that the optimal classification rule can be explicitly obtained for a class of Gaussian processes with "triangular" covariance functions. This explicit knowledge has two practical consequences. First, the consistency of the …
The paper extends calibration to sets of probabilistic classifiers, finding many ensembles are poorly calibrated.
Knowing when a classifier's prediction can be trusted is useful in many applications and critical for safely using AI. While the bulk of the effort in machine learning research has been towards improving classifier performance, understanding when a classifier's predictions should and should not be trusted has received …
Adversarially robust machine learning has received much recent attention. However, prior attacks and defenses for non-parametric classifiers have been developed in an ad-hoc or classifier-specific basis. In this work, we take a holistic look at adversarial examples for non-parametric classifiers, including nearest neig…
Detects changes in classifier scores to identify shifts in class priors.
We propose a high dimensional classification method that involves nonparametric feature augmentation. Knowing that marginal density ratios are the most powerful univariate classifiers, we use the ratio estimates to transform the original feature measurements. Subsequently, penalized logistic regression is invoked, taki…
Improved conformal prediction for better conditional coverage of classifier predictions.
Undersampling often outperforms other methods in nonparametric classification.
The study compares clustering risk in Hidden Markov and i.i.d. models, showing the Bayes classifier is nearly optimal.
Develops a new nonparametric trace regression model for high-dimensional data.
When applying the support vector machine (SVM) to high-dimensional classification problems, we often impose a sparse structure in the SVM to eliminate the influences of the irrelevant predictors. The lasso and other variable selection techniques have been successfully used in the SVM to perform automatic variable selec…
To restore the historical sea surface temperatures (SSTs) better, it is important to construct a good calibration model for the associated proxies. In this paper, we introduce a new model for alkenone () based on the heteroscedastic Gaussian process (GP) regression method. Our nonparametric app…
Paper tackles label noise in large datasets, purifying noisy data with a nonparametric framework.
We propose a unified game-theoretical framework to perform classification and conditional image generation given limited supervision. It is formulated as a three-player minimax game consisting of a generator, a classifier and a discriminator, and therefore is referred to as Triple Generative Adversarial Network (Triple…
Paper introduces a method to evaluate abstaining classifiers by considering missing predictions as counterfactuals.
Proposes a gradient-based variable selection method for binary classification in RKHS.
Bayesian nonparametrics adapt model complexity to diverse datasets.
NP-iMCMC algorithm for nonparametric models in universal PPLs.
A boosting method improves nonparametric density estimation without smoothing assumptions.
The study examines Fisher-Riemann geodesics for nonparametric probability densities.
Study on CNNs' learning rates and approximation capacities.
Let be a random variable consisting of an observed feature vector and an unobserved class label with unknown joint distribution. In addition, let be a training data set consisting of completely observed independent copies of . Usual classification…
Surveying nonparametric inference with shape constraints, past and future.
We propose a novel class of kernels to alleviate the high computational cost of large-scale nonparametric learning with kernel methods. The proposed kernel is defined based on a hierarchical partitioning of the underlying data domain, where the Nyström method (a globally low-rank approximation) is married with a locall…
Optimal nonparametric regression estimator adapts to unknown smoothness.