Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

20416181 · Jun 202019922001200920182026
48 results for Naive Bayes Nearest Neighbour

Proposes learning a metric for class-conditional KNN to improve performance.

problem Failure of traditional NN techniques when data representation does not capture perceptual similarity.
method Class Conditional Metric Learning (CCML) that optimizes a soft form of NBNN selection rule.
result CCML outperforms existing learned distance metrics across various datasets.

Dynamic classifier chains improve multi-label classification efficiency.

problem Building efficient multi-label classification models.
method Dynamic ensemble of chain classifiers using Naive Bayes and nearest neighbor approaches, with heuristic for label order optimization.
result The proposed dynamic chain model based on Naive Bayes classifier and heuristic is efficient for multi-label classification.

A review of nearest neighbour classifiers, focusing on similarity measures and computational efficiency.

problem Improving the performance of nearest neighbour classifiers.
method Explains mechanisms for assessing similarity (distance), identifying nearest neighbours, and reducing data dimensionality.
result Added new sections on time-series similarity measures, retrieval speed-up, and intrinsic dimensionality.

Study minimax rates for cost-sensitive classification on manifolds using approximate nearest neighbours.

problem Minimizing classification error on manifolds embedded in high-dimensional spaces.
method Approximate nearest neighbour algorithm in a randomly projected low-dimensional space.
result Minimax rates for cost-sensitive classification on manifolds are achieved by the approximate nearest neighbour algorithm.

A new adaptive kNN classifier outperforms Random Forests.

problem Improving classification accuracy using nearest neighbors.
method Finding discriminant subspaces for efficient nearest neighbor classification, leveraging bagging for diversity.
result The proposed method outperforms Random Forests and other nearest neighbors ensembles.

Enhances classifier performance through feature space transformations and model selection.

problem Improving the accuracy of classifiers by reducing complexity.
method Combining feature mapping, prototype selection, and kernel function transformations to transform data into a more convenient distribution.
result Our methods produce competitive classifiers and are statistically different among them.

New algorithm improves topological stability in non-linear dimensionality reduction.

problem Topological instability in choosing nearest neighbors in Isomap.
method Uses point and its two nearest neighbors to find subspace and orthogonal complement, then adds new points based on distance and angle.
result Improves topological stability and reduces short-circuit errors.

As a consequence of the strong and usually violated conditional independence assumption (CIA) of naive Bayes (NB) classifier, the performance of NB becomes less and less favorable compared to sophisticated classifiers when the sample size increases. We learn from this phenomenon that when the size of the training data …

2014-12-21abs ↗pdf ↗

This paper tackles spam detection on Twitter by analyzing correlated features.

problem Spam detection on social media, especially Twitter, to improve user experience.
method Extracted tweet-based and user-based features, identified correlated features, and used artificial neural networks for classification.
result Achieved 97.57% accuracy in classifying tweets as spam or non-spam.

New clustering algorithm uses reverse nearest neighbour for better density-based clustering.

problem Density-based clustering of separated high-density regions.
method Uses reverse nearest neighbour (RNN) queries to estimate densities and recover clusters.
result Outperforms DBSCAN and ISDBSCAN on synthetic and real-world data.

Optimizes one-class classification methods for better performance.

problem Improving one-class classification accuracy through hyperparameter optimization.
method Hyperparameter optimization for five one-class classification methods (SVM, NND, LNND, LOF, ALP) using various datasets.
result ALP and SVM perform best after hyperparameter optimization, with ALP being more efficient.

Unified framework evaluates different nearest neighbor classification methods.

problem Evaluating and comparing classical, fuzzy, and fuzzy rough nearest neighbor classification methods.
method Standardized nearest neighbor weighting with kernel functions applied to distance and/or rank values of nearest neighbors.
result NN, FNN, and FRNN perform best with Boscovich distance, and NN and FRNN perform best with specific combinations of weights and scaling measures.

Despite its simplicity, the naive Bayes classifier has surprised machine learning researchers by exhibiting good performance on a variety of learning problems. Encouraged by these results, researchers have looked to overcome naive Bayes primary weakness - attribute independence - and improve the performance of the algo…

2012-10-19abs ↗pdf ↗

The paper combines supervised and unsupervised learning to predict financial market movements.

problem Predicting profitable opportunities in financial markets using machine learning.
method The paper uses linear models and Gaussian Mixture Models (GMM) to extract features from Bitcoin, Pepecoin, and Nasdaq markets.
result GMM filtering improved the performance of KNN and RF algorithms, leading to higher average returns.

Study classifies liability insurance policies using machine learning.

problem Classifying liability insurance policies with or without claims.
method Used machine learning models like nearest neighbour and logistic regression on Actuarial Challenge dataset.
result Models accurately classified policies into claims and non-claims groups.

Develops a theoretical framework for scalable Gaussian Process regression methods.

problem Limited scalability of Gaussian Process regression for large datasets.
method Introduces and analyzes Nearest Neighbour Gaussian Process (NNGP) and scalable GPnn methods.
result Derives almost sure pointwise limits for predictive criteria and proves risk minimax rates.

A nearly tight convex relaxation for sparse Naive Bayes features.

problem Feature selection in large-scale Naive Bayes classification.
method Proposes a convex relaxation for the combinatorial maximum-likelihood problem of feature selection in Naive Bayes.
result The convex relaxation bounds become tight as marginal feature contributions decrease, providing a nearly optimal solution.

New algorithm improves k-NN search efficiency and adaptability.

problem Inefficiency of existing k-NN search methods due to space partitioning.
method Dynamic continuous indexing, randomized algorithm, fine-grained control, adaptability to data density, dynamic updates.
result Outperforms LSH in approximation quality, speed, and space efficiency.

The paper compares one-hot encoding to Naïve Bayes for categorical variables.

problem Incorrect one-hot encoding affects Naïve Bayes performance.
method Mathematical and experimental analysis of PoB vs. categorical Naïve Bayes.
result Posterior probabilities are usually greater in the PoB case, but agree on the maximum a posteriori class label.

Proposes a non-convex optimization method for a parsimonious weighted naive Bayes classifier.

problem Improving naïve Bayes classifier performance with a large number of input variables.
method Sparse regularization of model log-likelihood for direct estimation of variable weights.
result Optimization-based weighted naïve Bayes classifiers achieve equivalent performance to averaging-based classifiers.

Improved k-NN search via Prioritized DCI reduces dependence on intrinsic dimensionality.

problem Exponential query time increase in k-NN search due to dimensionality.
method Proposes Prioritized DCI, a variant of DCI, to reduce query time dependence on intrinsic dimensionality.
result Significant improvement in query time, reducing dependence on intrinsic dimensionality.

DINOSAUR improves retrieval by accounting for embedding uncertainty in recommender systems.

problem Retrieval bias towards popular items due to noisy embeddings.
method Samples multiple embeddings per item and queries with sampled embeddings to account for uncertainty.
result Improves coverage of long-tail niche content without sacrificing recall.

A new ensemble method using random projections for kNN classification.

problem Improving kNN classification accuracy through ensemble methods.
method Random projection of bootstrap samples into lower dimensions, using extended neighbourhood rule for base learners.
result Enhanced classification accuracy compared to traditional kNN and other ensembles.

Naive Bayes model performs best in classifying seismological articles about precursory seismicity.

problem Classifying seismological articles about precursory seismicity using machine learning.
method Various supervised machine learning classifiers (Naive Bayes, k-Nearest Neighbors, Support Vector Machines, Random Forests) were tested on a seismological corpus of 100 articles.
result Naive Bayes model performs best with cross-validation accuracies of 86% for binary classification and up to 78% for multiclass classification.

Naive Bayes can be used as a discriminative classifier, matching the definition of logistic regression.

problem The definition of generative and discriminative classifiers.
method Comparing Naive Bayes and logistic regression, showing they can be used in either generative or discriminative ways.
result Naive Bayes can be used as a discriminative classifier.

Draft proposes adapting neural networks to match naive Bayes classifiers.

problem Bridge between neural networks and naive Bayes classifiers.
method Class-conditional compression and disentanglement using variational bounds.
result Latent representations enable naive Bayes classifier performance.

Smart Bayes integrates generative and discriminative features for improved classification.

problem Improving classification performance by combining generative and discriminative modeling.
method Integrates generative likelihood-ratio features into a logistic-regression-style classifier.
result Often outperforms logistic regression and Naive Bayes in simulations and real data.

Proposes a new k-NN algorithm to improve classification accuracy by removing noise and pseudo-neighbours.

problem Noise and pseudo-neighbours in large-scale databases affect k-NN performance.
method Introduces a weighted mutual k-Nearest Neighbour algorithm to detect and remove noise, and minimize distant neighbours' influence.
result The proposed algorithm provides comparative better results compared to standard k-NN.