Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

1122 · Dec 201319922001200920182026
48 results for Low-Sample-Size

Paper addresses SVM bias in high-dimension, low-sample-size settings.

problem Bias in SVM performance in high-dimension, low-sample-size settings.
method Proposes a bias-corrected SVM (BC-SVM) to improve SVM performance.
result BC-SVM gives preferable performances in high-dimension, low-sample-size settings.

High-dimensional data models, often with low sample size, abound in many interdisciplinary studies, genomics and large biological systems being most noteworthy. The conventional assumption of multinormality or linearity of regression may not be plausible for such models which are likely to be statistically complex due …

2008-05-21abs ↗pdf ↗

Proposes a novel classification criterion for high-dimensional data with few samples.

problem Challenges in classifying high-dimensional data with limited samples.
method Tolerance similarity criterion and No-separated Data Maximum Dispersion classifier (NPDMD).
result NPDMD outperforms state-of-the-art methods in various real-world applications.

Improved average distance classifier for HDLSS settings with multiple population differences.

problem Poor performance of average distance classifier in HDLSS settings with location and scale differences.
method Proposed transformations to the average distance classifier to handle multiple population differences.
result The proposed classifiers perform well even when populations differ in other aspects than location and scale.

New method selects better graphs for GGM inference in small sample sizes.

problem Inference of conditional correlations in high-dimensional data with limited samples.
method Composite procedure combining nodewise edge selection and penalised likelihood maximisation.
result Our method produces graphs closer to the true distribution with better KL divergence.

Locally sparse neural networks improve interpretability for biomedical tabular data.

problem Overfitting and lack of interpretability in neural networks for tabular biomedical data.
method Locally sparse neural network with a gating network to select relevant features.
result The method outperforms state-of-the-art models in synthetic and real-world biomedical datasets.

DeepFS uses deep neural networks to select significant features in ultra high-dimensional data.

problem Challenges in traditional feature selection methods for high-dimensional, low-sample-size data.
method Two-step nonparametric approach combining deep neural networks and feature screening.
result DeepFS effectively identifies significant features with high precision for ultra high-dimensional data.

Proposes a VAE for HDLSS data augmentation.

problem Data augmentation in HDLSS settings with small sample sizes.
method Geometry-based variational autoencoder with latent space modeling.
result Significant improvement in classification metrics (e.g., balanced accuracy from 66.3% to 74.3%).

PGLMC tackles HDLSS problems with improved linear classifier.

problem Challenges in high-dimensional low-sample-size data sets.
method Population-guided large margin classifier (PGLMC) with comprehensive consideration of local structural information and training samples.
result PGLMC outperforms state-of-the-art methods in most cases.

PSC classifier improves HDLSS classification on class-imbalanced data.

problem Classification on high-dimension low-sample-size data with class imbalance.
method Population Structure-learned Classifier (PSC) maximizing inter-class and intra-class scatter matrices.
result PSC outperforms state-of-the-art methods on IHDLSS.

DKN adapts to medical imaging data with limited samples and interpretable models.

problem Medical imaging data's unique nature makes general methods like CNN unsuitable.
method DKN uses a Kronecker product structure to adapt to low sample size and provide interpretable models.
result DKN achieves prediction power comparable to CNN and provides model interpretability.

Machine learning outperforms statistical methods with larger data sets.

problem Lower predictive performance of machine learning methods compared to statistical methods under low sample size.
method Learning curve method to analyze predictive performance across different sample sizes.
result Machine learning methods improve their predictive performance as sample size increases.

GOTabPFN improves tabular model performance with compact tokenization for HDLSS data.

problem Making tabular models effective for high-dimensional, low-sample size data without retraining.
method Introducing Graph-guided Ordering with Local Refinement (GO-LR) and Neuro-Inspired Subunit Compression (NSC) to create compact meta-features.
result GOTabPFN improves stability and accuracy in tabular benchmarks with compact tokenization.

Proposes a new method to improve regression models with reweighted samples.

problem Improves regression models' performance under low sample sizes and covariate perturbations.
method Reparametrizes sample weights using a doubly non-negative matrix and solves the reweighted estimate efficiently.
result Adversarial reweighting strategy delivers promising results on various datasets.

In clinical and neuroscientific studies, systematic differences between two populations of brain networks are investigated in order to characterize mental diseases or processes. Those networks are usually represented as graphs built from neuroimaging data and studied by means of graph analysis methods. The typical mach…

2015-11-19abs ↗pdf ↗

Real world systems typically feature a variety of different dependency types and topologies that complicate model selection for probabilistic graphical models. We introduce the ensemble-of-forests model, a generalization of the ensemble-of-trees model. Our model enables structure learning of Markov random fields (MRF) …

2013-12-17abs ↗pdf ↗

New estimator robust to adversarial noise and data heterogeneity.

problem Sensitive to adversarial noise and poor performance with heterogeneous data.
method Distributionally robust estimator minimizing worst-case conditional expected loss over adversarial distributions.
result Efficiently finds non-parametric local estimates via convex optimization.

The paper optimizes autoencoder latent spaces for one-class learning with controlled connectivity.

problem Learning representations with controllable connectivity for better upstream tasks.
method A novel loss function based on persistent homology controls the connectivity of autoencoder latent spaces.
result The controlled connectivity in latent space improves one-class learning performance, especially in low sample size scenarios.

Probabilistic graphical models are graphical representations of probability distributions. Graphical models have applications in many fields including biology, social sciences, linguistic, neuroscience. In this paper, we propose directed acyclic graphs (DAGs) learning via bootstrap aggregating. The proposed procedure i…

2014-06-09abs ↗pdf ↗

Proposes a new method for selecting regularization parameters in sparse precision matrix estimation.

problem Selecting an appropriate regularization parameter for sparse precision matrix estimation.
method Developed a closed-form matrix-valued regularization parameter based on the sampling distribution of optimality conditions.
result The proposed method achieves comparable estimation accuracy and superior support recovery to cross-validation, with significant runtime improvements.

Paper improves Bayesian network learning from related data sets.

problem Learning from heterogeneous data sets with different probabilistic structures.
method Mixed-effects models to pool information across related data sets.
result Mixed-effects models outperform traditional methods in accuracy.

PbP strategy improves logistic model prediction with missing values.

problem Predicting with missing inputs in logistic models.
method Pattern-by-Pattern (PbP) strategy for logistic models with missing values.
result PbP accurately approximates Bayes probabilities under GPMM across various missing data scenarios.

Efficiently solves large-scale SVMs with sparse semismooth Newton method.

problem Numerical difficulties in solving large-scale SVMs.
method Sparse semismooth Newton based augmented Lagrangian method.
result Outperforms state-of-the-art solvers for large-scale SVMs.

This paper analyzes the generalization risk of unrolled neural networks using Stein's Unbiased Risk Estimator.

problem Analyzing the generalization risk of unrolled neural networks and its relationship to network design and train sample size.
method Using Stein's Unbiased Risk Estimator (SURE), the paper analyzes the generalization risk with bias and variance components for recurrent unrolled networks, focusing on the degrees-of-freedom (DOF) component and the trace of the end-to-end network Jacobian.
result DOF is well-approximated by the weighted path sparsity of the network under incoherence conditions on the trained weights, and DOF increases with train sample size and converges to the generalization risk for both recurrent and non-recurrent schemes.

Guided adaptive shrinkage uses co-data to improve feature selection in genomic studies.

problem Feature selection challenges in high-dimensional genomics data, especially in clinical settings.
method Guided adaptive shrinkage methods that use co-data to adapt shrinkage parameters.
result Improves feature selection in genomic studies, demonstrated through comparisons and examples.

Feature selection from wide datasets leads to misleading results.

problem Feature selection in wide datasets with few samples can lead to misleading results.
method Derived sample size requirement for declaring features different, used real datasets to illustrate issues.
result Feature selection from very wide datasets may lead to misleading results.

New refit strategy improves probability estimation for multicategory angle-based classifiers.

problem Improving probability estimation for multicategory angle-based classifiers in high dimensional applications.
method Proposes a new refit strategy for multicategory angle-based classifiers, adding small computation cost.
result Significant improvement in probability estimation with minimal additional computation.