SPREV simplifies visualization of complex labeled datasets.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Learnable multiclass hypothesis classes don't always have a sample compression scheme of fixed size.
Adversarial training can hurt robust accuracy in small sample size scenarios.
The paper revisits and improves on a Bayesian relevance vector machine method for small sample sizes.
A new algorithm for robust causal discovery in small sample sizes.
The paper shows that relaxing assumptions about causal graphs can lead to exponentially large equivalence classes.
Image classifiers are sensitive to small changes, affecting most images in a class.
Exact learning improves naive Bayes classifier performance for small samples.
Mini-batch stochastic gradient descent (SGD) and variants thereof approximate the objective function's gradient with a small number of training examples, aka the batch size. Small batch sizes require little computation for each model update but can yield high-variance gradient estimates, which poses some challenges for…
Fewer obstructions for small graphs in knotless embedding.
Proving a conjecture of Dennis Johnson, we show that the Torelli subgroup of the mapping class group has a finite generating set whose size grows cubically with respect to the genus of the surface. Our main tool is a new space called the handle graph on which the Torelli group acts cocompactly.
Privacy affects how much data is needed for CVaR optimization.
Unified Bayesian model for multi-modal, small sample size biomedical data classification.
Classification and clustering are both important topics in statistical learning. A natural question herein is whether predefined classes are really different from one another, or whether clusters are really there. Specifically, we may be interested in knowing whether the two classes defined by some class labels (when t…
Kernel adaptive filters (KAF) are a class of powerful nonlinear filters developed in Reproducing Kernel Hilbert Space (RKHS). The Gaussian kernel is usually the default kernel in KAF algorithms, but selecting the proper kernel size (bandwidth) is still an open important issue especially for learning with small sample s…
The paper investigates why GNNs struggle to generalize from small to large graphs.
SLdisco uses supervised learning to discover causal models from observational data.
The ability to learn from a small number of examples has been a difficult problem in machine learning since its inception. While methods have succeeded with large amounts of training data, research has been underway in how to accomplish similar performance with fewer examples, known as one-shot or more generally few-sh…
The two-sample hypothesis testing problem is studied for the challenging scenario of high dimensional data sets with small sample sizes. We show that the two-sample hypothesis testing problem can be posed as a one-class set classification problem. In the set classification problem the goal is to classify a set of data …
Study identifies key metrics for small and large tick assets in LOBs.
Scalable methods integrate multiview data for clinical outcomes.
Recent advances in graph convolutional networks have significantly improved the performance of chemical predictions, raising a new research question: "how do we explain the predictions of graph convolutional networks?" A possible approach to answer this question is to visualize evidence substructures responsible for th…
Bob predicts a future observation based on a sample of size one. Alice can draw a sample of any size before issuing her prediction. How much better can she do than Bob? Perhaps surprisingly, under a large class of loss functions, which we refer to as the Cover-Hart family, the best Alice can do is to halve Bob's risk. …
In biospectroscopy, suitably annotated and statistically independent samples (e. g. patients, batches, etc.) for classifier training and testing are scarce and costly. Learning curves show the model performance as function of the training sample size and can help to determine the sample size needed to train good classi…
Class imbalance is an intrinsic characteristic of multi-label data. Most of the labels in multi-label data sets are associated with a small number of training examples, much smaller compared to the size of the data set. Class imbalance poses a key challenge that plagues most multi-label learning methods. Ensemble of Cl…
We employ techniques of machine-learning, exemplified by support vector machines and neural classifiers, to initiate the study of whether AI can "learn" algebraic structures. Using finite groups and finite rings as a concrete playground, we find that questions such as identification of simple groups by "looking" at the…
Naive Bayes estimator is widely used in text classification problems. However, it doesn't perform well with small-size training dataset. We propose a new method based on Naive Bayes estimator to solve this problem. A correlation factor is introduced to incorporate the correlation among different classes. Experimental r…
Proposes a new signal model for high-dimensional, small-sample-size data.
Two methods reduce BN and DNN complexity, balancing size and accuracy.
New conformal prediction methods for long-tailed classification problems.
Regularized EM algorithm improves clustering performance with small sample sizes.
Simple private estimators for mean and covariance outperform existing methods.
New SDP algorithm recovers large clusters in SBM with small clusters of any size.
Gradient descent fails to learn simple neural networks efficiently.
Deep model tackles claim size modeling with quantile-based regression.
Policy gradient methods achieve linear convergence in simple MDPs.
Feature selection from wide datasets leads to misleading results.
Study evaluates synthetic data augmentation for small datasets, highlighting inconsistencies in traditional metrics.
A new test optimizes detecting small communities in large networks.
We find the wealth distribution for an economic agent in the financial market, in analogy with standard derivation of generaliz Boltzman (Tsallis) factor in statistical mechanics. In this respect, Tsallis entropic index separates two different regimes, the large and small size market. The Pareto like wealth distributio…
SFCNeXt estimates brain age from small MRI datasets.
Study compares under-bagging with other methods for imbalanced data.
We propose a novel method to train deep convolutional neural networks which learn from multiple data sets of varying input sizes through weight sharing. This is an advantage in chemometrics where individual measurements represent exact chemical compounds and thus signals cannot be translated or resized without disturbi…
We demonstrate that the lowest possible price change (tick-size) has a large impact on the structure of financial return distributions. It induces a microstructure as well as it can alter the tail behavior. On small return intervals, the tick-size can distort the calculation of correlations. This especially occurs on s…
Decision trees have been a very popular class of predictive models for decades due to their interpretability and good performance on categorical features. However, they are not always robust and tend to overfit the data. Additionally, if allowed to grow large, they lose interpretability. In this paper, we present a mix…
Study examines mean estimation in high dimensions with small data.
Modern deep neural network training is typically based on mini-batch stochastic gradient optimization. While the use of large mini-batches increases the available computational parallelism, small batch training has been shown to provide improved generalization performance and allows a significantly smaller memory footp…
The classification of multi-class microarray datasets is a hard task because of the small samples size in each class and the heavy overlaps among classes. To effectively solve these problems, we propose novel Error Correcting Output Code (ECOC) algorithm by Enhance Class Separability related Data Complexity measures du…