Many classification applications require accurate probability estimates in addition to good class separation but often classifiers are designed focusing only on the latter. Calibration is the process of improving probability estimates by post-processing but commonly used calibration algorithms work poorly on small data…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
GOAL algorithm reduces and rotates feature space for small data classification.
In this paper, we confront the problem of deep learning's big labeled data requirements, offer a rule based strategy for extreme augmentation of small data sets and apply that strategy with the image to image translation model by Isola et al. (2016) to automate cel style cartoon coloring with very limited training data…
VAE improves semi-supervised learning in small data.
Estimates policy performance in small-data settings without sacrificing data.
The Era of Big Data has forced researchers to explore new distributed solutions for building fuzzy classifiers, which often introduce approximation errors or make strong assumptions to reduce computational and memory requirements. As a result, Big Data classifiers might be expected to be inferior to those designed for …
Bayesian methods reduce variance in subspace identification for small data sets.
This paper explores object detection in the small data regime, where only a limited number of annotated bounding boxes are available due to data rarity and annotation expense. This is a common challenge today with machine learning being applied to many new tasks where obtaining training data is more challenging, e.g. i…
A new meta-meta classification method tackles few-shot learning tasks.
Gaussian Processes (GPs) are known to provide accurate predictions and uncertainty estimates even with small amounts of labeled data by capturing similarity between data points through their kernel function. However traditional GP kernels are not very effective at capturing similarity between high dimensional data poin…
In this paper, we prove that the Schrödinger map flows from with to compact Kähler manifolds with small initial data in critical Sobolev spaces are global. This is a companion work of our previous paper [23] where the energy critical case was solved. In the first part of this paper, for heat f…
New regularization method corrects over-shrinkage in small data regression.
Physics-enhanced NNs improve predictive accuracy in small data scenarios.
The paper explores how smaller data sets can lead to better model selection decisions.
In modern supervised learning, there are a large number of tasks, but many of them are associated with only a small amount of labeled data. These include data from medical image processing and robotic interaction. Even though each individual task cannot be meaningfully trained in isolation, one seeks to meta-learn acro…
Convolutional neural networks (CNNs) work well on large datasets. But labelled data is hard to collect, and in some applications larger amounts of data are not available. The problem then is how to use CNNs with small data -- as CNNs overfit quickly. We present an efficient Bayesian CNN, offering better robustness to o…
Overfitting and treatment of "small data" are among the most challenging problems in the machine learning (ML), when a relatively small data statistics size is not enough to provide a robust ML fit for a relatively large data feature dimension . Deploying a massively-parallel ML analysis of generic classificatio…
SmallML predicts customer churn for SMEs with small data, improving accuracy by 24.2 points.
Model learns to select relevant clinical variables for disease subtype prediction from small data.
The CLT fails for LLM evaluations with small data, leading to underestimation of uncertainty.
Method predicts spatial values with few data using GP framework.
New method for density estimation without approximating posterior distributions.
Study examines mean estimation in high dimensions with small data.
Scaling clustering algorithms to massive data sets is a challenging task. Recently, several successful approaches based on data summarization methods, such as coresets and sketches, were proposed. While these techniques provide provably good and small summaries, they are inherently problem dependent - the practitioner …
Study shows low-complexity models can perform as well as state-of-the-art on small datasets.
Bayesian model updating uses VAEs to approximate likelihood with small data.
In the majority of molecular optimization tasks, predictive machine learning (ML) models are limited due to the unavailability and cost of generating big experimental datasets on the specific task. To circumvent this limitation, ML models are trained on big theoretical datasets or experimental indicators of molecular s…
This paper tackles the challenge presented by small-data to the task of Bayesian inference. A novel methodology, based on manifold learning and manifold sampling, is proposed for solving this computational statistics problem under the following assumptions: 1) neither the prior model nor the likelihood function are Gau…
Paper introduces a method to infer causal direction from limited data.
This article provides the role of big idea statisticians in future of Big Data Science. We describe the `United Statistical Algorithms' framework for comprehensive unification of traditional and novel statistical methods for modeling Small Data and Big Data, especially mixed data (discrete, continuous).
Proposes a model to generate high-dimensional financial returns using latent factor structure.
Classifier calibration does not always go hand in hand with the classifier's ability to separate the classes. There are applications where good classifier calibration, i.e. the ability to produce accurate probability estimates, is more important than class separation. When the amount of data for training is limited, th…
The automated construction of coarse-grained models represents a pivotal component in computer simulation of physical systems and is a key enabler in various analysis and design tasks related to uncertainty quantification. Pertinent methods are severely inhibited by the high-dimension of the parametric input and the li…
Physics-informed neural networks (PINNs) encode physical conservation laws and prior physical knowledge into the neural networks, ensuring the correct physics is represented accurately while alleviating the need for supervised learning to a great degree. While effective for relatively short-term time integration, when …
The paper develops a physics-aware method for modeling multiscale dynamics with reduced data.
Convolutional Neural Processes improve data efficiency in neural processes.
In this paper, we introduce a new gait segmentation method based on accelerometer data and develop a new distance function between two time series, showing novel and effectiveness in simultaneously identifying user and adversary. Comparing with the normally used Neural Network methods, our approaches use geometric feat…
PSiLON Net uses weight normalization and 1-path-norm regularization for efficient learning and sparsity.
This work combines machine learning with physical models to solve inverse problems efficiently.
In this paper, we consider the problem of partitioning a small data sample drawn from a mixture of product distributions. We are interested in the case that individual features are of low average quality , and we want to use as few of them as possible to correctly partition the sample. We analyze a spectral tech…
We study large-scale classification problems in changing environments where a small part of the dataset is modified, and the effect of the data modification must be quickly incorporated into the classifier. When the entire dataset is large, even if the amount of the data modification is fairly small, the computational …
Minkowski space is shown to be globally stable as a solution to the Einstein--Vlasov system in the case when all particles have zero mass. The proof proceeds by showing that the matter must be supported in the "wave zone", and then proving a small data semi-global existence result for the characteristic initial value p…
A new neural network model simulates financial markets without assuming underlying dynamics.
In this article we prove a family of local (in time) weighted Strichartz estimates with derivative losses for the Klein-Gordon equation on asymptotically de Sitter spaces and provide a heuristic argument for the non-existence of a global dispersive estimate on these spaces. The weights in the estimates depend on the ma…
Recent research shows that the following two models are equivalent: (a) infinitely wide neural networks (NNs) trained under l2 loss by gradient descent with infinitesimally small learning rate (b) kernel regression with respect to so-called Neural Tangent Kernels (NTKs) (Jacot et al., 2018). An efficient algorithm to c…
ALBU improves LDA performance on small datasets.
In a 2004 paper, Lindblad demonstrated that the minimal surface equation on describing graphical time-like minimal surfaces embedded in enjoy small data global existence for compactly supported initial data, using Christodoulou's conformal method. Here we give a different, geometr…
We consider supervised dimension reduction problems, namely to identify a low dimensional projection of the predictors $\-x$ which can retain the statistical relationship between $\-x$ and the response variable . We follow the idea of the sliced inverse regression (SIR) and the sliced average variance estimation (SA…