SmallML predicts customer churn for SMEs with small data, improving accuracy by 24.2 points.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
GOAL algorithm reduces and rotates feature space for small data classification.
Method predicts spatial values with few data using GP framework.
VAE improves semi-supervised learning in small data.
Meta-learning helps use small data from many tasks to compensate for lack of big data.
The Era of Big Data has forced researchers to explore new distributed solutions for building fuzzy classifiers, which often introduce approximation errors or make strong assumptions to reduce computational and memory requirements. As a result, Big Data classifiers might be expected to be inferior to those designed for …
Bayesian methods reduce variance in subspace identification for small data sets.
In this paper, we confront the problem of deep learning's big labeled data requirements, offer a rule based strategy for extreme augmentation of small data sets and apply that strategy with the image to image translation model by Isola et al. (2016) to automate cel style cartoon coloring with very limited training data…
Improves probability estimates for small datasets in multi-class problems.
New regularization method corrects over-shrinkage in small data regression.
Physics-enhanced NNs improve predictive accuracy in small data scenarios.
The paper explores how smaller data sets can lead to better model selection decisions.
The CLT fails for LLM evaluations with small data, leading to underestimation of uncertainty.
Better calibration for small datasets improves probability estimates.
This paper explores object detection in the small data regime, where only a limited number of annotated bounding boxes are available due to data rarity and annotation expense. This is a common challenge today with machine learning being applied to many new tasks where obtaining training data is more challenging, e.g. i…
New method for density estimation without approximating posterior distributions.
Estimates policy performance in small-data settings without sacrificing data.
Gaussian Processes (GPs) are known to provide accurate predictions and uncertainty estimates even with small amounts of labeled data by capturing similarity between data points through their kernel function. However traditional GP kernels are not very effective at capturing similarity between high dimensional data poin…
Study examines mean estimation in high dimensions with small data.
Soft-constrained PINN solves ODEs with minimal data, improving efficiency and robustness.
Bayesian model updating uses VAEs to approximate likelihood with small data.
Paper introduces a method to infer causal direction from limited data.
This article provides the role of big idea statisticians in future of Big Data Science. We describe the `United Statistical Algorithms' framework for comprehensive unification of traditional and novel statistical methods for modeling Small Data and Big Data, especially mixed data (discrete, continuous).
Convolutional neural networks (CNNs) work well on large datasets. But labelled data is hard to collect, and in some applications larger amounts of data are not available. The problem then is how to use CNNs with small data -- as CNNs overfit quickly. We present an efficient Bayesian CNN, offering better robustness to o…
In this paper, we prove that the Schrödinger map flows from with to compact Kähler manifolds with small initial data in critical Sobolev spaces are global. This is a companion work of our previous paper [23] where the energy critical case was solved. In the first part of this paper, for heat f…
Proposes a model to generate high-dimensional financial returns using latent factor structure.
A new meta-meta classification method tackles few-shot learning tasks.
A new framework for optimizing interventions with limited data.
Model learns to select relevant clinical variables for disease subtype prediction from small data.
Paper tackles robust prediction of nuclear reactor materials under scarce data.
In this paper, we introduce a new gait segmentation method based on accelerometer data and develop a new distance function between two time series, showing novel and effectiveness in simultaneously identifying user and adversary. Comparing with the normally used Neural Network methods, our approaches use geometric feat…
Study shows low-complexity models can perform as well as state-of-the-art on small datasets.
PSiLON Net uses weight normalization and 1-path-norm regularization for efficient learning and sparsity.
In the majority of molecular optimization tasks, predictive machine learning (ML) models are limited due to the unavailability and cost of generating big experimental datasets on the specific task. To circumvent this limitation, ML models are trained on big theoretical datasets or experimental indicators of molecular s…
This paper tackles the challenge presented by small-data to the task of Bayesian inference. A novel methodology, based on manifold learning and manifold sampling, is proposed for solving this computational statistics problem under the following assumptions: 1) neither the prior model nor the likelihood function are Gau…
Analyzes Willmore flow for graphs with boundary data, proving existence and convergence.
eSPA breaches overfitting barriers in ML with 10^-12 cost.
A new neural network model simulates financial markets without assuming underlying dynamics.
In this article we prove a family of local (in time) weighted Strichartz estimates with derivative losses for the Klein-Gordon equation on asymptotically de Sitter spaces and provide a heuristic argument for the non-existence of a global dispersive estimate on these spaces. The weights in the estimates depend on the ma…
Recent research shows that the following two models are equivalent: (a) infinitely wide neural networks (NNs) trained under l2 loss by gradient descent with infinitesimally small learning rate (b) kernel regression with respect to so-called Neural Tangent Kernels (NTKs) (Jacot et al., 2018). An efficient algorithm to c…
ALBU improves LDA performance on small datasets.
Increasing volume of Electronic Health Records (EHR) in recent years provides great opportunities for data scientists to collaborate on different aspects of healthcare research by applying advanced analytics to these EHR clinical data. A key requirement however is obtaining meaningful insights from high dimensional, sp…
Increasing volume of Electronic Health Records (EHR) in recent years provides great opportunities for data scientists to collaborate on different aspects of healthcare research by applying advanced analytics to these EHR clinical data. A key requirement however is obtaining meaningful insights from high dimensional, sp…
In a 2004 paper, Lindblad demonstrated that the minimal surface equation on describing graphical time-like minimal surfaces embedded in enjoy small data global existence for compactly supported initial data, using Christodoulou's conformal method. Here we give a different, geometr…
In the following short article we adapt a new and popular machine learning model for inference on medical data sets. Our method is based on the Variational AutoEncoder (VAE) framework that we adapt to survival analysis on small data sets with missing values. In our model, the true health status appears as a set of late…
The automated construction of coarse-grained models represents a pivotal component in computer simulation of physical systems and is a key enabler in various analysis and design tasks related to uncertainty quantification. Pertinent methods are severely inhibited by the high-dimension of the parametric input and the li…
In this work, we perform an exploratory study on synthesizing deep neural networks using biological synaptic strength distributions, and the potential influence of different distributions on modelling performance particularly for the scenario associated with small data sets. Surprisingly, a CNN with convolutional layer…
We propose a novel method to train deep convolutional neural networks which learn from multiple data sets of varying input sizes through weight sharing. This is an advantage in chemometrics where individual measurements represent exact chemical compounds and thus signals cannot be translated or resized without disturbi…