VAE improves semi-supervised learning in small data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Meta-learning helps use small data from many tasks to compensate for lack of big data.
The Era of Big Data has forced researchers to explore new distributed solutions for building fuzzy classifiers, which often introduce approximation errors or make strong assumptions to reduce computational and memory requirements. As a result, Big Data classifiers might be expected to be inferior to those designed for …
Bayesian methods reduce variance in subspace identification for small data sets.
In this paper, we confront the problem of deep learning's big labeled data requirements, offer a rule based strategy for extreme augmentation of small data sets and apply that strategy with the image to image translation model by Isola et al. (2016) to automate cel style cartoon coloring with very limited training data…
GOAL algorithm reduces and rotates feature space for small data classification.
Improves probability estimates for small datasets in multi-class problems.
New regularization method corrects over-shrinkage in small data regression.
Physics-enhanced NNs improve predictive accuracy in small data scenarios.
The paper explores how smaller data sets can lead to better model selection decisions.
SmallML predicts customer churn for SMEs with small data, improving accuracy by 24.2 points.
The CLT fails for LLM evaluations with small data, leading to underestimation of uncertainty.
Method predicts spatial values with few data using GP framework.
Better calibration for small datasets improves probability estimates.
This paper explores object detection in the small data regime, where only a limited number of annotated bounding boxes are available due to data rarity and annotation expense. This is a common challenge today with machine learning being applied to many new tasks where obtaining training data is more challenging, e.g. i…
New method for density estimation without approximating posterior distributions.
Estimates policy performance in small-data settings without sacrificing data.
Gaussian Processes (GPs) are known to provide accurate predictions and uncertainty estimates even with small amounts of labeled data by capturing similarity between data points through their kernel function. However traditional GP kernels are not very effective at capturing similarity between high dimensional data poin…
Study examines mean estimation in high dimensions with small data.
Bayesian model updating uses VAEs to approximate likelihood with small data.
Paper introduces a method to infer causal direction from limited data.
This article provides the role of big idea statisticians in future of Big Data Science. We describe the `United Statistical Algorithms' framework for comprehensive unification of traditional and novel statistical methods for modeling Small Data and Big Data, especially mixed data (discrete, continuous).
Convolutional neural networks (CNNs) work well on large datasets. But labelled data is hard to collect, and in some applications larger amounts of data are not available. The problem then is how to use CNNs with small data -- as CNNs overfit quickly. We present an efficient Bayesian CNN, offering better robustness to o…
In this paper, we prove that the Schrödinger map flows from with to compact Kähler manifolds with small initial data in critical Sobolev spaces are global. This is a companion work of our previous paper [23] where the energy critical case was solved. In the first part of this paper, for heat f…
Proposes a model to generate high-dimensional financial returns using latent factor structure.
A new meta-meta classification method tackles few-shot learning tasks.
Model learns to select relevant clinical variables for disease subtype prediction from small data.
In this paper, we introduce a new gait segmentation method based on accelerometer data and develop a new distance function between two time series, showing novel and effectiveness in simultaneously identifying user and adversary. Comparing with the normally used Neural Network methods, our approaches use geometric feat…
Study shows low-complexity models can perform as well as state-of-the-art on small datasets.
PSiLON Net uses weight normalization and 1-path-norm regularization for efficient learning and sparsity.
In the majority of molecular optimization tasks, predictive machine learning (ML) models are limited due to the unavailability and cost of generating big experimental datasets on the specific task. To circumvent this limitation, ML models are trained on big theoretical datasets or experimental indicators of molecular s…
This paper tackles the challenge presented by small-data to the task of Bayesian inference. A novel methodology, based on manifold learning and manifold sampling, is proposed for solving this computational statistics problem under the following assumptions: 1) neither the prior model nor the likelihood function are Gau…
eSPA breaches overfitting barriers in ML with 10^-12 cost.
A new neural network model simulates financial markets without assuming underlying dynamics.
In this article we prove a family of local (in time) weighted Strichartz estimates with derivative losses for the Klein-Gordon equation on asymptotically de Sitter spaces and provide a heuristic argument for the non-existence of a global dispersive estimate on these spaces. The weights in the estimates depend on the ma…
Recent research shows that the following two models are equivalent: (a) infinitely wide neural networks (NNs) trained under l2 loss by gradient descent with infinitesimally small learning rate (b) kernel regression with respect to so-called Neural Tangent Kernels (NTKs) (Jacot et al., 2018). An efficient algorithm to c…
ALBU improves LDA performance on small datasets.
In a 2004 paper, Lindblad demonstrated that the minimal surface equation on describing graphical time-like minimal surfaces embedded in enjoy small data global existence for compactly supported initial data, using Christodoulou's conformal method. Here we give a different, geometr…
In the following short article we adapt a new and popular machine learning model for inference on medical data sets. Our method is based on the Variational AutoEncoder (VAE) framework that we adapt to survival analysis on small data sets with missing values. In our model, the true health status appears as a set of late…
The automated construction of coarse-grained models represents a pivotal component in computer simulation of physical systems and is a key enabler in various analysis and design tasks related to uncertainty quantification. Pertinent methods are severely inhibited by the high-dimension of the parametric input and the li…
In this work, we perform an exploratory study on synthesizing deep neural networks using biological synaptic strength distributions, and the potential influence of different distributions on modelling performance particularly for the scenario associated with small data sets. Surprisingly, a CNN with convolutional layer…
We propose a novel method to train deep convolutional neural networks which learn from multiple data sets of varying input sizes through weight sharing. This is an advantage in chemometrics where individual measurements represent exact chemical compounds and thus signals cannot be translated or resized without disturbi…
The paper develops a physics-aware method for modeling multiscale dynamics with reduced data.
Scaling clustering algorithms to massive data sets is a challenging task. Recently, several successful approaches based on data summarization methods, such as coresets and sketches, were proposed. While these techniques provide provably good and small summaries, they are inherently problem dependent - the practitioner …
New method creates coresets for deep neural networks efficiently.
We derive a statistical model for estimation of a dendrogram from single linkage hierarchical clustering (SLHC) that takes account of uncertainty through noise or corruption in the measurements of separation of data. Our focus is on just the estimation of the hierarchy of partitions afforded by the dendrogram, rather t…
In this paper we outline initial concepts for an immune inspired algorithm to evaluate price time series data. The proposed solution evolves a short term pool of trackers dynamically through a process of proliferation and mutation, with each member attempting to map to trends in price movements. Successful trackers fee…
We analyze differences between two information-theoretically motivated approaches to statistical inference and model selection: the Minimum Description Length (MDL) principle, and the Minimum Message Length (MML) principle. Based on this analysis, we present two revised versions of MML: a pointwise estimator which give…