VAE improves semi-supervised learning in small data.
problem Limited data in biological samples.
method Applied VAE to capture features in a lower-dimension.
result Improved performance in semi-supervised learning.
Estimates mass of static vacuum metrics with small Bartnik data.
problem Estimating mass of static vacuum metrics with small perturbations.
method Second-order mass estimation using Bartnik data.
result New upper bound on Bartnik mass to fifth order.
Large corporate credit models may be adapted for small business risk assessment.
problem Limited data and lack of credit analysts for small businesses.
method Adapting large corporate credit risk models for small businesses.
result Adapted models can predict small business credit risk effectively.
Augmentation improves machine learning model performance on small datasets.
problem Suboptimal generalization performance of machine learning models on small datasets.
method Data augmentation to increase sample size and diversity.
result Augmentation improves AUC by 15.55% on average for small datasets.
Unified Bayesian model for multi-modal, small sample size biomedical data classification.
problem Classifying high-dimensional, multi-modal biomedical data with small sample sizes.
method Combines multi-modal data views into a latent space, prunes irrelevant features, and uses dual kernels for small sample size scenarios.
result Outperforms state-of-the-art models and identifies features aligned with existing markers.
A new method detects small holes in noisy data.
problem Detecting small holes in high-density regions from noise.
method Robust Density-Aware Distance (RDAD) filtration, incorporating distance-to-measure concept.
result The RDAD filtration prolongs the persistences of small holes, making them distinguishable from noise.
Study evaluates synthetic data augmentation for small datasets, highlighting inconsistencies in traditional metrics.
problem Inconsistent validation of synthetic data generated for small sample sizes.
method Proposes a normalized Bottleneck distance metric to evaluate synthetic tabular data.
result Common metrics like propensity scoring and MMD fail for small datasets, showing instability and high variability.
The Era of Big Data has forced researchers to explore new distributed solutions for building fuzzy classifiers, which often introduce approximation errors or make strong assumptions to reduce computational and memory requirements. As a result, Big Data classifiers might be expected to be inferior to those designed for …
This article provides the role of big idea statisticians in future of Big Data Science. We describe the `United Statistical Algorithms' framework for comprehensive unification of traditional and novel statistical methods for modeling Small Data and Big Data, especially mixed data (discrete, continuous).
Transfer learning improves chaotic dynamics predictions with less data.
problem Efficiently predicting chaotic dynamics with limited data.
method Transfer learning for nonlinear dynamics, optimizing transfer rate and leveraging small-scale turbulence universality.
result Significantly more accurate inference of chaotic dynamics achieved.
Framework for few-shot relation classification with minimal training data.
problem Few-shot relation classification with limited training data.
method Meta-learning framework that combines instance and support knowledge.
result Framework outperforms state-of-the-art results and achieves competitive performance with large training data.
In this paper, we confront the problem of deep learning's big labeled data requirements, offer a rule based strategy for extreme augmentation of small data sets and apply that strategy with the image to image translation model by Isola et al. (2016) to automate cel style cartoon coloring with very limited training data…
S2GPT-PINNs solve PDEs with sparse, small models.
problem Efficiently solving parametric PDEs with minimal resources.
method Sparse and small architecture, mathematically rigorous greedy algorithm, knowledge distillation, down-sampling.
result Achieves high efficiency with significantly fewer parameters.
In this paper we develop a Bayesian optimization based hyperparameter tuning framework inspired by statistical learning theory for classifiers. We utilize two key facts from PAC learning theory; the generalization bound will be higher for a small subset of data compared to the whole, and the highest accuracy for a smal…
We study large-scale classification problems in changing environments where a small part of the dataset is modified, and the effect of the data modification must be quickly incorporated into the classifier. When the entire dataset is large, even if the amount of the data modification is fairly small, the computational …
In modern supervised learning, there are a large number of tasks, but many of them are associated with only a small amount of labeled data. These include data from medical image processing and robotic interaction. Even though each individual task cannot be meaningfully trained in isolation, one seeks to meta-learn acro…
Proposes Causal-Batle for estimating treatment effects in small high-dimensional datasets.
problem Estimating treatment effects with small high-dimensional datasets.
method Adopts transfer learning techniques for causal inference.
result Improves treatment effect estimates in small high-dimensional datasets.
Augment small datasets with synthetic backgrounds to train lightweight CNNs for human pose estimation.
problem Training CNNs from limited real-world data for human pose estimation.
method Synthetic background substitution for data augmentation.
result Improves generalization to unseen environments.
We study the topology of small covers from their fundamental groups. We find a way to obtain explicit presentations of the fundamental group of a small cover. Then we use these presentations to study the relations between the fundamental groups of a small cover and its facial submanifolds. In particular, we can determi…
Analyzes Willmore flow for graphs with boundary data, proving existence and convergence.
problem Willmore flow of graphs with boundary conditions over bounded domains.
method Developed low-regularity theory, reformulated graphical equation, used time-weighted parabolic Hölder spaces.
result Proved short-time and global existence for initial data in C1+α(Ω) and Lipschitz, with exponential convergence. Physics-enhanced NNs improve predictive accuracy in small data scenarios.
problem Predicting physical systems dynamics with limited data.
method Integrating physical principles (Hamiltonian/Lagrangian) and regularization term based on energy level into neural networks.
result Significant gains in predictive accuracy for small data cases.
In this paper, we proved the mass angular momentum inequality\cite{D1}\cite{ChrusLiWe}\cite{SZ} for axisymmetric, asymptotically flat, vacuum constraint data sets with small trace. Given an initial data set with small trace, we construct a boost evolution spacetime of the Einstein vacuum equations as \cite{ChOM}. Then …
SmallML predicts customer churn for SMEs with small data, improving accuracy by 24.2 points.
problem AI exclusion of SMEs due to data scale mismatch.
method Bayesian transfer learning with hierarchical pooling and conformal prediction.
result 96.7% AUC on 100 obs SMEs, 24.2 point improvement over logistic regression.
Differential privacy for simple linear regression protects small datasets from individual data leaks.
problem Protecting sensitive personal information in small datasets from individual data leaks.
method Differential privacy algorithms for simple linear regression tailored for small datasets (tens to hundreds of datapoints).
result Robust estimators like Theil-Sen perform well on small datasets, but standard algorithms improve as dataset size increases.
Generative Latent Implicit Conditional Optimization (GLICO) learns from small samples.
problem Learning from small labeled datasets.
method Generative Latent Implicit Conditional Optimization (GLICO) learns a latent space and generator from small labeled data.
result GLICO synthesizes new samples for every class using as few as 10 examples per class.
Learning with noisy labels is one of the hottest problems in weakly-supervised learning. Based on memorization effects of deep neural networks, training on small-loss instances becomes very promising for handling noisy labels. This fosters the state-of-the-art approach "Co-teaching" that cross-trains two deep neural ne…
We propose a penalized orthogonal-components regression (POCRE) for large p small n data. Orthogonal components are sequentially constructed to maximize, upon standardization, their correlation to the response residuals. A new penalization framework, implemented via empirical Bayes thresholding, is presented to effecti…
A method for clustering small datasets in high dimensions using random projections.
problem Challenges in clustering small datasets in high-dimensional spaces.
method Random projection followed by binary clustering in one-dimensional space.
result Statistically significant clustering structures can be found with as few as 100-200 points.
Gaussian Processes (GPs) are known to provide accurate predictions and uncertainty estimates even with small amounts of labeled data by capturing similarity between data points through their kernel function. However traditional GP kernels are not very effective at capturing similarity between high dimensional data poin…
Bayesian methods reduce variance in subspace identification for small data sets.
problem High variance in traditional subspace identification methods for large models or small sample sizes.
method Investigation of Bayesian estimation solutions (regularized and shrinkage estimators) for subspace identification.
result Bayesian estimators reduce estimation risk by up to 40% compared to traditional methods.
Many classification applications require accurate probability estimates in addition to good class separation but often classifiers are designed focusing only on the latter. Calibration is the process of improving probability estimates by post-processing but commonly used calibration algorithms work poorly on small data…
FedFaiREE addresses fairness in decentralized learning with small samples.
problem Ensuring fairness in decentralized federated learning with limited data.
method FedFaiREE is a post-processing algorithm for distribution-free fair learning in decentralized settings with small samples.
result FedFaiREE provides theoretical guarantees for both fairness and accuracy in decentralized environments.
Better calibration for small datasets improves probability estimates.
problem Limited data hinders traditional calibration methods.
method Generating additional calibration data improves classifier performance.
result The proposed approach enhances calibration for small datasets.
HAR regression improves performance on small datasets.
problem Small datasets with complex functions.
method Data-adaptive kernel ridge regression using tensor-product spline basis.
result Achieves n−1/3 convergence rate for right-continuous functions. Estimates policy performance in small-data settings without sacrificing data.
problem Poor performance of cross-validation in small-data optimization.
method Uses sensitivity analysis to estimate gradient of optimal objective value.
result Explicit high-probability bounds on error of estimator for small-data, large-scale problems.
Paper proposes a method to evaluate SME credit risk using meta paths.
problem Evaluate credit risk of small and medium-sized enterprises with limited data.
method Exploits the representative power of information networks and meta paths to infer SME financial status.
result Meta path feature effectively identifies SMEs with credit risks.
We focus on developing a novel scalable graph-based semi-supervised learning (SSL) method for a small number of labeled data and a large amount of unlabeled data. Due to the lack of labeled data and the availability of large-scale unlabeled data, existing SSL methods usually encounter either suboptimal performance beca…
DeFi TrustBoost uses blockchain and AI to assess small business loans.
problem Assessing small business loans from low-wealth households.
method Combines blockchain and Explainable AI to ensure confidentiality, compliance, and security.
result Tamper-proof auditing and on-chain/off-chain data storage for financial organizations.
The study tackles forgery in machine unlearning, showing that forging is limited and can be detected.
problem Adversarial crafting of data to mimic model behavior without removing information.
method Developed a framework to analyze ε-forging sets and proved their measure decay. result The forging set measure decays as ε(d−r)/2, providing evidence against false unlearning claims. GBEST model improves survival analysis for small datasets.
problem Challenges in survival analysis, especially with small data.
method Bayesian bootstrap and Beta Stacy bootstrap methods integrated into bagging tree models.
result GBEST model outperforms classical survival models in predictive performance and stability.
GOAL algorithm reduces and rotates feature space for small data classification.
problem Challenges in identifying important features for classification in small data settings.
method GOAL algorithm reduces and rotates feature space in a lower-dimensional gauge, providing an analytically tractable solution.
result GOAL algorithm outperforms state-of-the-art ML tools in synthetic and real-world applications.
Novel approach for SEM in small samples with p>n.
problem Small sample size and p>n issues in factor-based SEM. method Reformulates covariance structure into self-covariance and cross-covariance, defines a feasible set with relative error constraint.
result Improved stability and directional information in small-sample settings.
SFM resolves small-scale physics challenges in weather data.
problem Challenges in super-resolving small-scale details in physical sciences like weather.
method Encoding inputs to a latent base distribution, flow matching for stochastic details, adaptive noise scaling.
result SFM framework significantly outperforms existing methods.
Dirichlet process mixture (DPM) models tend to produce many small clusters regardless of whether they are needed to accurately characterize the data - this is particularly true for large data sets. However, interpretability, parsimony, data storage and communication costs all are hampered by having overly many clusters…
Geometric Brownian motion simulates stock prices for Brazilian small caps index.
problem Simulating stock prices for the Brazilian small caps index.
method Used geometric Brownian motion to simulate stock prices of Brazilian small caps index using historical data.
result Simulated prices better for portfolios with higher returns, lower risks, and higher Sharpe Indexes.
Paper explains why small-loss criterion works for learning from noisy labels.
problem Learning from noisy labels in deep learning with limited labeled data.
method Theoretical analysis and reformulation of the small-loss criterion.
result Theoretical explanation and reformulation of the small-loss criterion.
Data augmentation improves financial prediction models, especially for small datasets.
problem Improving financial prediction models on small, noisy, non-stationary datasets.
method Evaluation of data augmentation methods combined with deep learning models on financial datasets.
result Data augmentation significantly improves financial performance, up to 400% improvement in risk-adjusted return.
One significant challenge to scaling entity resolution algorithms to massive datasets is understanding how performance changes after moving beyond the realm of small, manually labeled reference datasets. Unlike traditional machine learning tasks, when an entity resolution algorithm performs well on small hold-out datas…