Adaptive sampling detects local concept drift with limited labels.
problem Detecting local concept drift in dynamic environments with scarce labels.
method Combines residual-based exploration and exploitation with EWMA monitoring.
result Superior performance in label efficiency and drift detection accuracy.
DiffsFormer uses AI-generated samples to improve stock forecasting accuracy.
problem Data scarcity in stock forecasting, including low signal-to-noise ratio and data homogeneity.
method DiffsFormer employs a Diffusion Model with a Transformer architecture to generate augmented stock factors.
result DiffsFormer achieves significant improvements in stock forecasting accuracy (7.2% and 27.8% relative improvements for CSI300 and CSI800 datasets, respectively).
Enhances PHM solutions by augmenting scarce multivariate time series data.
problem Data scarcity in failure prognostics.
method Extends autoregressive models for data augmentation.
result Significantly improves PHM solution performance.
Personalizes pre-trained models for nonparametric regression with limited data.
problem Improving data efficiency in nonparametric regression with few samples.
method Develops a theoretical framework and algorithms for few-shot personalization of black-box models.
result Achieves minimax optimal rate for personalization in nonparametric regression.
GCNs help in diagnosing label scarcity and feature quality on graphs.
problem Understanding when GCNs improve node classification.
method Simulated label scarcity, feature ablation, and per-class analysis.
result GCNs provide largest gains under extreme label scarcity, matching original performance with noisy features, but hurt when homophily is low and features are strong.
Local Gaussian correlation struggles in tails but a new method improves it.
problem Local Gaussian correlation's limitations in tail dependence.
method A new adaptive bandwidth method for LGC, optimizing for local effective sample size.
result Adaptive bandwidths outperform global ones in moderate dependence, but not in strong or weak dependence.
Paper tackles RUL prediction with scarce data using indirect supervision.
problem Predicting RUL with indirect supervision and scarce time series data.
method Unified framework called parameterized static regression, handling data scarcity without interpolation.
result Competitive performance in prediction accuracy with simulated data scarcity.
BLADE uses Bayesian methods to discover complex systems from scarce data.
problem Efficiently discovering governing equations of complex dynamical systems from limited data.
method Combines replica-exchange stochastic gradient Langevin Monte Carlo with active learning.
result Reduces measurement requirements by 60% for Lotka-Volterra and 40% for Burgers' equation.
Proposes a new VAE framework for anomaly detection in time series data.
problem Data scarcity leads to latent holes and discontinuous regions in latent space, causing non-robust reconstructions.
method Combines VAEs with self-supervised learning to address data scarcity and improve anomaly detection.
result Improves robustness of anomaly detection in time series data by addressing latent holes and discontinuities.
Synthetic augmentation improves financial machine learning performance in variance-dominant regimes.
problem Data scarcity in financial machine learning.
method Formalized synthetic augmentation, introduced size-matched null augmentation, and developed a non-parametric block permutation test.
result Synthetic augmentation is beneficial only in variance-dominant regimes, such as persistent volatility forecasting.
Paper explores VRM for PSMLC with partially labeled medical images.
problem Improving PSMLC with limited labeled data.
method Applies VRM to PSMLC for better model performance.
result VRM improves PSMLC performance with partial labels.
PINNs can learn trivial solutions; new approach improves performance.
problem PINNs fail to learn non-trivial solutions in data-scarce settings.
method Developed an alternative sampling approach and new penalty term.
result PINNs can be improved to learn non-trivial solutions with up to 80% fewer collocation points.
MarketGAN generates financial returns using GANs to match empirical stylized facts.
problem Generating financial returns under data scarcity and preserving stylized facts.
method Generative adversarial learning with a TCN backbone.
result MarketGAN outperforms conventional methods in portfolio applications.
The goal of this paper is to deal with a data scarcity scenario where deep learning techniques use to fail. We compare the use of two well established techniques, Restricted Boltzmann Machines and Variational Auto-encoders, as generative models in order to increase the training set in a classification framework. Essent…
FedACS uses attention to select clients with similar data for federated learning.
problem Non-IID data and data scarcity in federated learning.
method FedACS integrates an attention mechanism to prioritize clients with similar data distributions.
result FedACS improves federated learning performance by addressing non-IID data and data scarcity.
Generative algorithms learn high-dimensional data efficiently and generate new samples.
problem Learning from scarce high-dimensional data.
method Lipschitz-regularized gradient flows and particle-based algorithms.
result Correctly transports gene expression data points with high dimensionality.
RDLI integrates domain logic and context grounding to detect crypto anomalies under scarce labels.
problem Extreme label scarcity and evasion strategies in crypto networks.
method Relational Domain Logic Integration (RDLI) with Retrieval Grounded Context (RGC).
result RDLI outperforms GNN baselines by 28.9% in F1 score under 0.01% label scarcity.
Paper develops flood forecasting system for data scarce regions.
problem Flood forecasting in developing countries with data scarcity.
method Operational system for flood extent forecasts in India.
result Scalable and cost-efficient flood forecasting system.
Synthetic datasets help study bias in ML, overcoming data scarcity.
problem Lack of relevant datasets for bias research in ML.
method Presented a family of synthetic datasets with adjustable bias levels.
result Demonstrated an experiment using synthetic data to study bias.
The committor function is a central object of study in understanding transitions between metastable states in complex systems. However, computing the committor function for realistic systems at low temperatures is a challenging task, due to the curse of dimensionality and the scarcity of transition data. In this paper,…
Model predicts political ideology using context vectors to mitigate bias and scarcity.
problem Scarcity and selection bias in political ideology prediction.
method Proposes a statistical model decomposing embeddings into context and position vectors, training an end-to-end model for deployment.
result Model can predict ideological labels even with minimal biased data, outperforming state-of-the-art methods.
DSI improves tail-risk estimation in generative models by averaging checkpoints.
problem Generative models' instability in rare adverse scenarios.
method Diachronic Sample Integration (DSI) ensembles generated samples across checkpoints.
result DSI reduces tail-estimation error compared to single-checkpoint baselines.
Pessimistic Q-learning improves sample efficiency in offline reinforcement learning.
problem Insufficient coverage and sample scarcity in offline reinforcement learning datasets.
method Pessimistic Q-learning algorithm for offline reinforcement learning, focusing on variance reduction.
result Near-optimal sample complexity achieved with the proposed algorithm.
Fused Encoder Networks improve momentum strategies on crypto data.
problem Deploying momentum strategies on crypto data with limited samples leads to over-fitted models.
method Hybrid transfer learning model combining source and target datasets.
result Fused Encoder Networks outperform classical momentum strategies and benchmarks.
CDSSL improves representation quality by integrating linear and nonlinear dependencies.
problem Scarcity of labeled data and neglect of nonlinear dependencies in SSL.
method CDSSL combines linear correlations and nonlinear dependencies using HSIC in RKHS.
result CDSSL enhances representation quality on diverse benchmarks.
Semi-pessimistic RL tackles distributional shift and data scarcity in offline RL.
problem Distributional shift and scarcity of labeled data in offline RL.
method Proposes a semi-pessimistic RL method that simplifies learning by seeking a lower bound of the reward function.
result Demonstrates clear competitiveness and improved policy learning with vast unlabeled data.
Optimal sampling strategy improves prediction accuracy with surrogate variables under measurement constraints.
problem Measurement-constrained datasets and lack of labeled data.
method A-optimality criterion for optimal sampling, leveraging surrogate variables.
result Achieves lower asymptotic variance and reduced empirical mean squared error.
SPI uses synthetic data to improve predictive inference efficiency.
problem Inefficient predictive inference with scarce calibration data.
method Integrates synthetic data to align nonconformity scores and improve coverage guarantees.
result SPI yields substantially tighter and more informative prediction sets.
Improved ranking method for scarce data with feature info.
problem Ranking items with limited comparisons and feature data.
method Modified RankCentrality using diffusion methods for feature info.
result Meaningful rankings even with scarce comparisons.
Improves functional linear regression with shape transfer learning.
problem Data scarcity in functional linear models.
method Shape-based transfer learning from auxiliary to target domains.
result Enhances robustness and generalizability of functional linear models.
New framework studies policy learning problems under data scarcity.
problem Learning improving policies when data is insufficient.
method Developed a mathematical framework for policy learning problems.
result Reduced policy learning problems to simpler ones in sample complexity.
FunDPS improves PDE solution recovery from sparse data.
problem Recovering whole solutions from sparse or noisy measurements in PDEs.
method Function-space diffusion model with gradient-based guidance.
result FunDPS achieves 32% accuracy improvement over state-of-the-art methods.
Improved preterm prediction using synthetic EHG signals.
problem Prediction bias towards term labor in preterm EHG data.
method Quantifying synthetic samples' effect, optimizing feature weights, and combining activation functions.
result Substantial improvement in prediction precision.
One of the major challenges in training deep architectures for predictive tasks is the scarcity and cost of labeled training data. Active Learning (AL) is one way of addressing this challenge. In stream-based AL, observations are continuously made available to the learner that have to decide whether to request a label …
A new method uses deep learning to efficiently sample rare transitions for estimating committor functions.
problem Efficiently sampling rare transitions to estimate committor functions in high-dimensional problems.
method DASTR (Deep Adaptive Sampling on Transition Paths) method using deep generative models.
result Significantly improved accuracy in approximating committor functions through efficient sampling.
Unified approach combines prediction-powered inference and variance reduction for semi-supervised optimization.
problem Scarcity of labeled data in semi-supervised optimization.
method PPI-SVRG, combining PPI and SVRG methods.
result Unified convergence bound with improved performance under label scarcity.
Proposes a new method to estimate individual treatment effects using unlabeled data.
problem Difficult estimation of individual treatment effects due to high costs of intervention studies.
method Combines causal inference matching and semi-supervised learning label propagation.
result Demonstrates successful mitigation of data scarcity in ITE estimation.
New method detects money laundering in Bitcoin using minimal labels.
problem Detecting money laundering in Bitcoin transactions with scarce labels.
method Active learning approach to anomaly detection.
result 5% of labels are sufficient to match supervised baseline performance.
Improved method for unbiased causal discovery in presence of unobserved confounding.
problem Unbiased data synthesis for causal discovery algorithms in the presence of unobserved confounding.
method Explicit block-hierarchical ancestral sampling to address limitations of implicit parameterization.
result Our approach fully covers the space of causal models, including those generated by implicit parameterization.
We investigate the geometry of word metrics on fundamental groups of manifolds associated with the generating sets consisting of elements represented by closed geodesics. We ask whether the diameter of such a metric is finite or infinite. The first answer we interpret as an abundance of closed geodesics, while the seco…
In domains such as health care and finance, shortage of labeled data and computational resources is a critical issue while developing machine learning algorithms. To address the issue of labeled data scarcity in training and deployment of neural network-based systems, we propose a new technique to train deep neural net…
Visual Speech Recognition (VSR) is the process of recognizing or interpreting speech by watching the lip movements of the speaker. Recent machine learning based approaches model VSR as a classification problem; however, the scarcity of training data leads to error-prone systems with very low accuracies in predicting un…
CAOS aggregates multiple one-shot predictors for efficient uncertainty quantification.
problem Lack of principled uncertainty quantification in one-shot prediction.
method CAOS, a conformal framework that aggregates multiple one-shot predictors and uses a leave-one-out calibration scheme.
result CAOS produces smaller prediction sets with reliable coverage compared to split conformal baselines.
This work tackles multivariate CDFs and copulas using tensor factorization.
problem Learning multivariate distributions, especially for mixed random variables, is challenging.
method Introducing a low-rank model for efficient sampling, inference, and uncertainty quantification.
result The proposed model outperforms traditional methods in various applications.
Detecting aggressive cancer tumors using ctDNA dynamics from few blood samples.
problem Early multi-cancer detection using circulating tumor DNA (ctDNA) levels.
method Combines continuous time Markov modelling and Signature theory for efficient testing procedures.
result Correctly addresses the challenge of data scarcity in cancer monitoring.
C-PP-COAD detects anomalies with limited real data, reducing dependency on real calibration data.
problem Limited real calibration data for online anomaly detection.
method Context-aware prediction-powered conformal online anomaly detection (C-PP-COAD).
result Significantly reduces dependency on real calibration data without compromising FDR control.
Mathematical Reinforcement Learning faces a 'Two-Hump' problem due to sparse rewards and a scarcity of intermediate 'hard-but-solvable' instances.
problem Mathematical search problems in Reinforcement Learning
method Novel data generation techniques and algorithmic enhancements
result Substantial performance improvements over previous baselines
DKPS provides guarantees for synthetic data from Transformer models, improving downstream tasks.
problem Lack of labeled data for building performant AI models.
method Data Kernel Perspective Space (DKPS) for mathematical analysis of synthetic data quality.
result Concrete statistical guarantees for the quality of transformer model outputs.