Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

2795588371,116 · Jun 202019922001200920172026
48 results for data scarcity

Proposes a new VAE framework for anomaly detection in time series data.

problem Data scarcity leads to latent holes and discontinuous regions in latent space, causing non-robust reconstructions.
method Combines VAEs with self-supervised learning to address data scarcity and improve anomaly detection.
result Improves robustness of anomaly detection in time series data by addressing latent holes and discontinuities.

FedACS uses attention to select clients with similar data for federated learning.

problem Non-IID data and data scarcity in federated learning.
method FedACS integrates an attention mechanism to prioritize clients with similar data distributions.
result FedACS improves federated learning performance by addressing non-IID data and data scarcity.

GCNs help in diagnosing label scarcity and feature quality on graphs.

problem Understanding when GCNs improve node classification.
method Simulated label scarcity, feature ablation, and per-class analysis.
result GCNs provide largest gains under extreme label scarcity, matching original performance with noisy features, but hurt when homophily is low and features are strong.

DiffsFormer uses AI-generated samples to improve stock forecasting accuracy.

problem Data scarcity in stock forecasting, including low signal-to-noise ratio and data homogeneity.
method DiffsFormer employs a Diffusion Model with a Transformer architecture to generate augmented stock factors.
result DiffsFormer achieves significant improvements in stock forecasting accuracy (7.2% and 27.8% relative improvements for CSI300 and CSI800 datasets, respectively).

Semi-pessimistic RL tackles distributional shift and data scarcity in offline RL.

problem Distributional shift and scarcity of labeled data in offline RL.
method Proposes a semi-pessimistic RL method that simplifies learning by seeking a lower bound of the reward function.
result Demonstrates clear competitiveness and improved policy learning with vast unlabeled data.

Paper tackles RUL prediction with scarce data using indirect supervision.

problem Predicting RUL with indirect supervision and scarce time series data.
method Unified framework called parameterized static regression, handling data scarcity without interpolation.
result Competitive performance in prediction accuracy with simulated data scarcity.

Model predicts political ideology using context vectors to mitigate bias and scarcity.

problem Scarcity and selection bias in political ideology prediction.
method Proposes a statistical model decomposing embeddings into context and position vectors, training an end-to-end model for deployment.
result Model can predict ideological labels even with minimal biased data, outperforming state-of-the-art methods.

Proposes a new method to estimate individual treatment effects using unlabeled data.

problem Difficult estimation of individual treatment effects due to high costs of intervention studies.
method Combines causal inference matching and semi-supervised learning label propagation.
result Demonstrates successful mitigation of data scarcity in ITE estimation.

C-PP-COAD detects anomalies with limited real data, reducing dependency on real calibration data.

problem Limited real calibration data for online anomaly detection.
method Context-aware prediction-powered conformal online anomaly detection (C-PP-COAD).
result Significantly reduces dependency on real calibration data without compromising FDR control.

DKPS provides guarantees for synthetic data from Transformer models, improving downstream tasks.

problem Lack of labeled data for building performant AI models.
method Data Kernel Perspective Space (DKPS) for mathematical analysis of synthetic data quality.
result Concrete statistical guarantees for the quality of transformer model outputs.

RDLI integrates domain logic and context grounding to detect crypto anomalies under scarce labels.

problem Extreme label scarcity and evasion strategies in crypto networks.
method Relational Domain Logic Integration (RDLI) with Retrieval Grounded Context (RGC).
result RDLI outperforms GNN baselines by 28.9% in F1 score under 0.01% label scarcity.

Unified approach combines prediction-powered inference and variance reduction for semi-supervised optimization.

problem Scarcity of labeled data in semi-supervised optimization.
method PPI-SVRG, combining PPI and SVRG methods.
result Unified convergence bound with improved performance under label scarcity.

Paper tackles robust prediction of nuclear reactor materials under scarce data.

problem Challenges of data scarcity and uncertainty in nuclear reactor design.
method Meta-learning approach informed by uncertainty and prior knowledge.
result Achieves superior performance in rupture life prediction.

This work discovers governing equations from limited data using physics-informed deep learning.

problem Discovering governing equations from scarce and noisy data for complex systems.
method Physics-informed deep learning framework integrating neural networks, physics embedding, and sparse regression.
result The method effectively identifies governing equations from various spatiotemporal systems with different levels of data scarcity and noise.

BLADE uses Bayesian methods to discover complex systems from scarce data.

problem Efficiently discovering governing equations of complex dynamical systems from limited data.
method Combines replica-exchange stochastic gradient Langevin Monte Carlo with active learning.
result Reduces measurement requirements by 60% for Lotka-Volterra and 40% for Burgers' equation.

Flood forecasts are crucial for effective individual and governmental protective action. The vast majority of flood-related casualties occur in developing countries, where providing spatially accurate forecasts is a challenge due to scarcity of data and lack of funding. This paper describes an operational system provid…

2019-10-11abs ↗pdf ↗

Mathematical Reinforcement Learning faces a 'Two-Hump' problem due to sparse rewards and a scarcity of intermediate 'hard-but-solvable' instances.

problem Mathematical search problems in Reinforcement Learning
method Novel data generation techniques and algorithmic enhancements
result Substantial performance improvements over previous baselines

Local Gaussian correlation struggles in tails but a new method improves it.

problem Local Gaussian correlation's limitations in tail dependence.
method A new adaptive bandwidth method for LGC, optimizing for local effective sample size.
result Adaptive bandwidths outperform global ones in moderate dependence, but not in strong or weak dependence.

Self-supervised learning improves RUL prediction with limited data in fatigue damage prognosis.

problem Limited labelled data for RUL prediction in fatigue damage prognosis.
method Pre-training deep learning models on unlabelled sensor data using self-supervised learning.
result Self-supervised pre-trained models significantly outperform non-pre-trained models in RUL prediction with scarce labelled data.

Personalizes pre-trained models for nonparametric regression with limited data.

problem Improving data efficiency in nonparametric regression with few samples.
method Develops a theoretical framework and algorithms for few-shot personalization of black-box models.
result Achieves minimax optimal rate for personalization in nonparametric regression.

Fused Encoder Networks improve momentum strategies on crypto data.

problem Deploying momentum strategies on crypto data with limited samples leads to over-fitted models.
method Hybrid transfer learning model combining source and target datasets.
result Fused Encoder Networks outperform classical momentum strategies and benchmarks.

Generative Adversarial Networks create synthetic data for structural damage detection.

problem Data scarcity in structural damage detection.
method 1-D Wasserstein Deep Convolutional Generative Adversarial Networks (1-D WDCGAN-GP) for synthetic data generation.
result Generated synthetic data improves damage detection accuracy in 1-D Deep Convolutional Neural Networks.

Synthetic augmentation improves financial machine learning performance in variance-dominant regimes.

problem Data scarcity in financial machine learning.
method Formalized synthetic augmentation, introduced size-matched null augmentation, and developed a non-parametric block permutation test.
result Synthetic augmentation is beneficial only in variance-dominant regimes, such as persistent volatility forecasting.

New method tackles dynamic data labeling issues with limited labels.

problem Dynamic data labeling with scarce labeled instances.
method Instance exploitation technique for aggressive model adaptation.
result Aggressive model adaptation leads to better performance than standard methods.

This work addresses fairness in ML models by training and evaluating attribute classifiers under uncertain and incomplete data.

problem Challenges in fairness metrics due to uncertain and incomplete data.
method Developed a theoretical and empirical analysis to understand and improve bias estimation in the data-scarce regime.
result The test accuracy of the attribute classifier is not always correlated with its effectiveness in bias estimation.

Transfer learning improves loan recovery rate forecasting under data scarcity.

problem Data scarcity in loan portfolios limits RR modeling accuracy.
method Introduces FT-MDN-Transformer, a mixture-density tabular Transformer architecture for TL.
result FT-MDN-Transformer outperforms baseline models in RR forecasting, especially under covariate and conditional shifts.

Generative algorithms learn high-dimensional data efficiently and generate new samples.

problem Learning from scarce high-dimensional data.
method Lipschitz-regularized gradient flows and particle-based algorithms.
result Correctly transports gene expression data points with high dimensionality.

Self-supervised learning improves ECG classification performance.

problem Label scarcity in clinical 12-lead ECG data.
method Adapted self-supervised methods to ECG domain, focusing on contrastive representations and latent forecasting.
result Contrastive predictive coding adaptation yields linear evaluation performance only 0.5% below supervised performance.

Continuous time models in the theory of real options give explicit formulas for optimal exercise strategies when options are simple and the price of an underlying asset follows a geometric Brownian motion. This paper suggests a general, computationally simple approach to real options in discrete time. Explicit formulas…

2004-04-05abs ↗pdf ↗

The goal of this paper is to deal with a data scarcity scenario where deep learning techniques use to fail. We compare the use of two well established techniques, Restricted Boltzmann Machines and Variational Auto-encoders, as generative models in order to increase the training set in a classification framework. Essent…

2019-03-21abs ↗pdf ↗

We investigate the geometry of word metrics on fundamental groups of manifolds associated with the generating sets consisting of elements represented by closed geodesics. We ask whether the diameter of such a metric is finite or infinite. The first answer we interpret as an abundance of closed geodesics, while the seco…

2019-04-25abs ↗pdf ↗

SAGDA generates synthetic agricultural datasets to improve ML in African farming.

problem Data scarcity in African agriculture limits machine learning innovations.
method SAGDA is an open-source Python toolkit that generates, augments, and validates synthetic agricultural datasets.
result SAGDA enhances ML applications in agriculture, such as yield prediction and fertilizer recommendation.

Develops methods to simulate rare transitions in molecular systems.

problem Rare transitions between metastable states in molecular systems are difficult to study due to limited data.
method Two novel methods: chain-based and midpoint-based approaches.
result Demonstrates effectiveness of methods in both data-rich and data-scarce scenarios.

In this work we study the problem of inferring a discrete probability distribution using both expert knowledge and empirical data. This is an important issue for many applications where the scarcity of data prevents a purely empirical approach. In this context, it is common to rely first on an initial domain knowledge …

2019-10-20abs ↗pdf ↗

In domains such as health care and finance, shortage of labeled data and computational resources is a critical issue while developing machine learning algorithms. To address the issue of labeled data scarcity in training and deployment of neural network-based systems, we propose a new technique to train deep neural net…

2018-10-14abs ↗pdf ↗