The study improves theoretical understanding of using multiple synthetic datasets for better model accuracy.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
FedSyn generates synthetic data from multiple organizations' datasets.
Enhances anomaly detection using multiple reference datasets.
Synthetic telematics dataset created from insurance claims data.
Dataset distillation is a method for reducing dataset sizes by learning a small number of synthetic samples containing all the information of a large dataset. This has several benefits like speeding up model training, reducing energy consumption, and reducing required storage space. Currently, each synthetic sample is …
Paper proposes HCDC to improve hyperparameter search efficiency.
DIGEN benchmark provides synthetic datasets for ML algorithm evaluation.
Adaptive methods learn from multiple datasets, leveraging similarities and robust to outliers.
Online data augmentation improves forecasting performance in deep learning.
M3E2 neural network estimates multiple treatment effects.
Non-orthogonal joint diagonalization (NJD) free of prewhitening has been widely studied in the context of blind source separation (BSS) and array signal processing, etc. However, NJD is used to retrieve the jointly diagonalizable structure for a single set of target matrices which are mostly formulized with a single da…
Principal component analysis (PCA) is widely used for feature extraction and dimensionality reduction, with documented merits in diverse tasks involving high-dimensional data. Standard PCA copes with one dataset at a time, but it is challenged when it comes to analyzing multiple datasets jointly. In certain data scienc…
We introduce a method called multi-scale local shape analysis, or MLSA, for extracting features that describe the local structure of points within a dataset. The method uses both geometric and topological features at multiple levels of granularity to capture diverse types of local information for subsequent machine lea…
Paper benchmarks machine learning for detecting process curve drifts.
Improves accuracy and fairness in prediction systems with multiple domain experts.
SYNC generates synthetic data from aggregated sources using Gaussian copulas.
SCARY dataset generates complex causal scenarios for causality research.
Deep neural networks (DNNs) have recently received vast attention in applications requiring classification of radar returns, including radar-based human activity recognition for security, smart homes, assisted living, and biomedicine. However,acquiring a sufficiently large training dataset remains a daunting task due t…
Study shows current image classification models lack robustness to real-world dataset shifts.
Detect hidden confounding in observational data using multiple environments.
Sparse GCA finds linear relationships in multiple datasets, using gradient descent.
This paper considers a semi-supervised learning framework for weakly labeled polyphonic sound event detection problems for the DCASE 2019 challenge's task4 by combining both the tri-training and adversarial learning. The goal of the task4 is to detect onsets and offsets of multiple sound events in a single audio clip. …
Private PGB boosts synthetic data quality using GANs and privacy techniques.
Principal component analysis (PCA) has well-documented merits for data extraction and dimensionality reduction. PCA deals with a single dataset at a time, and it is challenged when it comes to analyzing multiple datasets. Yet in certain setups, one wishes to extract the most significant information of one dataset relat…
In this paper, we propose a mixture of probabilistic partial canonical correlation analysis (MPPCCA) that extracts the Causal Patterns from two multivariate time series. Causal patterns refer to the signal patterns within interactions of two elements having multiple types of mutually causal relationships, rather than a…
AutoSimulate efficiently optimizes synthetic data generation.
Nonnegative matrix factorization (NMF), a dimensionality reduction and factor analysis method, is a special case in which factor matrices have low-rank nonnegative constraints. Considering the stochastic learning in NMF, we specifically address the multiplicative update (MU) rule, which is the most popular, but which h…
New method constructs synthetic treatment groups without mean exchangeability assumption.
SynthBH uses synthetic data to control FDR in multiple testing.
Properties of data are frequently seen to vary depending on the sampled situations, which usually changes along a time evolution or owing to environmental effects. One way to analyze such data is to find invariances, or representative features kept constant over changes. The aim of this paper is to identify one such fe…
Motivated by electricity consumption metering, we extend existing nonnegative matrix factorization (NMF) algorithms to use linear measurements as observations, instead of matrix entries. The objective is to estimate multiple time series at a fine temporal scale from temporal aggregates measured on each individual serie…
Integrates MRF into multimodal VAE for better complex intermodal interactions.
The rising interest in pattern recognition and data analytics has spurred the development of innovative machine learning algorithms and tools. However, as each algorithm has its strengths and limitations, one is motivated to judiciously fuse multiple algorithms in order to find the "best" performing one, for a given da…
AKO improves stability and power of Knockoff inference.
The problem of missing values in multivariable time series is a key challenge in many applications such as clinical data mining. Although many imputation methods show their effectiveness in many applications, few of them are designed to accommodate clinical multivariable time series. In this work, we propose a multiple…
Approaches for approximating persistent homology for large datasets.
In this paper, we propose a novel method for generating a synthetic dataset obeying Gaussian distribution. Compared to the commonly used benchmark datasets with unknown distribution, the synthetic dataset has an explicit distribution, i.e., Gaussian distribution. Meanwhile, it has the same characteristics as the benchm…
Overparameterized MLR fits hyper-curves, improving model robustness.
Multiple clustering aims at discovering diverse ways of organizing data into clusters. Despite the progress made, it's still a challenge for users to analyze and understand the distinctive structure of each output clustering. To ease this process, we consider diverse clusterings embedded in different subspaces, and ana…
Interactive learning is a process in which a machine learning algorithm is provided with meaningful, well-chosen examples as opposed to randomly chosen examples typical in standard supervised learning. In this paper, we propose a new method for interactive learning from multiple noisy labels where we exploit the disagr…
CardiCat generates synthetic data for high-cardinality tabular datasets.
Synthetic population generation is the process of combining multiple socioeconomic and demographic datasets from different sources and/or granularity levels, and downscaling them to an individual level. Although it is a fundamental step for many data science tasks, an efficient and standard framework is absent. In this…
DP-CDA generates synthetic data to enhance privacy in high-dimensional datasets.
Granger causality analysis, as one of the most popular time series causality methods, has been widely used in the economics, neuroscience. However, unobserved confounders is a fundamental problem in the observational studies, which is still not solved for the non-linear Granger causality. The application works often de…
Enhanced synthetic dataset improves asset allocation analysis.
DP-FedTabDiff generates private synthetic tabular data using diffusion models and differential privacy.
Paper proposes AdaDetect for FDR-controlled novelty detection.
RAMEN corrects observational data biases for multiple environments.