Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

146291437582 · Jun 202019922001200920182026
48 results for multi-site datasets

The paper presents a test to determine when pooling datasets from multiple sites is beneficial for regression analysis.

problem Determining when pooling multi-site datasets improves regression analysis power.
method Developed a hypothesis test for classical and high-dimensional linear regression to assess when pooling datasets is sensible.
result Identified specific conditions under which pooling datasets improves power in regression analysis.

HBR improves normative modeling of neuroimaging data across multiple sites.

problem Dealing with nuisance variation in neuroimaging data across different sites.
method Hierarchical Bayesian regression (HBR) for multi-site normative modeling.
result HBR provides more accurate normative ranges compared to existing methods.

Paper tackles robust decision-making from multiple sites with shared structure.

problem Learning robust sequential decisions from heterogeneous multi-site datasets.
method Group-Robust MDPs with d-rectangular uncertainty sets, feature-wise worst-case aggregation, and cluster-level pooling.
result Proves suboptimality bound for robust planning policy under robust partial coverage assumption.

New federated method preserves privacy and estimates treatment effects.

problem Privacy-preserving causal inference for multi-site studies.
method Multiply robust nuisance function estimation, transfer learning.
result Efficient and optimal treatment effect estimation under different scenarios.

FONT clusters patients across health systems with privacy and efficiency.

problem Challenges in multi-site cluster analysis due to data-sharing restrictions.
method Federated One-shot Ensemble Clustering (FONT) algorithm that requires only a single round of communication and exchanges only fitted model parameters and class labels.
result FONT improves consistency of patient clusters across sites compared to locally fitted clusters.

DWC consolidates neural networks trained on separate datasets, improving performance.

problem Training deep neural networks on large, distributed datasets is challenging.
method DWC: a continual learning method to consolidate weights of separate neural networks trained on independent datasets.
result DWC led to increased performance on test sets from different sites compared to an ensemble baseline.

Novel method uses image descriptors to harmonize MRI brain volumes across centers.

problem Inconsistencies in MRI brain volume measurements across different centers and scanners.
method Trained a Relevance Vector Machine (RVM) model using image descriptors to harmonize brain volumes.
result Decreases scanner and center variability while preserving measurements for longitudinal studies.

Deep chest X-ray classifiers show bias in predicting diagnoses.

problem Bias in deep learning classifiers predicting diagnoses from chest X-rays.
method Trained convolutional neural networks on multiple public datasets to predict 14 diagnostic labels.
result True positive rates vary significantly among different protected attributes, indicating bias.

New model extracts shared brain activity patterns from fMRI data.

problem Challenges in aggregating multi-subject fMRI data due to variability.
method Shared Gaussian Process Factor Analysis (S-GPFA) incorporating temporal information.
result Model reveals ground truth latent structures and replicates experimental performance.

Paper benchmarks CF mitigation in federated time series forecasting.

problem Catastrophic forgetting in federated learning for time series forecasting.
method Comprehensive evaluation of CF mitigation strategies in federated time series forecasting.
result Introduction of a new benchmark for CF in time series federated learning.

Study uses machine learning to classify autism based on brain connectivity variability.

problem Classifying autism using brain functional connectivity.
method Machine learning models trained on brain imaging data from ABIDE database.
result Increased FC variability in brain regions associated with low variability in ASD patients.

Improved 3D MRI classification using contrastive learning with continuous proxy metadata.

problem Insufficient labelled data for 3D medical image classification.
method Proposed a new loss function (y-Aware InfoNCE) to leverage continuous proxy metadata in contrastive learning.
result 3D CNN model pre-trained on 10^4 multi-site healthy brain MRI scans outperforms fully-supervised methods.

Transformer model pretrains on synthetic graphs for AD detection.

problem Limited labeled data and class imbalance in AD diagnosis.
method Diffusion-generated synthetic graphs, Graph Transformers, transfer learning.
result Framework outperforms baselines in AD diagnosis metrics.

Physics-Informed Neural Networks improve N2O flux predictions over classical models.

problem Predicting N2O flux emissions from agricultural processes.
method Constructed a rigorously derived physics residual from DayCent models and trained an MLP-based PINN on agricultural data.
result Physics-Informed Neural Networks consistently outperform classical models in predicting N2O flux emissions.

Infrastructure monitors AI/ML radiology models across multiple sites.

problem Monitoring and improving AI/ML radiology models across multiple sites.
method Interactive radiology reporting, centralized cloud system, post-marketing surveillance.
result Efficient monitoring and iterative development of AI/ML models without radiologist burden.

New fair regression method improves fairness in chronic kidney disease classification.

problem Mitigating societal bias in health care for multiple groups.
method Penalized fair regression framework for multiple groups, with penalties for true positive rate disparity.
result Achieves fairness-accuracy frontier beyond existing methods in simulations and real-world data.

FedRD improves risk difference estimation in federated learning for clinical outcomes.

problem Privacy-preserving model co-training in medical research is hindered by server-dependent architectures and focus on relative effect measures.
method FedRD is a server-independent, communication-efficient framework for federated risk difference estimation in distributed survival data.
result FedRD provides valid confidence intervals and hypothesis testing, and is asymptotically equivalent to pooled individual-level analysis.

RESPIRE calibrates low-cost air-quality sensors for CO levels, resistant to outliers.

problem Calibrating LCAQ sensors against regulatory-grade monitors is expensive and time-consuming.
method PROvably outlier-resistant semi-parametric regression technique.
result RESPIRE offers improved prediction in cross-site, cross-season, and cross-sensor settings.

New method calibrates asynchronous, error-prone covariates for longitudinal data.

problem Estimation biases and slow convergence in analyzing time-varying covariates with measurement error.
method Functional calibration approach based on functional principal component analysis.
result Asymptotically unbiased and consistent estimators for time-invariant coefficients; optimal convergence rate for time-varying coefficients.

Study models extreme skew surges along French Atlantic coast.

problem Appropriate modelling of extreme skew surges for coastal risk management.
method Peak-over-threshold framework, multivariate generalized Pareto distribution, extreme regression framework.
result Reconstructed historical skew surge time series at stations with limited data.

Synthetic learning improves neonatal brain MRI segmentation robustness.

problem Challenges in neonatal brain MRI segmentation due to image contrast and anatomical variations.
method Synthetic learning model trained on few T2-weighted volumes, then enhanced with motion artifacts and over-segmentation.
result Synthetic learning robust to image contrast and improves segmentation of both T1- and T2-weighted images.

Develops adaptive algorithms for sustainable fertilizer use in agriculture.

problem Sustaining high yields while reducing environmental impacts of fertilizer use.
method Nonlinear model-based bandit algorithms linking biological processes to decision-making.
result Faster learning and higher profits with interpretable recommendations.

Proposes a method to analyze distributed datasets without sharing original data.

problem Difficulty in centralizing large, distributed datasets due to size and privacy concerns.
method Centralizes intermediate representations instead of original datasets.
result Achieves higher prediction performance compared to individual analyses.

MusPy is a toolkit for symbolic music generation, providing tools for dataset management and analysis.

problem Facilitating the creation and analysis of symbolic music datasets.
method Development of an open-source Python library (MusPy) with features for dataset management, data I/O, preprocessing, and model evaluation. Demonstrated through statistical analysis and cross-dataset generalizability experiments.
result MusPy's dataset analysis reveals varying degrees of cross-genre representation across different music datasets.

StyleDiff compares unlabeled datasets using disentangled image spaces.

problem Mismatches between development and real-world datasets lead to inaccurate predictions.
method Uses disentangled image spaces and focuses on attributes to compare datasets.
result Accurately detects and presents differences between datasets.

New framework assesses graph-learning datasets for better evaluation.

problem Insufficient evaluation of graph-learning datasets and methods.
method Introduces Rings framework for dataset ablations and proposes performance separability and mode complementarity measures.
result Demonstrates utility of Rings framework for graph-learning dataset evaluation.

Improved dataset distillation for images and texts boosts model accuracy.

problem Reducing dataset size for faster and more energy-efficient model training.
method Simultaneous distillation of images and soft labels, extending to text datasets.
result 2-4% increase in accuracy for image classification tasks, 20% reduction in distilled samples.

Paper introduces ToyADMOS dataset for detecting anomalous machine sounds.

problem Lack of large-scale datasets for ADMOS anomaly detection.
method Collected anomalous sounds of miniature machines by deliberate damage.
result Released dataset includes over 180 hours of normal and 4,000 anomalous sounds.

MTL method uses unlabeled data with pseudo labels to improve classification with disjoint datasets.

problem Improving classification performance with disjoint labeled datasets using unlabeled data.
method Proposes MTL-SA method to select and augment unlabeled data with confident pseudo labels and close distribution to labeled data.
result Extensive experiments show the effectiveness of MTL-SA method in improving classification performance.

Study shows pruning datasets can improve machine learning model performance.

problem Improving machine learning model performance through dataset pruning.
method Comparison of different algorithms on unpruned and iteratively pruned datasets.
result Algorithms that perform better on unpruned datasets also perform better on pruned datasets.

The study examines dataset usage patterns in machine learning research.

problem Lack of attention to dataset dynamics in machine learning research.
method Analysis of dataset usage patterns across machine learning subcommunities and time periods (2015-2020).
result Increasing concentration on fewer and fewer datasets, significant adoption from other tasks, and concentration across the field on datasets introduced by elite institutions.