The paper presents a test to determine when pooling datasets from multiple sites is beneficial for regression analysis.
problem Determining when pooling multi-site datasets improves regression analysis power.
method Developed a hypothesis test for classical and high-dimensional linear regression to assess when pooling datasets is sensible.
result Identified specific conditions under which pooling datasets improves power in regression analysis.
Resting-state functional Magnetic Resonance Imaging (R-fMRI) holds the promise to reveal functional biomarkers of neuropsychiatric disorders. However, extracting such biomarkers is challenging for complex multi-faceted neuropatholo-gies, such as autism spectrum disorders. Large multi-site datasets increase sample sizes…
HBR improves normative modeling of neuroimaging data across multiple sites.
problem Dealing with nuisance variation in neuroimaging data across different sites.
method Hierarchical Bayesian regression (HBR) for multi-site normative modeling.
result HBR provides more accurate normative ranges compared to existing methods.
Paper tackles robust decision-making from multiple sites with shared structure.
problem Learning robust sequential decisions from heterogeneous multi-site datasets.
method Group-Robust MDPs with d-rectangular uncertainty sets, feature-wise worst-case aggregation, and cluster-level pooling.
result Proves suboptimality bound for robust planning policy under robust partial coverage assumption.
New federated method preserves privacy and estimates treatment effects.
problem Privacy-preserving causal inference for multi-site studies.
method Multiply robust nuisance function estimation, transfer learning.
result Efficient and optimal treatment effect estimation under different scenarios.
FONT clusters patients across health systems with privacy and efficiency.
problem Challenges in multi-site cluster analysis due to data-sharing restrictions.
method Federated One-shot Ensemble Clustering (FONT) algorithm that requires only a single round of communication and exchanges only fitted model parameters and class labels.
result FONT improves consistency of patient clusters across sites compared to locally fitted clusters.
New method corrects MRI biases across scanners and sites.
problem Site and scanner biases in diffusion MRI data.
method Learning invariant representations using variational auto-encoders (VAE).
result Improvements on test data relative to a baseline method.
DWC consolidates neural networks trained on separate datasets, improving performance.
problem Training deep neural networks on large, distributed datasets is challenging.
method DWC: a continual learning method to consolidate weights of separate neural networks trained on independent datasets.
result DWC led to increased performance on test sets from different sites compared to an ensemble baseline.
Novel method uses image descriptors to harmonize MRI brain volumes across centers.
problem Inconsistencies in MRI brain volume measurements across different centers and scanners.
method Trained a Relevance Vector Machine (RVM) model using image descriptors to harmonize brain volumes.
result Decreases scanner and center variability while preserving measurements for longitudinal studies.
Deep chest X-ray classifiers show bias in predicting diagnoses.
problem Bias in deep learning classifiers predicting diagnoses from chest X-rays.
method Trained convolutional neural networks on multiple public datasets to predict 14 diagnostic labels.
result True positive rates vary significantly among different protected attributes, indicating bias.
New model extracts shared brain activity patterns from fMRI data.
problem Challenges in aggregating multi-subject fMRI data due to variability.
method Shared Gaussian Process Factor Analysis (S-GPFA) incorporating temporal information.
result Model reveals ground truth latent structures and replicates experimental performance.
Paper benchmarks CF mitigation in federated time series forecasting.
problem Catastrophic forgetting in federated learning for time series forecasting.
method Comprehensive evaluation of CF mitigation strategies in federated time series forecasting.
result Introduction of a new benchmark for CF in time series federated learning.
Paper proposes federated offline RL for personalized medicine.
problem Privacy constraints and heterogeneity in healthcare data.
method Multi-site Markov decision process model and first federated policy optimization algorithm.
result The proposed algorithm achieves comparable suboptimality to centralized RL.
Study uses machine learning to classify autism based on brain connectivity variability.
problem Classifying autism using brain functional connectivity.
method Machine learning models trained on brain imaging data from ABIDE database.
result Increased FC variability in brain regions associated with low variability in ASD patients.
Causal analysis reveals regional discrepancies in TOPCAT trial results.
problem Inconclusive results in TOPCAT trial for heart failure treatment.
method Causal discovery methods with domain knowledge integration.
result Significant causal effects shown for some subgroups globally.
Improved 3D MRI classification using contrastive learning with continuous proxy metadata.
problem Insufficient labelled data for 3D medical image classification.
method Proposed a new loss function (y-Aware InfoNCE) to leverage continuous proxy metadata in contrastive learning.
result 3D CNN model pre-trained on 10^4 multi-site healthy brain MRI scans outperforms fully-supervised methods.
Transformer model pretrains on synthetic graphs for AD detection.
problem Limited labeled data and class imbalance in AD diagnosis.
method Diffusion-generated synthetic graphs, Graph Transformers, transfer learning.
result Framework outperforms baselines in AD diagnosis metrics.
Methods for prediction and tolerance intervals in non-normal models.
problem Constructing prediction and tolerance intervals for non-normal data.
method Two approaches: pivotal quantity approximation and confidence interval for mean.
result Intuitive, simple, efficient methods with proper operating characteristics.
Paper proposes a method to improve DAG structure reconstruction using auxiliary DAGs.
problem Improving DAG structure reconstruction with limited data.
method Introduces structural similarity measures and a transfer learning framework.
result Significant improvement in DAG reconstruction, even with dissimilar auxiliary DAGs.
Physics-Informed Neural Networks improve N2O flux predictions over classical models.
problem Predicting N2O flux emissions from agricultural processes.
method Constructed a rigorously derived physics residual from DayCent models and trained an MLP-based PINN on agricultural data.
result Physics-Informed Neural Networks consistently outperform classical models in predicting N2O flux emissions.
Infrastructure monitors AI/ML radiology models across multiple sites.
problem Monitoring and improving AI/ML radiology models across multiple sites.
method Interactive radiology reporting, centralized cloud system, post-marketing surveillance.
result Efficient monitoring and iterative development of AI/ML models without radiologist burden.
New fair regression method improves fairness in chronic kidney disease classification.
problem Mitigating societal bias in health care for multiple groups.
method Penalized fair regression framework for multiple groups, with penalties for true positive rate disparity.
result Achieves fairness-accuracy frontier beyond existing methods in simulations and real-world data.
FedRD improves risk difference estimation in federated learning for clinical outcomes.
problem Privacy-preserving model co-training in medical research is hindered by server-dependent architectures and focus on relative effect measures.
method FedRD is a server-independent, communication-efficient framework for federated risk difference estimation in distributed survival data.
result FedRD provides valid confidence intervals and hypothesis testing, and is asymptotically equivalent to pooled individual-level analysis.
Imputation-Powered Inference improves subpopulation efficiency in missing data settings.
problem Complex missing data patterns challenge standard inference methods.
method Imputation-Powered Inference (IPI) combines blackbox imputation with bias correction.
result IPI provides valid and efficient M-estimation under MCAR blockwise missingness.
RESPIRE calibrates low-cost air-quality sensors for CO levels, resistant to outliers.
problem Calibrating LCAQ sensors against regulatory-grade monitors is expensive and time-consuming.
method PROvably outlier-resistant semi-parametric regression technique.
result RESPIRE offers improved prediction in cross-site, cross-season, and cross-sensor settings.
New method calibrates asynchronous, error-prone covariates for longitudinal data.
problem Estimation biases and slow convergence in analyzing time-varying covariates with measurement error.
method Functional calibration approach based on functional principal component analysis.
result Asymptotically unbiased and consistent estimators for time-invariant coefficients; optimal convergence rate for time-varying coefficients.
Study models extreme skew surges along French Atlantic coast.
problem Appropriate modelling of extreme skew surges for coastal risk management.
method Peak-over-threshold framework, multivariate generalized Pareto distribution, extreme regression framework.
result Reconstructed historical skew surge time series at stations with limited data.
Method synthesizes 4D CMR images from XCAT model using GAN and SPADE.
problem Synthesizing realistic 4D CMR images with annotations and adaptable styles.
method Hybrid GAN approach with XCAT anatomical ground truth and SPADE for semantic preservation.
result Synthesized images with modality-specific features learned from real CMR data.
Synthetic learning improves neonatal brain MRI segmentation robustness.
problem Challenges in neonatal brain MRI segmentation due to image contrast and anatomical variations.
method Synthetic learning model trained on few T2-weighted volumes, then enhanced with motion artifacts and over-segmentation.
result Synthetic learning robust to image contrast and improves segmentation of both T1- and T2-weighted images.
Efficient method for shape modeling invariant to rigid motion.
problem Statistical shape modeling for rigidly moving shapes.
method Non-Euclidean Lie group analysis of metric distortion and curvature.
result Outperforms state-of-the-art classifiers in sparse data.
Develops adaptive algorithms for sustainable fertilizer use in agriculture.
problem Sustaining high yields while reducing environmental impacts of fertilizer use.
method Nonlinear model-based bandit algorithms linking biological processes to decision-making.
result Faster learning and higher profits with interpretable recommendations.
Method detects batch heterogeneity in genomic data.
problem Batch effects confound genomic diagnostics.
method Bayesian model evidence clustering.
result Detects batch effects without known labels.
Dependent MMD coresets help compare multiple related datasets.
problem Comparing multiple related datasets for insights into model generalization.
method Dependent MMD coresets for collections of datasets.
result Dependent MMD coresets facilitate comparison and understanding of multiple related datasets.
Synthetic dataset for deep learning with known Gaussian distribution.
problem Lack of datasets with known distribution for deep learning verification.
method Proposes a method to generate a synthetic dataset with Gaussian distribution.
result Synthetic dataset mimics MNIST and can be used with DNNs.
Proposes a method to analyze distributed datasets without sharing original data.
problem Difficulty in centralizing large, distributed datasets due to size and privacy concerns.
method Centralizes intermediate representations instead of original datasets.
result Achieves higher prediction performance compared to individual analyses.
SCARY dataset generates complex causal scenarios for causality research.
problem Lack of complexity in existing causal datasets.
method Synthetic dataset with 40 scenarios, three seeds, and two data generation mechanisms.
result Provides a valuable resource for realistic causal discovery.
MusPy is a toolkit for symbolic music generation, providing tools for dataset management and analysis.
problem Facilitating the creation and analysis of symbolic music datasets.
method Development of an open-source Python library (MusPy) with features for dataset management, data I/O, preprocessing, and model evaluation. Demonstrated through statistical analysis and cross-dataset generalizability experiments.
result MusPy's dataset analysis reveals varying degrees of cross-genre representation across different music datasets.
New handwritten digits dataset for Kannada script.
problem Lack of datasets for Kannada numeral digits.
method Developed Kannada-MNIST and Dig-MNIST datasets.
result Initial CNN accuracy is lower than MNIST, indicating a challenge in generalization.
StyleDiff compares unlabeled datasets using disentangled image spaces.
problem Mismatches between development and real-world datasets lead to inaccurate predictions.
method Uses disentangled image spaces and focuses on attributes to compare datasets.
result Accurately detects and presents differences between datasets.
Paper introduces a new pedestrian dataset for adverse weather conditions.
problem Lack of controlled and annotated pedestrian datasets for adverse weather conditions.
method Presentation of a new dataset and baseline results for various machine learning tasks.
result Baseline results for various tasks on the Cerema AWP dataset.
New framework assesses graph-learning datasets for better evaluation.
problem Insufficient evaluation of graph-learning datasets and methods.
method Introduces Rings framework for dataset ablations and proposes performance separability and mode complementarity measures.
result Demonstrates utility of Rings framework for graph-learning dataset evaluation.
Improved dataset distillation for images and texts boosts model accuracy.
problem Reducing dataset size for faster and more energy-efficient model training.
method Simultaneous distillation of images and soft labels, extending to text datasets.
result 2-4% increase in accuracy for image classification tasks, 20% reduction in distilled samples.
Paper introduces ToyADMOS dataset for detecting anomalous machine sounds.
problem Lack of large-scale datasets for ADMOS anomaly detection.
method Collected anomalous sounds of miniature machines by deliberate damage.
result Released dataset includes over 180 hours of normal and 4,000 anomalous sounds.
Meta-Dataset benchmarks few-shot learning models with diverse datasets.
problem Lack of diverse and realistic datasets for evaluating few-shot learning models.
method Meta-Dataset: a new benchmark with diverse datasets and realistic tasks.
result Meta-Dataset uncovers important research challenges in few-shot learning.
Fairness GAN generates fair datasets for decision making.
problem Creating fair datasets for allocative decision making.
method Novel auxiliary classifier GAN aiming for demographic parity or equality of opportunity.
result Improves demographic parity and equality of opportunity in generated images.
MTL method uses unlabeled data with pseudo labels to improve classification with disjoint datasets.
problem Improving classification performance with disjoint labeled datasets using unlabeled data.
method Proposes MTL-SA method to select and augment unlabeled data with confident pseudo labels and close distribution to labeled data.
result Extensive experiments show the effectiveness of MTL-SA method in improving classification performance.
Study shows pruning datasets can improve machine learning model performance.
problem Improving machine learning model performance through dataset pruning.
method Comparison of different algorithms on unpruned and iteratively pruned datasets.
result Algorithms that perform better on unpruned datasets also perform better on pruned datasets.
The study examines dataset usage patterns in machine learning research.
problem Lack of attention to dataset dynamics in machine learning research.
method Analysis of dataset usage patterns across machine learning subcommunities and time periods (2015-2020).
result Increasing concentration on fewer and fewer datasets, significant adoption from other tasks, and concentration across the field on datasets introduced by elite institutions.