Model learns from multiple data sources for D2T and T2D tasks.
problem Limited performance due to single-source corpora.
method Variational auto-encoder with disentangled style and content variables.
result Model outperforms single-source counterpart on multiple datasets.
The paper tackles robust policy learning from multiple data sources.
problem Learning a policy that generalizes across diverse settings from multiple heterogeneous data sources.
method Proposes a minimax regret optimization objective and a policy learning algorithm combining doubly robust offline policy evaluation and no-regret learning.
result Achieves minimal worst-case mixture regret up to a moderated vanishing rate of the total data across all sources.
Study shows algorithms benefit from limited target data with many source domains.
problem Adapting to new domains with scarce labeled target data.
method New family of model selection algorithms.
result Beneficial guarantees in scenarios with limited target data.
New technique for multiple-source adaptation without density estimation.
problem Multiple-source adaptation problem.
method Discriminative technique that uses conditional probabilities from unlabeled data.
result Our technique outperforms previous generative solutions and other domain adaptation baselines.
Proposes using Wasserstein barycenters for robust optimization with multiple data sources.
problem Distributionally robust optimization with multiple heterogeneous data sources.
method Construct nominal distribution through Wasserstein barycenter of multiple data samples, reformulates as a finite convex program.
result Proposed scheme outperforms other estimators in sparse inverse covariance matrix estimation.
The paper aggregates predictions from multiple undisclosed datasets using conformal prediction.
problem Making valid predictions from multiple non-disclosed datasets without data sharing.
method Each dataset uses a transductive conformal predictor independently, and predictions are aggregated.
result The method produces valid and less variable aggregated predictions.
A new method for analyzing multi-source, multi-way data reduces dimensionality and reveals shared and individual structures.
problem Analyzing multi-source, multi-way data from different high-throughput technologies.
method Multiple Linked Tensor Factorization (MULTIFAC) extending CP decomposition with L2 penalties and EM algorithm for incomplete data.
result MULTIFAC approximates underlying signal, identifies shared and unshared structures, and imputes missing data.
The paper develops methods to identify stable associations across multiple studies.
problem Identifying stable associations across multiple studies with possible distributional shifts.
method Modeling heterogeneous multi-source data with multiple high-dimensional regressions and devising a novel sampling method for valid confidence intervals of maximin effects.
result Significant maximin effects indicate stable associations that can be generalized to target populations.
Adaptive kernel approach learns causal effects from diverse data sources.
problem Learning causal effects from multiple, decentralized data sources in a federated setting.
method Adaptive transfer algorithm using Random Fourier Features to estimate similarities and disentangle loss function components.
result Empirically outperforms baselines on decentralized data sources with different distributions.
Bayesian model predicts iron deficiency from multi-source multi-way molecular data.
problem Predicting iron deficiency in rhesus monkeys from multi-source multi-way molecular data.
method Developed a Bayesian approach with a linear model incorporating multi-way dependence and varying signal sizes across sources.
result Model accurately classifies iron deficiency in monkeys and outperforms simpler models.
Melanie improves predictive performance in non-stationary data streams by transferring knowledge between multiple sources.
problem Concept drift in data streams leads to poor predictive performance.
method Melanie uses multiple sub-classifiers to learn different aspects from various sources and compose an ensemble for the target concept.
result Melanie improves predictive performance over existing algorithms by leveraging multiple sources.
Improves understanding of PWS by calculating influence of sources and data.
problem Understanding the influence of each component in PWS.
method Proposes source-aware Influence Function (IF) to decompose and calculate influence.
result Improves end model's generalization performance and identifies mislabeling.
Motivation: Modelling methods that find structure in data are necessary with the current large volumes of genomic data, and there have been various efforts to find subsets of genes exhibiting consistent patterns over subsets of treatments. These biclustering techniques have focused on one data source, often gene expres…
New method combines multiple data sources for optimal decision-making with limited outcomes.
problem Optimal decision-making with limited outcome data from multiple heterogeneous sources.
method Calibrated optimal decision-making method leveraging common intermediate outcomes.
result Proposed estimator of conditional mean outcome is asymptotically normal and more efficient.
New method clusters medical codes using multiple data sources.
problem Challenges in community detection with multiple data sources.
method Combines multiple data sources with prior distance knowledge.
result Yields insightful clustering structure for ICD9 codes.
New PCA method handles multiple datasets and detects sparse patterns robustly.
problem Handling multi-source data with sparse and outlier-robust PCA.
method Developed a regularization problem with a penalty for structured sparsity and outlier resistance.
result The method detects global and local patterns across multiple data sources robustly.
EnMDAP aligns conditional distributions for multi-source domain adaptation using pseudolabels.
problem Training a target model with no labeled data in the absence of target data labels.
method EnMDAP uses label-wise moment matching and ensemble learning with multiple feature extractors.
result EnMDAP achieves state-of-the-art performance in multi-source domain adaptation tasks.
This work tackles robust multi-source domain adaptation under label shift.
problem Label shift and data contamination in multi-source domain adaptation.
method Domain-weighted empirical risk minimization framework with refinement procedure.
result The proposed method achieves superior performance in multi-category classification problems.
Proposes MDDA for multi-source domain adaptation.
problem Performance decay in deep neural networks due to domain shift between labeled and unlabeled data.
method Multi-source distilling domain adaptation (MDDA) network considering multiple source distributions and target similarities.
result Significantly outperforms state-of-the-art approaches on public DA benchmarks.
Combines prediction intervals from multiple non-disclosed sources.
problem Creating valid prediction intervals from multiple non-disclosed data sources.
method Train a conformal predictor on each data source independently and combine intervals.
result Produces valid prediction intervals with improved efficiency.
In the era of big data, a large amount of noisy and incomplete data can be collected from multiple sources for prediction tasks. Combining multiple models or data sources helps to counteract the effects of low data quality and the bias of any single model or data source, and thus can improve the robustness and the perf…
Study shows multi-source learning is more resilient to adversarial corruption than single-source learning.
problem Learning from multiple untrusted data sources, especially when some are adversarially corrupted.
method Analyzed the scenario where an adversary can corrupt a fixed fraction of data sources, derived a generalization bound for this setting.
result PAC-learnability is possible in the multi-source setting even when some data sources are adversarially corrupted.
Rugby-Bot predicts multiple metrics from a single source using fine-grain data.
problem Complexity of sporting events requires multiple metrics for accurate analysis.
method Multi-task learning with fine-grain spatial data and wide-and-deep learning.
result Predictions are consistent and can be in distribution form.
A method for integrating multiple cancer data sources using kernel principal component analysis.
problem Lack of comprehensive analysis of cancer subtypes from multiple data sources.
method Unsupervised data integration method based on kernel principal component analysis with a scoring function to determine input matrix impact.
result Enables visualization and clustering of integrated data for cancer subtype identification.
A new method for domain generalization using source-specific classifiers.
problem Generalizing across multiple sources for any target domain.
method Multiple domain-specific classifiers and a domain agnostic component.
result Improved performance on public benchmarks.
Paper proposes a new method to aggregate multiple sources with different label distributions.
problem Aggregating from multiple target-shifted sources with different label distributions.
method Unified framework to select relevant sources for domain adaptation with limited label, unsupervised, and label partial unsupervised scenarios.
result Empirical results significantly outperform baselines.
Proposes MinimaxFCM for multi-view clustering of data from multiple sources.
problem Clustering data from multiple heterogeneous views.
method Minimax optimization-based fuzzy c means clustering.
result MinimaxFCM outperforms other multi-view clustering methods.
Improves domain adaptation by combining multiple source domains and target domain data.
problem Poor performance of empirical risk minimization in distributionally shifted target domains.
method Distributionally robust model optimizing adversarial reward based on explained variance across multiple source domains.
result The robust model is a weighted average of conditional outcome models from source domains.
The paper tackles distribution-free prediction intervals for multi-source data.
problem Challenges in achieving valid inferences due to distribution shifts and privacy concerns.
method Derives efficient influence functions, incorporates machine learning, and proposes data-adaptive strategies.
result Achieves parametric rates of convergence to nominal coverage probabilities for prediction intervals.
Survey explores methods to adapt deep learning models across multiple labeled domains.
problem Difficulty in obtaining labeled data for deep learning models.
method Multi-source domain adaptation (MDA) to transfer knowledge from labeled to unlabeled or sparsely labeled target domains.
result MDA methods improve performance by minimizing domain shift.
The Rashomon effect shows many models can perform similarly, explored in this paper.
problem Why do many models perform similarly in machine learning?
method Categorized causes into statistical, structural, and procedural sources.
result Structural multiplicity persists and cannot be resolved without additional assumptions.
Bayesian approach for aggregating unreliable data sources to create accurate heatmaps.
problem Classifying regions with sparse, unreliable data from multiple sources.
method Bayesian Gaussian Process classifier that models reliability and bias of each data source.
result Reduces crowdsourced data needed and improves accuracy of heatmaps.
CAMul forecasts with calibrated and accurate multi-view time-series data.
problem Combining diverse data sources for reliable time-series forecasting.
method CAMul integrates multi-modal data views dynamically, assigning importance based on context.
result CAMul outperforms state-of-the-art models by 25% in accuracy and calibration.
Unified framework for multi-source data analysis improves network structure identification.
problem High dimensionality and heterogeneity in large-scale network data.
method msLBM framework combining multiple data sources for simultaneous grouping and connectivity analysis.
result Statistically optimal rates achieved for consensus knowledge graph learning.
Method constructs uniformly valid prediction sets across multiple distributions.
problem Uniformly valid prediction sets across multiple distributions.
method Max-p aggregation scheme and optimization programs.
result Optimal and efficient prediction sets for multiple distributions.
This work tackles collective matrix completion with multiple and heterogeneous data sources.
problem Reconstructing data from multiple heterogeneous matrices.
method Estimation based on minimizing goodness-of-fit and nuclear norm penalization of the whole collective matrix.
result Proposed estimators achieve fast rates of convergence under two settings.
Much information available on the web is copied, reused or rephrased. The phenomenon that multiple web sources pick up certain information is often called trend. A central problem in the context of web data mining is to detect those web sources that are first to publish information which will give rise to a trend. We p…
Enhances stock market prediction using multi-sourced data.
problem Improving stock market prediction by considering multiple data sources.
method Extended Coupled Hidden Markov Model incorporating historical trading data and news events, with correlations between stocks incorporated.
result Superior performance on China A-share market data in 2016 compared to previous methods.
The paper tackles personalized policy learning from diverse data sources in a federated setting.
problem Learning personalized decision policies from observational bandit feedback across multiple heterogeneous data sources.
method Introduces a novel regret analysis for distinguishing global and local regret, and presents a federated policy learning algorithm using local policies trained with doubly robust offline policy evaluation strategies.
result Establishes finite-sample upper bounds on global and local regret, characterizing them by source heterogeneity and distribution shift.
This study tackles offline RL with perturbed data sources, deriving a lower bound and proposing an optimal algorithm.
problem Understanding offline RL with multiple perturbed data sources.
method Derives an information-theoretic lower bound, proposes HetPEVI algorithm considering sample and source uncertainties.
result HetPEVI is optimal up to a polynomial factor of the horizon length and can solve offline RL tasks.
New method uses data characteristics for better random projections.
problem Improving random projections for better data reduction.
method Data-dependent random projections for superior performance.
result Proves superior performance in matrix multiplication, regression, and classification.
Face verification remains a challenging problem in very complex conditions with large variations such as pose, illumination, expression, and occlusions. This problem is exacerbated when we rely unrealistically on a single training data source, which is often insufficient to cover the intrinsically complex face variatio…
A decentralized approach for multi-source domain adaptation.
problem Transfer knowledge from multiple related domains to an unlabeled target domain.
method Federated Dataset Dictionary Learning (FedDaDiL) framework, eliminating central server, using Wasserstein barycenters.
result Our decentralized approach effectively adapts source domains to an unlabeled target domain.
New method recovers latent sources from multiple noisy views using deep neural networks.
problem Recovering a common latent source from multiple nonlinearly mixed views.
method Novel identifiability proofs using deep neural networks.
result Independent latent sources can be recovered from multiple noisy views using deep neural networks.
The paper proposes a method to infer user profiles from multiple sources of social media data.
problem Mining user profiles from social media data using a single type of information.
method Hinge-loss Markov Random Fields (HL-MRFs) integrated with multiple sources of UGC and social relations.
result HL-MRFs successfully incorporate multiple sources of information and outperform competing methods.
MRTL transfers knowledge across multiple target domains using shared latent factors.
problem Transfer learning in label-scarce target domains.
method MRTL uses collective nonnegative matrix tri-factorization to transfer knowledge from multiple sources to multiple targets.
result MRTL achieves better performance than state-of-the-art methods.
CoDATS improves DA on time series data with weak supervision.
problem Improving domain adaptation for time series data with limited labeled data.
method CoDATS model for Time Series data, DA-WS method with weak supervision.
result Significant accuracy improvements over state-of-the-art methods.
Deep learning predicts real-time parking occupancy using multiple data sources.
problem Predicting real-time parking occupancy in spatio-temporal networks.
method Graph-Convolutional Neural Networks (GCNN) for spatial relations, Recurrent Neural Networks (RNN) with Long-Short Term Memory (LSTM) for temporal features, multiple data sources.
result The model outperforms other methods with an average testing MAPE of 10.6%.