Combines classifiers from different types to improve ensemble accuracy.
problem Improving ensemble accuracy by combining classifiers of different types.
method Builds heterogeneous ensembles by pooling classifiers from multiple homogeneous ensembles, using cross-validation or out-of-bag data for optimal composition.
result Optimal heterogeneous ensemble compositions can be determined using cross-validation or out-of-bag data.
MNIST-NET10 fusion improves MNIST classification to 0.1% error rate.
problem Improving MNIST classification accuracy.
method Complex heterogeneous fusion architecture using degree of certainty aggregation.
result MNIST-NET10 achieves 0.1% error rate with 10 misclassifications.
Theory and method for reducing prediction variance in noisy feature-subsampled ridge ensembles.
problem Reduction of prediction variance in noisy data with feature bagging.
method Developed analytical learning curves for noisy ridge ensembles, introduced heterogeneous feature ensembling.
result Subsampling shifts the double-descent peak, leading to improved performance over a single linear predictor.
We propose a novel "tree-averaging" model that utilizes the ensemble of classification and regression trees (CART). Each constituent tree is estimated with a subset of similar data. We treat this grouping of subsets as Bayesian ensemble trees (BET) and model them as an infinite mixture Dirichlet process. We show that B…
Paper uses ensemble learning for IoT cybersecurity anomaly detection.
problem Anomaly detection in IoT data is challenging due to heterogeneous device types.
method Bayesian hyperparameter optimisation for ensemble learning.
result Ensemble learning with Bayesian optimisation improves anomaly detection accuracy.
CRE discovers interpretable subgroups with heterogeneous treatment effects.
problem Identifying subgroups with notable treatment effect heterogeneity.
method Causal Rule Ensemble (CRE) using an ensemble-of-trees approach.
result CRE offers interpretable decision rules and high stability in subgroup discovery.
Analyzes merging vs. ensembling for multi-study prediction, showing transition point for better performance.
problem Choosing between merging or ensembling multiple studies for prediction.
method Analyzes ridge regression approaches, comparing merging and ensembling methods.
result There is a transition point where ensembling outperforms merging as cross-study heterogeneity increases.
New ensemble methods improve time series forecasting accuracy.
problem Global Forecasting Models (GFM) lack localisation for heterogeneous datasets.
method Ensemble techniques with clustering and varied GFM models.
result Significantly higher accuracy achieved compared to baseline models.
HRF enhances tree diversity in random forests to improve performance.
problem Selection bias and lack of diversity in random forests.
method Introducing heterogeneity during tree construction by assigning lower weights to features used in previous trees.
result HRF outperforms other ensemble methods in accuracy across 52 datasets.
Boosting strategies for merging vs. ensembling studies analyzed.
problem Deciding between merging and ensembling studies for boosting.
method Analytical transition point and bias-variance decomposition for boosting with linear learners.
result Theoretical guidelines for merging vs. ensembling studies.
Optimal ensemble construction improves prediction accuracy for multi-study tasks, especially in pandemic scenarios.
problem Poor out-of-study prediction performance due to heterogeneous datasets.
method Optimal ensemble construction using a two-stage stacking strategy that jointly estimates ensemble weights and study-specific model parameters.
result Our method outperforms multi-study stacking and other standard methods in predicting excess mortality during the pandemic.
Tree ensemble method tackles multi-objective constrained optimization in energy systems.
problem Complex, multi-objective, and constrained optimization problems in energy systems.
method Data-driven tree ensemble approach for black-box problems with heterogeneous variable spaces.
result Competitive performance and sampling efficiency compared to state-of-the-art tools.
Study interbank lending and borrowing dynamics with heterogeneous mean field model.
problem Modeling systemic risk in a network of banks with varying capitalization.
method Developed a mean field type model with coupled diffusions to describe log-capitalization evolution.
result Existence of Nash equilibria in large-scale heterogeneous interbank networks.
Ensemble unsupervised anomaly detection using IRT for hidden ground truth.
problem Challenges in constructing an ensemble from unsupervised anomaly detection methods.
method Use Item Response Theory to compute latent traits and construct an ensemble that downplays noisy methods.
result Demonstrated effectiveness of IRT ensemble on extensive data repository.
New algorithms improve ensemble diversity, leading to more accurate and smaller models.
problem Building accurate predictive models with diverse base predictors.
method Integrates ensemble diversity into a reinforcement learning framework for ensemble selection.
result Diversity-incorporating ensembles are more accurate and smaller in size.
Fed-ensemble improves FL by averaging predictions from multiple models.
problem Improving generalization in federated learning.
method Random permutations to update K models, averaging predictions.
result Predictions from all K models have the same predictive posterior distribution.
The combination of multiple classifiers using ensemble methods is increasingly important for making progress in a variety of difficult prediction problems. We present a comparative analysis of several ensemble methods through two case studies in genomics, namely the prediction of genetic interactions and protein functi…
Bayesian design improves accuracy without extra cost.
problem Nested inference in complex systems limits BED accuracy and efficiency.
method Grouped geometric pooled posterior with EKI formulation.
result Improved accuracy and stable estimators at comparable cost.
We define a random-matrix ensemble given by the infinite-time covariance matrices of Ornstein-Uhlenbeck processes at different temperatures coupled by a Gaussian symmetric matrix. The spectral properties of this ensemble are shown to be in qualitative agreement with some stylized facts of financial markets. Through the…
Gestalt combines two models to improve SQuAD2.0 performance.
problem Improving the accuracy of answering questions in context paragraphs.
method A stacking ensemble of ALBERT and RoBERTa models, combined with a CNN-based meta-model.
result Best ensemble achieved 87.117 EM and 90.306 F1 scores, improving baseline by 0.55% and 0.61% respectively.
Bayesian spatial predictive synthesis improves spatial data predictions.
problem Model misspecification and heterogeneity in spatial data.
method Bayesian ensemble methodology capturing spatially-varying model uncertainty and performance heterogeneity.
result Synthesized predictions outperform standard methods in accuracy and uncertainty quantification.
ROME improves algorithmic fairness by learning latent group structure robustly.
problem Latent subgroup disparities and distribution shifts in machine learning models.
method ROME uses an Expectation-Maximization algorithm for linear models and a neural Mixture-of-Experts for nonlinear settings.
result ROME significantly improves fairness compared to standard methods while maintaining average performance.
The motivation of this work is to improve the performance of standard stacking approaches or ensembles, which are composed of simple, heterogeneous base models, through the integration of the generation and selection stages for regression problems. We propose two extensions to the standard stacking approach. In the fir…
Pattern ensembling fills in missing or inaccurate trajectory data.
problem Incompleteness, missing information, and inaccuracies in geolocation data.
method Probabilistically ensemble similar trajectory patterns from the vicinity.
result Reconstructs missing or unreliable trajectory segments effectively.
New graph-based method selects outlier ensemble components.
problem Poor components negatively affect consensus results in outlier ensembles.
method Mapping rankings to graphs, mining to identify subsets.
result Our method outperforms state-of-the-art techniques.
This work improves Gaussian process regression for large, non-stationary data.
problem Scalability issues and performance degradation for non-stationary data.
method Combines variational free energy approximations with online expectation propagation and local splitting steps.
result Incremental adaptation to locality, heterogeneity, and non-stationarity in training data.
A method for clustering using transfer learning from similar labeled data.
problem Clustering with datasets having different features and labeled data.
method Constructing meta-features to describe structural characteristics of data and transferring them between source and target domains.
result The method is efficient and works under arbitrary feature descriptions of source and target domains with smaller complexity.
HIVE-COTE v1.0 improves time series classification with enhanced usability.
problem Improving time series classification accuracy and usability.
method Presented a walkthrough guide and extensive experimental evaluation of HIVE-COTE v1.0.
result HIVE-COTE v1.0 outperforms three recently proposed algorithms in predictive performance and resource usage.
High-capacity neural network ensembles often benefit more from high-capacity models than from increased diversity.
problem The performance of high-capacity neural network ensembles is often harmed by interventions that promote predictive diversity.
method A large-scale study of nearly 600 neural network classification ensembles, examining various interventions and architectures.
result Discouraging predictive diversity can be benign in large-network ensembles, and higher-capacity models often yield better performance than diverse architectures.
Model predicts short-term Amazon rainforest fires with high accuracy.
problem Accurate short-term forecasting of Amazon rainforest fires is challenging.
method Used Seasonal and Trend decomposition based on Loess combined with multi-month-ahead load forecasting algorithms.
result Proposed decomposition-ensemble models provide more accurate forecasts than other models.
Develops a new classical network ensemble framework.
problem Lack of information-theoretic frameworks for complex networks.
method Optimal trade-off between compressed representation and actual network ensemble.
result Power-law degree distribution is optimal for networks with only expected degrees as constraints.
Gradient-free ensemble learns sector forecasts from diverse models.
problem Predicting sector returns in a volatile market.
method Dynamic model combination using out-of-sample R-squared.
result Ensemble outperforms individual models in sector rotation.
Paper proves global convergence of NCELM model.
problem Ensuring global convergence of NCELM model.
method Two-stage process: random base learners and NCL penalty term updates.
result Global convergence of NCELM proved using Banach theorem.
Paper introduces Hedged Bandits for ensemble learning with performance guarantees.
problem Lack of performance guarantees in existing ensemble learning methods.
method Hedged Bandits method providing asymptotic and short run performance guarantees.
result Hedged Bandits method outperforms existing ensemble learning methods.
Bayesian tree ensemble model for estimating treatment effects in high-dimensional survival data.
problem Estimating heterogeneous treatment effects in censored survival data with many covariates.
method Developed a Bayesian tree ensemble model with a horseshoe prior for adaptive shrinkage.
result Accurately estimates treatment effects in high-dimensional covariate spaces and non-linear functions.
HESCA combines simpler models from different families to outperform individual models.
problem Choosing the best classifier for a classification problem.
method Building ensembles of simpler models from different families of classifiers.
result HESCA significantly outperforms individual models and represents a strong benchmark.
A diverse system combines CNNs and meta-nets for handwritten digit recognition.
problem Handwritten digit recognition using diverse classification hypotheses.
method Generate diverse classification hypotheses using CNNs and other techniques, then combine them with Meta-Nets.
result Achieved state-of-the-art performance in handwritten digit recognition.
Proposes an interpretable machine learning framework for multi-arm HTE estimation.
problem Challenges in estimating heterogeneous treatment effects in multi-arm settings.
method Rule-based ensemble approach for HTE estimation in multi-arm trials.
result Achieved lower bias and higher estimation accuracy compared to existing methods.
Federated learning improves by training central model with client model outputs.
problem Direct averaging of client models is limited in FL.
method Ensemble distillation for model fusion.
result Central model trained faster with fewer communication rounds.
Study forecasts monthly electricity demand using pattern similarity-based methods.
problem Forecasting monthly electricity demand accurately.
method Pattern similarity-based forecasting methods (PSFMs) including k-NN, fuzzy, kernel regression, and GRNN.
result Ensemble models outperform individual PSFMs in forecasting accuracy.
USNRT uses tree-structured learning to improve uncertainty quantification of variance networks.
problem Improving uncertainty quantification of variance networks.
method Tree-structured local neural network model that partitions feature space into regions for training region-specific neural networks to predict mean and variance.
result USNRT shows superior performance in estimating uncertainty with variances on UCI datasets compared to recent methods.
Bayesian optimization improves with transfer learning for aircraft design.
problem Cold start problem in Bayesian optimization for aircraft design.
method Ensemble of surrogate models using transfer learning in a constrained Bayesian optimization framework.
result Significant improvement in convergence and prediction accuracy.
Funnelling improves cross-lingual text classification accuracy.
problem Classifying documents in multiple languages more accurately than individual language classifiers.
method A two-tier classification system using posterior probabilities from language-dependent classifiers.
result Funnelling significantly outperforms state-of-the-art baselines in multilingual text classification.
Boost-R uses gradient boosted trees for analyzing recurrence data.
problem Analyzing recurrence data with static and dynamic features.
method Gradient boosted additive trees with time-dependent functions.
result Estimates the cumulative intensity function of recurrent event processes.
A method for inferring motility models and heterogeneity from particle trajectories.
problem Understanding motility patterns from discrete trajectory data of biological agents.
method Maximum likelihood approach for second-order Langevin models with population heterogeneity.
result The proposed method outperforms alternative approaches for short trajectories.
Study develops ensemble machine learning framework for predicting groundwater heavy metal pollution.
problem Statistical complexity and spatial heterogeneity of heavy metal contamination in groundwater.
method Nested cross-validated ensemble machine learning with response transformations (raw, log, Gaussian copula).
result Copula-based models with DBSCAN clustering diagnostics provide the most reliable and interpretable assessments of groundwater contamination.
LESS combines local predictors for subsets to learn from heterogeneous input-output pairs.
problem Learning from heterogeneous input-output pairs in populations with varied behavior.
method LESS algorithm: generates subsets, trains local predictors, combines them.
result LESS is highly competitive compared to state-of-the-art methods.
Combines Xgboost and transductive SVM for semi-supervised learning.
problem Improving semi-supervised learning performance with heterogeneous tabular data.
method Proposes an optimization-based ensemble method to adaptively combine Xgboost and transductive SVM.
result Significantly improves classification accuracy over state-of-the-art methods.