Random forests reduce bias and variance, especially in low SNR settings.
problem Reducing bias and variance in machine learning models, particularly in low SNR scenarios.
method Empirical study of random forests and bagging ensembles, focusing on the importance of mtry tuning. result Random forests reduce both bias and variance, outperforming bagging ensembles in high SNR settings.
Models of bags of words typically assume topic mixing so that the words in a single bag come from a limited number of topics. We show here that many sets of bag of words exhibit a very different pattern of variation than the patterns that are efficiently captured by topic mixing. In many cases, from one bag of words to…
Improved TSC with BOSS and SP techniques.
problem Comparing BOP and BOSS for time series classification.
method Deconstructed and measured components of BOP and BOSS, adapted CV techniques.
result SP with BOSS significantly more accurate than benchmarks.
Proposes a new MIL formulation using infinitely many shapelets.
problem Weakness of single shapelet classifiers and lack of theoretical guarantee for multiple shapelets.
method Formulates a new MIL approach with infinitely many shapelets and provides an efficient algorithm.
result Empirical study shows effectiveness in MIL and Shapelet Learning.
Feature learning and deep learning have drawn great attention in recent years as a way of transforming input data into more effective representations using learning algorithms. Such interest has grown in the area of music information retrieval (MIR) as well, particularly in music audio classification tasks such as auto…
New formulation of MIL using shapelets for better classifier of bags.
problem Finding a good classifier of bags based on shapelets.
method Formulation using all possible shapelets, reduced to DC programs, and heuristic options.
result Richer class of classifiers with theoretical justification and empirical validation.
Many objects in the real world are difficult to describe by a single numerical vector of a fixed length, whereas describing them by a set of vectors is more natural. Therefore, Multiple instance learning (MIL) techniques have been constantly gaining on importance throughout last years. MIL formalism represents each obj…
Paper discovers shifting patterns in sequence classification and proposes a method to improve performance.
problem Discriminative patterns in sequential data are time-dependent and degrade traditional classification methods.
method Proposes a novel sequence classification method using multi-instance learning and LSTM models to detect and model shifting patterns.
result Demonstrates superior sequence classification performance and detection of shifting patterns in cropland mapping and affective state recognition.
Paper proposes a method to generate instance labels from weakly supervised data.
problem Weakly supervised instance labeling in medical image analysis.
method Uses multiple instance learning (MIL) and knowledge distillation to generate instance-level predictions.
result Significantly outperforms state-of-the-art MIL methods in instance-level prediction.
A new approach to MI learning using bag-to-class divergence.
problem Sparse MI training sets and difficulty in classifying bags.
method Introducing bag-to-class divergence to MI learning, emphasizing hierarchical random vectors.
result Bag-to-class divergence is a more effective classifier for MI learning.
Bagging stabilizes models without distributional assumptions.
problem Stability of machine learning models without distributional assumptions.
method Derives a finite-sample guarantee on bagging stability for any model.
result Guarantee applies to many bagging variants and is optimal.
Feature bagging improves stability through random feature subsampling.
problem Improving the stability of ensemble learning methods.
method Introducing feature instability (FI) and analyzing feature bagging in parametric and model-free settings.
result Feature bagging provides stronger stability than non-bagged methods, especially with aggressive subsampling.
The paper analyzes bagging in overparameterized learning, deriving risk properties and optimal subsample sizes.
problem Characterizing the risk of bagged predictors in overparameterized settings.
method General strategy using classical results on simple random sampling, specialized for ridge and ridgeless predictors.
result Derives exact asymptotic risk of bagged ridge and ridgeless predictors under various conditions.
Multi-instance learning (MIL) has a wide range of applications due to its distinctive characteristics. Although many state-of-the-art algorithms have achieved decent performances, a plurality of existing methods solve the problem only in instance level rather than excavating relations among bags. In this paper, we prop…
Bagging reduces MSE for unstable estimators like trees, but only if the distribution has high kurtosis.
problem Theoretical properties of bagging applied to statistical estimators.
method Theoretical analysis of MSE reduction for bagged estimators, focusing on variance estimation.
result Bagging reduces MSE for unstable estimators like trees, but only if the distribution has high kurtosis.
Accelerometer measurements are the prime type of sensor information most think of when seeking to measure physical activity. On the market, there are many fitness measuring devices which aim to track calories burned and steps counted through the use of accelerometers. These measurements, though good enough for the aver…
Multiple instance learning (MIL) is concerned with learning from sets (bags) of objects (instances), where the individual instance labels are ambiguous. In this setting, supervised learning cannot be applied directly. Often, specialized MIL methods learn by making additional assumptions about the relationship of the ba…
The paper compares aggregated data labels in curated and random bags for machine learning models.
problem Protecting user privacy in machine learning systems with aggregated data.
method Examined curated and random bags for training machine learning models and compared their performance.
result Gradient-based learning can be performed on aggregated data without performance degradation.
Bagging is a device intended for reducing the prediction error of learning algorithms. In its simplest form, bagging draws bootstrap samples from the training sample, applies the learning algorithm to each bootstrap sample, and then averages the resulting prediction rules. We extend the definition of bagging from stati…
A new large-scale tabular benchmark for Learning from Label Proportions.
problem Lack of a large-scale open benchmark for tabular Learning from Label Proportions.
method Proposed LLP-Bench, a suite of 70 datasets (62 feature bag and 8 random bag) from real-world tabular data.
result Demonstrated the effectiveness of 9 SOTA and popular tabular LLP techniques on 62 feature bag datasets.
Paper introduces r-DEP classifier for binary classification tasks.
problem No natural ordering for feature patterns in practical situations.
method Introduces reduced dilation-erosion (r-DEP) classifier using multi-valued mathematical morphology.
result r-DEP classifiers outperform traditional SVCs in balanced accuracy.
Under-bagging k-NN improves performance on imbalanced classification.
problem Imbalanced classification problems where one class is significantly underrepresented.
method Proposes an under-bagging k-NN ensemble learning algorithm, analyzing convergence rates and efficiency. result Achieves optimal convergence rates under mild assumptions and reduces sub-sample size and k for highly imbalanced data. Automatically classifying the tissues types of Region of Interest (ROI) in medical imaging has been an important application in Computer-Aided Diagnosis (CAD), such as classification of breast parenchymal tissue in the mammogram, classify lung disease patterns in High-Resolution Computed Tomography (HRCT) etc. Recently…
Bagging improves sparse regression performance, especially with reduced sampling ratios.
problem Improving sparse regression performance in low measurement scenarios.
method Generalized Bagging with various bootstrap sampling ratios.
result Bagging outperforms L1 minimization and Bolasso in challenging sparse regression cases.
Improved Bayesian uncertainty quantification using variational bagging.
problem Inefficient and underestimating uncertainty in mean-field variational Bayes.
method Integrates bagging with variational Bayes for improved inference.
result Bagged variational posterior provides proper uncertainty quantification.
Graph neural networks improve MIL performance without losing interpretability.
problem Learning bag-level labels from bags of instances.
method Proposed a GNN-based algorithm treating bags as graphs to learn embeddings and predict labels.
result Achieves state-of-the-art performance on MIL datasets.
Develops RL algorithm for non-Markovian, non-stationary reward streams.
problem Maximizing rewards from non-Markovian, non-stationary reward streams.
method Uses causal DAG to construct Markov states, solves periodic MDP.
result Optimal state construction maximizes discounted rewards.
In the supervised learning setting termed Multiple-Instance Learning (MIL), the examples are bags of instances, and the bag label is a function of the labels of its instances. Typically, this function is the Boolean OR. The learner observes a sample of bags and the bag labels, but not the instance labels that determine…
A new method for MIR in remote sensing without assuming a prime instance per bag.
problem Multiple Instance Regression in remote sensing with high variability.
method Treats each bag as a set of instances and learns to map each bag to its unique label using all instances.
result Outperforms previous state-of-the-art on three real-world datasets.
In multiple instance learning, objects are sets (bags) of feature vectors (instances) rather than individual feature vectors. In this paper we address the problem of how these bags can best be represented. Two standard approaches are to use (dis)similarities between bags and prototype bags, or between bags and prototyp…
Bagging stabilizes linear interpolators, improving their generalization performance.
problem Unstable linear interpolators fail on noisy data.
method Introduced multiplier-bootstrap-based bagged least square estimator.
result Bagging effectively mitigates variance, leading to bounded prediction risk.
Algorithm learns from label proportions in unlabeled bags.
problem Learning from unlabeled bags with known label proportions.
method Differentiable loss functions for deep neural networks.
result Deep neural networks can accurately classify images from unlabeled bags.
New method improves weakly-supervised action localization.
problem Locating action segments in videos with limited labels.
method Explicitly models key instance assignment as hidden variable using EM framework.
result Achieves state-of-the-art performance on THUMOS14 and ActivityNet1.2 benchmarks.
BRDAD uses bagging and regularization to improve anomaly detection without labeled data.
problem Anomaly detection in unlabeled data with sensitivity to k-nearest neighbors. method Bagged regularized k-distances (BRDAD) for anomaly detection, converting to convex optimization. result BRDAD addresses sensitivity to hyperparameter choice and improves performance on large datasets.
Proposes a new method for MIR using kernel mean embeddings.
problem Multiple instance regression (MIR) where bags contain multiple instances with a single label.
method Computes kernel mean embeddings of predicted label distributions and learns a regressor from these embeddings.
result Better results than baseline instance-MIR across all datasets, state-of-the-art on two.
Paper addresses instance-level label prediction in MIL, improving performance compared to existing methods.
problem Lack of instance-level label prediction in existing MIL approaches restricts their applicability.
method Proposes a novel algorithm with instance-level loss function, unbiasedly and consistently estimated.
result Empirically validated superior instance-level and bag-level performance compared to state-of-the-art methods.
Extends multi-instance learning to nested bags for better interpretation and classification.
problem Nested bags in multi-instance learning for diverse applications.
method Bag-layer neural network layer for aggregating bags of inputs.
result Bag-layer neural network can learn and interpret complex functions over sets of sets.
New framework learns labels at both bag and graph levels.
problem Learning multi-label classifiers from multi-graph bags.
method Designing scoring functions and rank-loss objective for graph and bag levels; developing sub-gradient descent algorithm.
result Superior performance over state-of-the-art algorithms.
This paper compares two loss functions for learning from aggregated responses and introduces an interpolating estimator.
problem Learning from aggregated responses in privacy-sensitive settings.
method Investigates bag-level and instance-level loss functions, and introduces an interpolating estimator.
result Instance-level loss can be seen as a regularized form of bag-level loss, leading to improved estimators.
Improves key instance detection in MIL models by using neural network inversion with sparseness constraint.
problem Limited key instance detection performance in attention-based deep MIL models due to skewed attention scores.
method Sparse network inversion with a sparseness constraint incorporated into neural network inversion, solved by proximal gradient method.
result Significantly improved key instance detection performance while maintaining bag-level prediction performance.
WildWood improves Random Forest predictions using bootstrap out-of-bag samples.
problem Improving Random Forest predictions for supervised learning.
method Uses bootstrap out-of-bag samples to compute improved predictions by aggregating all possible subtrees with exponential weights.
result WildWood produces faster and more competitive predictions compared to other ensemble methods.
Labeling data for classification requires significant human effort. To reduce labeling cost, instead of labeling every instance, a group of instances (bag) is labeled by a single bag label. Computer algorithms are then used to infer the label for each instance in a bag, a process referred to as instance annotation. Thi…
A new SVM method using bagging and importance sampling.
problem Solving SVM problems efficiently in large databases.
method Bagging and importance sampling applied to SVM.
result Faster solution of SVM problem with minimal loss in prediction error.
Paper proposes a neural network approach for MIL, achieving comparable and sometimes superior performance.
problem Learning bag labels from unlabeled bags of instances.
method Neural networks with attention mechanism for permutation-invariant aggregation.
result Attention-based aggregation improves performance on various MIL datasets.
We investigate the mapping class group of an orientable ω-bounded surface. Such a surface splits, by Nyikos's Bagpipe Theorem, into a union of a bag (a compact surface with boundary) and finitely many long pipes. The subgroup consisting of classes of homeomorphisms fixing the boundary of the bag is a normal subgroup …
Boosting weak learners to strong ones from aggregate labels is possible for LLP but not for MIL.
problem Boosting weak learners to strong ones from aggregate labels in learning from label proportions (LLP).
method Using a weak learner on large enough bags to obtain a strong learner for small bags in polynomial time.
result Boosting is possible for LLP but not for MIL.
SFBoW provides sentence embeddings with predefined dimensions.
problem Sentence embeddings problem at document-level.
method Refinement of Fuzzy Bag-of-Words, predefined dimension.
result Competitive performances in Semantic Textual Similarity benchmarks.
Learning from Label Proportions (LLP) is a learning setting, where the training data is provided in groups, or "bags", and only the proportion of each class in each bag is known. The task is to learn a model to predict the class labels of the individual instances. LLP has broad applications in political science, market…