Improved set prediction model using multiset-equivariant operations and approximate implicit differentiation.
problem Existing set prediction models struggle with multisets and cannot represent certain functions.
method Introduced multiset-equivariance, improved DSPN with approximate implicit differentiation, and applied to CLEVR object property prediction.
result Significantly improved object property prediction on CLEVR dataset.
New model predicts sets from feature vectors without discontinuity issues.
problem Discontinuity issues in predicting sets from feature vectors.
method General model that respects set structure, auto-encodes point sets, predicts bounding boxes, and attributes.
result Model successfully predicts sets from a single feature vector without discontinuity.
Study improves top-k set prediction with low cardinality.
problem Improving top-k set prediction accuracy with low cardinality.
method Introduces new target loss function and surrogate losses.
result Demonstrates effectiveness of cardinality-aware algorithms.
A new convex loss function optimizes set predictions with balanced size and coverage.
problem Optimizing set predictions with balanced size and coverage.
method Proposes a convex loss function using Choquet integrals for nondecreasing subset-valued functions.
result Optimal trade-offs between conditional probabilistic coverage and set size.
FSPool improves set prediction accuracy and convergence.
problem Set prediction models struggle with simple datasets due to the responsibility problem.
method Featurewise sort pooling to construct a permutation-equivariant auto-encoder.
result FSPool improves reconstructions and representations on various datasets.
Paper proposes an alternative to set losses for predicting unordered variables without imposing structure.
problem Learning unordered variables with unknown interrelations without imposing structure.
method Viewing set prediction as conditional density estimation and using deep energy-based models with gradient-guided sampling.
result Empirically demonstrates capability to learn multi-modal densities and produce different plausible predictions.
Paper presents unsupervised calibration for split conformal classification.
problem Inconvenient requirement of labeled calibration samples.
method Uses unsupervised calibration samples alongside supervised training samples.
result Achieves comparable performance to supervised calibration methods.
Meta-learning reduces set prediction size in conformal prediction for few-shot calibration.
problem Inefficient set prediction in conformal prediction for limited training data.
method Meta-learning approach using cross-validation-based conformal prediction.
result Meta-learning scheme reduces set prediction size and preserves formal guarantees.
New model predicts fine-grained types for high-multiplicity entities.
problem Fine-grained entity typing with high type multiplicity.
method Set-prediction approach to high-multiplicity fine-grained typing.
result Model outperforms baselines on Wikipedia-based corpus.
Develops deep neural network techniques for sets as input and output.
problem Bottlenecks in set representation and discontinuity issues in set prediction.
method Techniques for set representation and prediction, addressing unordered nature and relations.
result Improvements in set prediction and representation across various experiments.
Learning with non-modular losses is an important problem when sets of predictions are made simultaneously. The main tools for constructing convex surrogate loss functions for set prediction are margin rescaling and slack rescaling. In this work, we show that these strategies lead to tight convex surrogates iff the unde…
The paper tackles optimal set prediction in multi-class classification.
problem Finding the best set of classes for uncertain predictions.
method Formalized decision-theoretic framework, quantified uncertainty, Bayes-optimal prediction algorithms.
result Efficient algorithms for optimal set prediction in multi-class classification.
Deep Auto-Set learns sets of activities from wearable sensor data.
problem Recognizing multiple activities simultaneously from sensor data.
method Deep auto-encoder-set network for set prediction.
result Significant improvement over baseline models in HAR datasets.
Structured prediction is used in areas such as computer vision and natural language processing to predict structured outputs such as segmentations or parse trees. In these settings, prediction is performed by MAP inference or, equivalently, by solving an integer linear program. Because of the complex scoring functions …
Kernel Induced Random Survival Forests (KIRSF) is a statistical learning algorithm which aims to improve prediction accuracy for survival data. As in Random Survival Forests (RSF), Cumulative Hazard Function is predicted for each individual in the test set. Prediction error is estimated using Harrell's concordance inde…
New method improves ensemble diversity and generalization.
problem Ensemble diversity does not guarantee practical generalization.
method Introduced a new diversity metric and training method for extrapolating differently on local data patches.
result Improves generalization and diversity in practical settings, especially under data limits and covariate shift.
Two new methods improve efficiency of conformal predictive systems.
problem Efficiency of conformal predictive systems in regression problems.
method Split conformal predictive systems and cross-conformal predictive systems.
result Cross-conformal predictive systems are more efficient but not guaranteed valid.
Unified framework for generalized Venn and Venn-Abers calibration for reliable prediction.
problem Asymptotic guarantees of popular distribution-free methods in model calibration.
method Unified framework extending Vovk's approach to generic loss functions, transforming predictors into set-valued predictions.
result Finite-sample set predictions shrink to a single conditionally calibrated prediction, capturing epistemic uncertainty.
Proposes real-time risk monitoring for machine learning systems under unknown shifts.
problem Dynamic distribution shifts challenge real-world machine learning systems' risk assurances.
method Sequential hypothesis testing with 'testing by betting' to detect risk violations.
result Effective real-time risk monitoring under various unknown shifts.
Paper improves prediction sets for distribution shifts without labels.
problem Improving prediction sets effectiveness in the presence of distribution shifts.
method Develops ECP and EACP methods to adjust score function based on model uncertainty.
result Consistent improvement over existing baselines and nearly matches fully supervised methods.
L-ARC improves model fairness by localizing risk guarantees.
problem Improving model fairness in tasks like image segmentation and wireless networks.
method Localized Adaptive Risk Control (L-ARC) updates a threshold function in RKHS to target localized statistical risk guarantees.
result L-ARC produces prediction sets with improved fairness across different data subpopulations.
Wide neural networks simplify to linear models under gradient descent.
problem Understanding the training dynamics of deep neural networks.
method Analyzing wide neural networks in the infinite width limit and showing they evolve as linear models.
result Gradient-based training of wide neural networks results in predictions from a Gaussian process with a specific kernel.
Framework uses dropout to efficiently explore Rashomon set for multiplicity estimation.
problem Efficiently measuring and mitigating conflicting model outputs in classification tasks.
method Dropout-based exploration of Rashomon set for multiplicity estimation.
result Framework outperforms baselines in multiplicity metric estimation with significant runtime speedup.
FLOPART solves peak detection by creating accurate train and test set predictions.
problem Correctly detecting peaks in sequential data.
method Dynamic programming changepoint algorithm with zero train label errors.
result FLOPART provides highly accurate predictions on both train and test sets.
New method optimizes Lipschitz interpolation for robust predictions.
problem Lack of robust methods for estimating Lipschitz constants from data.
method Optimizes parameters of presupposed metrics to minimize prediction errors.
result Competitive approach for nonparametric black-box learning.
Algorithm identifies and corrects training set bugs using trusted items.
problem Training set flaws affect machine learning models.
method Algorithm uses a combination of combinatorial and continuous optimization to identify and correct bugs in the training set.
result The algorithm can effectively identify and suggest changes to correct training set bugs.
New approach tackles decision-making under predictions that shape outcomes.
problem Challenges in learning optimal decision rules when predictions influence outcomes.
method Introduces performative omniprediction, a predictor that encodes optimal decision rules for multiple objectives.
result Efficient performative omnipredictors exist under a natural restriction of outcome performativity.
The paper presents a method to compute trusted confidence bounds for LECs in CPS.
problem Non-transparent predictions of LECs make CPS safety challenging.
method Inductive Conformal Prediction (ICP) and Triplet Network architecture.
result Efficient real-time computation of trusted confidence bounds.
Learning from Label Proportions (LLP) is a learning setting, where the training data is provided in groups, or "bags", and only the proportion of each class in each bag is known. The task is to learn a model to predict the class labels of the individual instances. LLP has broad applications in political science, market…
Deep Learning model diagnoses four lymphoma categories with high accuracy.
problem Automated detection of lymphoma categories using digital pathology images.
method Convolutional neural network algorithm trained on 128 cases of lymph node images.
result Excellent diagnostic accuracy (95% image-by-image, 10% set-by-set).
GraphDETR detects subgraphs in large graphs using deep learning.
problem Detecting subgraphs in large graphs efficiently and accurately.
method Formulates subgraph detection as a set prediction problem using GraphDETR, a deep learning framework.
result GraphDETR can detect diverse patterns in large graphs, achieving strong performance on molecular functional group detection.
Paper proposes a cost-sensitive conformal training method with provably controllable learning bounds.
problem Uncertainty quantification and learning bounds in conformal prediction.
method Cost-sensitive conformal training algorithm that minimizes the expected size of prediction sets using rank weighting.
result Theoretical analysis shows tightness between weighted objective and expected size of conformal prediction sets.
This paper improves deep learning classifiers by integrating conformal prediction during training.
problem High-stake AI applications require reliable uncertainty estimates for safe deployment.
method Integrates conformal prediction (CP) during training of deep learning models.
result Reduces inefficiency and allows more control over confidence sets.
Paper develops an algorithm with PAC guarantees for detecting alien categories.
problem Detecting alien categories not seen in training data reliably.
method Develops an algorithm with PAC-style guarantees for alien detection under known upper bounds on alien fraction.
result Empirical results show the algorithm's effectiveness in detecting aliens.
Consistent supervised learning with missing values is possible using imputation or specialized models.
problem Predicting with missing values in both training and testing data.
method Two approaches: imputing with a constant and using a predictor for complete observations through multiple imputation. Decision trees can handle missing values naturally.
result Imputing with a constant can be consistent when missing values are not informative.
Transformers learn to solve various tasks without explicit design.
problem Training models with minimal inductive bias.
method Meta-learning approach to train general-purpose in-context learning algorithms.
result Transformers can be meta-trained to solve a wide range of tasks.
Framework learns inter-electronic potential for molecular dynamics.
problem Predicting time-dependent Hartree-Fock dynamics from electron density.
method Developed three models using four-index tensors, preserving symmetries.
result Model with eight-fold symmetry performs best across metrics.
A new method for combining multiple data views in supervised learning.
problem Combining multiple data views in supervised learning, especially in biology and medicine.
method Cooperative learning combines squared error loss with an agreement penalty to encourage predictions from different data views to agree.
result Cooperative learning achieves higher predictive accuracy on simulated and real multiomics data.
Bayesian CNNs with many channels are equivalent to Gaussian processes.
problem Understanding the behavior of deep convolutional networks in the infinite channel limit.
method Deriving an equivalence between multi-layer convolutional neural networks and Gaussian processes, introducing a Monte Carlo method for estimation.
result The GPs corresponding to CNNs with and without weight sharing are identical in the infinite channel limit.
COMA combines prediction sets from multiple models for online, adaptive prediction.
problem Combining multiple prediction models with uncertainty guarantees.
method Online model aggregation using weighted voting of conformal prediction sets.
result COMA retains coverage guarantees under negative correlation assumptions.
This work provides bounds on the performance of prediction models in the predict-then-optimize framework.
problem Generalizing the performance of prediction models in the predict-then-optimize framework with the SPO loss function.
method Deriving generalization bounds using the Natarajan dimension and exploiting the strength property of the feasible region.
result Improved generalization bounds for the SPO loss function, scaling logarithmically in the number of extreme points and linearly in the decision dimension.