Simplifies complex regression model interpretation.
problem Interpreting complex machine learning models.
method A method for grounding interpretations in actual learning examples.
result Validated on academic and industrial regression tasks.
Self-supervised methods learn from noisy data alone, useful for imaging problems.
problem Inferring signals from noisy and incomplete observations.
method Learning a solver from measurement data alone, without ground-truth references.
result Self-supervised methods can learn meaningful estimates from noisy data.
GCNs' performance linked to feature, graph, and ground truth alignment.
problem Improving GCNs' classification performance.
method Subspace alignment measure (SAM) based on Frobenius norm of chordal distances.
result SAM quantifies the alignment between features, graph, and ground truth.
Wasserstein GANs use different p-metrics to improve model performance.
problem Improving stability and performance of Wasserstein GANs.
method Introduce (q,p)-Wasserstein GANs using various p-metrics. result Different p-metrics can notably improve GAN performance. Unsupervised clustering can reproduce categorization systems if features and metrics are correctly selected.
problem Reproducing expert-provided categorization systems using unsupervised clustering.
method Investigated using toy datasets and real-world fund categorization. Used appropriate feature selection and a supervised Random Forest-based distance metric.
result Unsupervised clustering can reproduce ground truth classes if features and metrics are correctly selected.
SATNet solves the Symbol Grounding Problem, enabling self-supervised learning.
problem Mapping visual inputs to symbolic variables without explicit supervision.
method Self-supervised pre-training pipeline and proofreading method.
result SATNet achieves full accuracy with no label leakage, surpassing state-of-the-art.
Deep learning model estimates mutual information with low bias and variance.
problem Estimating mutual information between continuous variables.
method Supervised deep learning approach using Linfoot informational correlation as labels.
result Lower bias and variance compared to other methods.
Using authors's methods of 1980, 1981, some explicit finite sets of number fields containing ground fields of arithmetic hyperbolic reflection groups are defined, and good bounds of their degrees (over Q) are obtained. For example, degree of the ground field of any arithmetic hyperbolic reflection group in dimension at…
Logistic regression can handle noisy labels effectively when labels are imperfectly assigned by multiple experts.
problem Label noise in supervised classification due to manual labelling by multiple experts.
method Using approximate posterior probabilities of class membership from multiple experts to train logistic regression models.
result Logistic regression can be robust to label noise when classification difficulty is the only source of errors.
New bias found in cluster validity indices when ground truth distribution changes.
problem Bias in external cluster validity indices when ground truth distribution changes.
method Identified new type of bias (GT bias) and studied its empirical and theoretical implications.
result GT bias can change the bias status of cluster validity indices.
New algorithm improves sparse-view tomography without needing ground-truth data.
problem Poor image reconstructions with sparse projections and non-uniform sensors.
method Unsupervised deep learning with CNN and STN modules.
result Significantly outperforms filtered backprojection in sparse-view scenarios.
Study heat profiles and eigenfunctions using Brownian motion.
problem Investigate heat profiles and eigenfunctions of Laplace equations.
method Probabilistic tools based on Brownian motion and Feynman-Kac formulae.
result Supremum norm bounds for ground state Dirichlet eigenfunctions and comparison of maximum temperatures.
New pseudo-Hermitian models from non-semisimple TQFTs.
problem Constructing exactly solvable pseudo-Hermitian spin Hamiltonians.
method Identifying ground states on surfaces using non-semisimple TQFTs.
result Ground states depend only on spatial topology and can be assigned by non-semisimple TQFTs.
A new learning method for prosthetic arms without explicit rewards.
problem Learning a prosthetic arm to interact with users without explicit reward signals.
method Interaction-Grounded Learning, observing multidimensional context and feedback vectors, discovering latent reward signal.
result The algorithm can discover a latent reward signal and ground its policies for successful interaction.
A new method learns meaningful distances between samples using optimal transport.
problem Learning meaningful distances between samples in datasets without labeled data.
method Computes OT distances between samples and features using singular vectors of a function mapping ground metrics to OT distances.
result Wasserstein Singular Vectors provide a scalable solution for unsupervised ground metric learning.
Deep learning predicts mass from images with sparse ground truth.
problem Accurately estimating mass from images with limited ground truth.
method Semi-supervised deep learning with gradient aggregation and sparse ground truth.
result Deep neural network accurately predicts mass from images.
Crowdsourcing infers ground truth from multiple annotators, verified for supervised learning.
problem Obtaining universally valid ground truth for supervised learning is challenging and costly.
method Gather multiple annotations from diverse individuals, verify and aggregate for training classifiers.
result Inferred ground truth improves classifier performance in sensitive tasks like mitosis detection.
New method uses kernel methods to approximate ground states of quantum Hamiltonians efficiently.
problem Approximating ground states of quantum Hamiltonians using neural networks is computationally expensive.
method Introduces a statistical learning approach using kernel methods to make optimization trivial.
result Ground state properties of arbitrary gapped quantum Hamiltonians can be reached with polynomial resources.
Analogator learns to make analogies by example.
problem Developing a computer program to learn analogies.
method Connectionist approach using a recurrent network architecture trained to divide scenes into figure and ground.
result Analogator can make new analogies between novel situations.
Optimizes ground metric on graphs for evolving density models.
problem Optimizing ground metric for evolving density models.
method Adaptive ground metric learning constrained to geodesic distances on graphs.
result Efficiently learned geodesic distances align with observed density evolution.
Tracr compiles programs into transformer models for interpretability.
problem Uncertainty in understanding transformer model outputs due to unknown learned programs.
method Tracr compiles human-readable programs into known structure transformer models.
result Known structure of Tracr-compiled models serves as ground-truth for interpretability.
New method trains deep denoisers without ground truth data.
problem Training deep denoisers with high-quality ground truth images is often impractical.
method Uses Stein's unbiased risk estimator (SURE) for training deep neural networks with only noisy images.
result Trained deep denoisers perform similarly to those trained with ground truth images.
Transportation distances have been used for more than a decade now in machine learning to compare histograms of features. They have one parameter: the ground metric, which can be any metric between the features themselves. As is the case for all parameterized distances, transportation distances can only prove useful in…
Paper proposes a method to recover accurate labels from partially valid data in multi-label learning.
problem Tackles noisy supervision in multi-label learning with partially valid labels.
method Develops a two-stage method that estimates label enrichment and ground-truth confidences.
result Demonstrates improved performance over state-of-the-art PML methods.
End-to-end framework learns from imperfect annotations directly.
problem Training machine learning models on imperfect human annotations.
method End-to-end framework merging aggregation with model training and modeling annotator competencies.
result Accuracy gains of up to 25% over state-of-the-art annotation aggregation methods.
Study validates metrics for offline MBO using diffusion models.
problem Evaluate metrics for offline MBO without ground truth oracle.
method Propose and quantify validation metrics over datasets.
result Identify most effective validation metrics.
TemperatureGAN generates hourly atmospheric temperature data with high fidelity.
problem Generating accurate hourly atmospheric temperature data for climate risk assessment.
method Generative Adversarial Network (GAN) conditioned on months, locations, and time periods.
result TemperatureGAN produces high-fidelity hourly atmospheric temperature data with good spatial and temporal consistency.
CNNs improve particle identification in ground-based gamma-ray astronomy.
problem Identifying particles in gamma-ray astronomy images.
method Used convolutional neural networks (CNNs) with PyTorch and TensorFlow.
result Improved accuracy in identifying gamma-rays and background particles.
Model learns image-word associations from captions using contrastive learning.
problem Phrase grounding, associating image regions to caption words.
method Optimizing word-region attention to maximize mutual information, using language model guided word substitutions for negatives.
result Model achieves 76.7% accuracy on Flickr30K Entities benchmark, a 5.7% gain from weak supervision.
This research detects and identifies human-made objects in 3D point clouds using novel methods.
problem Detect and identify human-made objects in 3D point clouds.
method Ground filtering, local information extraction, clustering using Marked Point Fields (MPFs) and Hessian matrix.
result The proposed method outperforms previous techniques in detecting human-made objects.
SAMPLR optimizes for ground truth in aleatoric parameters to avoid curriculum-induced covariate shift.
problem Curriculum learning shifts training distribution, leading to suboptimal policies in aleatoric settings.
method SAMPLR optimizes ground-truth utility function, avoiding curriculum-induced covariate shift.
result SAMPLR preserves optimality under ground-truth distribution, promoting robustness across various environments.
This paper analyzes the difficulty of unsupervised domain adaptation using information theory.
problem The challenge of unsupervised domain adaptation under covariate shift.
method Formulates the problem using a distribution π in the ground-truth triples (p, q, f), defines optimal learner performance, and introduces PTLU for quantifying difficulty.
result Characterizes the optimal learner and introduces PTLU as a measure of UDA difficulty.
A new algorithm improves efficiency in selecting examples for deep learning.
problem Efficiently choosing multiple examples to mark up for deep learning on large datasets.
method Large BatchBALD algorithm, approximating BatchBALD with reduced computational complexity.
result Comparable quality in selection while significantly reducing computation time, especially for large batches.
VCML learns concepts and metaconcepts from images and questions.
problem Learning concepts and metaconcepts from visual data.
method Bidirectional connection between visual concepts and metaconcepts.
result VCML can generalize from limited data and noisy inputs.
New method tackles semi-unsupervised learning with ultra-sparse labels.
problem Learning from datasets where some classes have no labelled examples.
method Combining clustering and semi-supervised learning with deep generative models.
result Effective learning possible even when half of ground truth classes are unlabelled.
Deep learning improves mammography assessment with high accuracy.
problem Challenges in mammography assessment due to noise, resolution, and lack of ground truths.
method Proposes a classification approach using multi-scale deep tissue classifiers.
result Highest AUC of 0.9 achieved in classifying suspicious tissue patches.
Ground-A-Video edits videos without training, preserving intended changes.
problem Complex multi-attribute video editing with omitted or wrong changes.
method Grounding-guided video-to-video translation with Cross-Frame Gated Attention.
result Zero-shot multi-attribute video editing with improved accuracy and frame consistency.
CPPO learns policies from partial offline data in MDPs with structural assumptions.
problem Offline Reinforcement Learning with partial coverage assumption.
method Constrained Pessimistic Policy Optimization (CPPO) using a function class and model class constraint.
result CPPO achieves PAC guarantee with partial coverage, learning competitive policies.
Deep learning for Venus images uses high-res hyperspectral data to simulate ground truth.
problem Lack of accurate ground truth data for training deep neural networks in remote sensing.
method Unmixing high-resolution hyperspectral images to simulate ground truth for training a CNN.
result The model can classify mid-resolution Venus images successfully.
New method falsifies causal discovery results without ground truth.
problem Evaluation of causal discovery algorithms without ground truth data.
method Detects incompatibilities between causal graphs learned on different subsets of variables.
result Detection of incompatibilities can falsify wrongly inferred causal relations.
Unified framework for hyperparameter optimization and meta-learning.
problem Hyperparameter optimization and meta-learning.
method Bilevel programming, iterative inner objective solving, sufficient conditions for convergence.
result Software package Far-HO for hyperparameter optimization and meta-learning.
TCT learns multimodal sequence representations by translating from related sequences.
problem Challenges in learning semantic representations from multimodalities.
method Transformer based Cross-modal Translator (TCT) combined with Multimodal Transformer Network (MTN).
result Proposed method achieves new state-of-the-art performance on video-grounded dialogue.
FedSpace optimizes ML training on satellites and ground stations.
problem Training machine learning models on satellites with limited bandwidth.
method Federated Learning framework that dynamically schedules model aggregation based on satellite orbits.
result Reduces training time by 1.7 days over state-of-the-art algorithms.
This work tackles continual learning with semi-supervised data, showing that even with minimal labeled data, performance can match full-supervised methods.
problem Training deep networks on a stream of tasks without forgetting, especially when labeled data is scarce.
method Designing a novel CSSL method that leverages metric learning and consistency regularization to learn from both labeled and unlabeled data.
result Our method outperforms state-of-the-art methods trained with full supervision, achieving comparable performance with only 25% labeled data.
The paper analyzes GNNs with one hidden layer, proving their generalizability and convergence rate.
problem Theoretical guarantee on generalizability of GNNs with one hidden layer.
method Tensor initialization and accelerated gradient descent.
result The proposed learning algorithm converges to the ground-truth GNN model for regression and to a model close to the ground-truth for binary classification.
Active learning selects high-quality examples for text-to-SQL systems.
problem Efficiently annotate large language models for text-to-SQL systems.
method Formalizes example selection as a constrained experimental design problem over semantic query embeddings, proposing a stratified greedy algorithm that maximizes heteroscedastic mutual information.
result Proposed method significantly reduces labeling effort while maintaining high text-to-SQL retrieval accuracy.
Study reveals issues with neural autoregressive models and proposes mode recovery cost.
problem Unreasonable affinity of neural autoregressive models to short and long sequences.
method Investigates modes of ground-truth, empirical, and decoding-induced distributions via mode recovery cost.
result Mode recovery cost varies depending on ground-truth distribution and impacts decoding-induced distribution.
Study corrects misaligned cadaster maps using noisy supervision.
problem Correcting misaligned cadaster maps with noisy supervision data.
method Iterative training rounds to refine ground truth annotations.
result Reduces noise in cadaster map alignment datasets.