Bayesian method corrects for model selection multiplicity in regression.
problem Model selection multiplicity in regression analysis.
method Developed a Bayesian prior distribution based on Holm procedure analogy.
result Adequate multiplicity correction requires sparsity not provided by recommended priors.
Max-rank improves multiple testing in conformal prediction.
problem Simultaneous testing of multiple hypotheses in scientific inquiries.
method Introduces max-rank, a novel correction for positive dependencies in simultaneous testing.
result Max-rank efficiently controls family-wise error rate and improves predictive uncertainty estimates.
A method for inferring ground-truth signals from degraded sensor data.
problem Inferring ground-truth signals from multiple degraded sensor signals.
method Iterative correction of degraded signals using a Bayesian multi-sensor data fusion method.
result The method effectively infers ground-truth signals from noisy and degraded sensor data.
Closed formulas for η-corrections in the once-punctured torus identified.
problem Identifying η-corrections in the Kauffman bracket skein algebra of the once-punctured torus.
method Explicit closed formulas for Chebyshev-threaded families and η-corrections.
result Explicit Chebyshev expansions and coefficients for η-corrections.
RAMEN corrects observational data biases for multiple environments.
problem Bias in observational data for causal inference.
method RAMEN algorithm that leverages heterogeneity of multiple data sources.
result RAMEN produces unbiased treatment effect estimates.
We consider off-policy policy evaluation when the trajectory data are generated by multiple behavior policies. Recent work has shown the key role played by the state or state-action stationary distribution corrections in the infinite horizon context for off-policy policy evaluation. We propose estimated mixture policy …
Approach generates multiple correct predictions from single supervision.
problem Single correct prediction from multiple possible alternatives.
method Develops an approach to generate multiple high-quality predictions.
result Can generate high-quality outputs different from observed.
Paper finds methods to accurately determine the number of clusters in data.
problem Finding the correct number of clusters in a dataset.
method Penalized k-means algorithms with ideal clusters and multiplicative penalties.
result K-means with multiplicative penalties provides a clearer indication of the correct number of clusters.
Industrial recommender systems deal with extremely large action spaces -- many millions of items to recommend. Moreover, they need to serve billions of users, who are unique at any point in time, making a complex user state space. Luckily, huge quantities of logged implicit feedback (e.g., user clicks, dwell time) are …
We determine the sample complexity of pure exploration bandit problems with multiple good answers. We derive a lower bound using a new game equilibrium argument. We show how continuity and convexity properties of single-answer problems ensures that the Track-and-Stop algorithm has asymptotically optimal sample complexi…
The problem of finding itemsets that are statistically significantly enriched in a class of transactions is complicated by the need to correct for multiple hypothesis testing. Pruning untestable hypotheses was recently proposed as a strategy for this task of significant itemset mining. It was shown to lead to greater s…
Ogburn et al. (2019, arXiv:1910.05438) discuss "The Blessings of Multiple Causes" (Wang and Blei, 2018, arXiv:1805.06826). Many of their remarks are interesting. But they also claim that the paper has "foundational errors" and that its "premise is...incorrect." These claims are not substantiated. There are no foundatio…
We address the problem of learning vector representations for entities and relations in Knowledge Graphs (KGs) for Knowledge Base Completion (KBC). This problem has received significant attention in the past few years and multiple methods have been proposed. Most of the existing methods in the literature use a predefin…
Paper shows unique decomposition of 3-manifolds and multiplicative property of Reidemeister torsion.
problem Unique decomposition of compact 3-manifolds.
method Study of adjoint Reidemeister torsion on disk sum decompositions.
result Adjoint Reidemeister torsion has a multiplicative property on unique decompositions.
This paper certifies cluster assignments from sum-of-norms clustering algorithms.
problem Certifying the correct cluster assignments from approximate solutions of sum-of-norms clustering.
method Presented a clustering test that identifies and certifies the correct cluster assignment from an approximate solution.
result The correct cluster assignment is guaranteed to be certified by a primal-dual path following algorithm after sufficient iterations.
In machine learning the best performance on a certain task is achieved by fully supervised methods when perfect ground truth labels are available. However, labels are often noisy, especially in remote sensing where manually curated public datasets are rare. We study the multi-modal cadaster map alignment problem for wh…
Paper proposes a new meta-learning approach for correcting noisy labels.
problem Learning with noisy labels in machine learning models.
method Meta-learned instance re-weighting approach extended to label correction problem.
result Proposed MLC (Meta Label Correction) framework achieves large improvements over previous methods.
CRC improves multivariate forecasting accuracy without risking performance degradation.
problem Systematic errors and lack of guarantees in multivariate forecasters.
method CRC uses a causality-inspired encoder and hybrid corrector with a safety mechanism.
result CRC consistently improves accuracy and ensures high non-degradation rates.
New method corrects complex distortions in single view images.
problem Complex distortions in images, especially those caused by refractive surfaces.
method Differentiable image sampling and semantic information augmentation.
result Model can estimate and correct highly complex distortions.
New MCMC method corrects bias without extra cost.
problem Correcting bias in MCMC algorithms without additional computational cost.
method Generalized Markov Chain Importance Sampling methods.
result Proposed methods are more efficient than Metropolis-Hastings versions.
Tumors often contain multiple subpopulations of cancerous cells defined by distinct somatic mutations. We describe a new method, PhyloWGS, that can be applied to WGS data from one or more tumor samples to reconstruct complete genotypes of these subpopulations based on variant allele frequencies (VAFs) of point mutation…
The study corrects measurement error in evaluating health effects of multiple pollutants.
problem Bias in estimating health effects of air pollution constituents due to mismeasurement.
method Used a linear regression calibration model and extended DML approach to correct for measurement error.
result Identified two PM2.5 constituents (Br and Mn) that show a negative causal effect on cognitive function after correction.
This work explores test-time scaling strategies for LLMs, improving sample efficiency and expressiveness.
problem Understanding the sample efficiency and expressiveness of test-time scaling strategies for LLMs.
method Established separation and expressiveness results for self-consistency, best-of-n, and self-correction strategies. result Self-correction enables Transformers to simulate online learning over multiple tasks without prior knowledge.
Proposes a new algorithm to estimate invariant subspaces across multilayer networks.
problem Estimating invariant subspaces across heterogeneous multiple networks.
method Bias-corrected joint spectral embedding algorithm that recursively calibrates diagonal bias and iteratively updates the subspace estimator.
result Established entrywise subspace perturbation bound and entrywise eigenvector central limit theorem for the algorithm.
Training-free method improves text-to-image generation quality.
problem Error accumulation in simultaneous token updates of masked diffusion models.
method Training-free self-correction framework exploiting inductive biases.
result Significantly improved generation quality on text-to-image tasks.
Logit correction improves model performance by correcting spurious correlations.
problem Spurious correlations lead to poor model performance during inference.
method Proposes logit correction (LC) loss to mitigate spurious correlations.
result LC loss outperforms state-of-the-art solutions by 5.5% absolute improvement.
In this paper, we extend the β-CNMF to two dimensions and derive exact multiplicative updates for its factors. The new updates generalize and correct the nonnegative matrix factor deconvolution previously proposed by Schmidt and Mørup. We show by simulation that the updates lead to a monotonically decreasing β-dive…
Confidence intervals improve evaluation of binary prediction rules in data mining.
problem Uncertainty in performance measures estimation from finite datasets.
method Asymptotic normal approximations for confidence intervals, with a blurring correction.
result Improved finite sample coverage probabilities and general performance measures inference.
Two kernel Stein tests control decision errors in non-parametric model comparison.
problem Non-parametric multiple model comparison.
method Two statistical tests controlling false positive and false discovery rates.
result The first test has a higher true positive rate than the second under appropriate conditions.
A new method improves super learner validation efficiency.
problem Improving the efficiency of super learner validation.
method Bootstrap Bias Corrected Cross Validation applied to Super Learning.
result Bootstrap Bias Corrected Cross Validation proved efficient and cost-effective.
In this letter, we generalize the convolutional NMF by taking the β-divergence as the contrast function and present the correct multiplicative updates for its factors in closed form. The new updates unify the β-NMF and the convolutional NMF. We state why almost all of the existing updates are inexact and approximat…
Proposes a method to compute valid lower confidence bounds for multiple models selected based on their performance.
problem Model selection and evaluation in machine learning.
method Interprets model selection as a simultaneous inference problem, uses bootstrap tilting and maxT-type multiplicity correction.
result Yields valid lower confidence bounds that are at least as good as standard approaches and reliably reach nominal coverage probability.
This note corrects a mistake in the original book in the evolution equations of total curvature for the curve-shrinking flow in an ambient Ricci Flow. The resulting upper bound for the evolution of total curvature is an exponential bound in time. The change involves the multiplicative constant. Here we show that it dep…
Distributed learning is an effective way to analyze big data. In distributed regression, a typical approach is to divide the big data into multiple blocks, apply a base regression algorithm on each of them, and then simply average the output functions learnt from these blocks. Since the average process will decrease th…
Unified approach for multicalibration in weakly supervised learning.
problem Existing multicalibration methods require clean input-label pairs, which are unavailable in weakly supervised learning.
method Developed estimators and post-hoc correction methods for multicalibration under weak supervision.
result Unified framework for estimating and correcting multicalibration under weak supervision with finite-sample guarantees.
The paper develops methods to assess and correct model uncertainties in graphical models.
problem Model uncertainty in probabilistic graphical models.
method Information-theoretic and non-parametric stress tests.
result Ranking and correcting impactful sources of uncertainty in graphical models.
QC-ST and CoCo methods correct batch effects in metabolomics data.
problem Batch effects in metabolomics data obscure biological variations.
method QC-ST for simultaneous detection of QC samples' mean vectors and covariance matrices, CoCo for covariance correction.
result QC-ST and CoCo improve batch effect correction in metabolomics datasets.
This work uses adversarial learning to detect and correct feature shifts in various datasets.
problem Detecting and correcting feature shifts in real-world datasets.
method Adversarial learning applied to multiple discriminators to detect and correct feature shifts.
result Mainstream classifiers can effectively localize and correct feature shifts, outperforming existing techniques.
GrowNet uses shallow neural networks for gradient boosting, outperforming existing methods.
problem Improving gradient boosting performance through shallow neural networks.
method Unified gradient boosting framework with shallow neural networks as weak learners, incorporating corrective steps.
result GrowNet outperformed state-of-the-art boosting methods in classification, regression, and learning to rank tasks.
Algorithm identifies and corrects noisy labels using Gaussian process regression.
problem Detecting and correcting real-valued noisy labels from mixed data.
method Gaussian process regression with heteroscedastic noise model and leave-one-out cross-validation.
result The method can pinpoint corrupted sample points and improve regression models.
Crowdlab uses classifiers to estimate consensus labels and annotator quality.
problem Leveraging multiple annotators for data classification.
method Weighted ensemble approach using any trained classifier.
result Superior estimates for consensus labels and annotator quality.
Corrected graph convolutions improve node classification on graphs.
problem Oversmoothing in graph convolutions degrades performance.
method Theoretical analysis based on CSBM, spectral analysis for k rounds of corrected graph convolutions.
result Corrected graph convolutions can improve node classification performance exponentially.
Proposes DC-S3GD for efficient large-scale decentralized neural network training.
problem Training large-scale decentralized neural networks efficiently.
method Decentralized stale-synchronous version of DC-ASGD with gradient correction.
result Achieves state-of-the-art results in training Convolutional Neural Networks.
New method corrects bias in stochastic gradient samplers.
problem Bias in stochastic gradient samplers.
method Gradient-Guided Monte Carlo with stochastic gradients.
result Corrected sampler yields nonzero acceptance probabilities.
End-to-end KBQA system learns from multiple reasoning paths without labeled paths.
problem Lack of labeled reasoning paths limits KBQA system performance.
method End-to-end KBQA system using multiple reasoning paths.
result Demonstrates strong performance on various KBQA datasets.
In 1978 Brakke introduced the mean curvature flow in the setting of geometric measure theory. There exist multiple variants of the original definition. Here we prove that most of them are indeed equal. One central point is to correct the proof of Brakke's §3.5, where he develops an estimate for the evolution of the mea…
Geometric approach finds correspondences between different conditions.
problem Integrating multiple biological datasets.
method Fibered latent space with pull-back metric, diffeomorphism flows.
result Minimizing energy functional yields diffeomorphism flows.
New method learns to answer questions from correct demonstrations without assuming bounded complexity of the demonstrator.
problem Learning to generate answers from correct demonstrations with multiple correct answers.
method Formalizes imitation learning in contextual bandits, relying on reward model complexity, not policy complexity.
result Achieves nearly optimal performance with sample complexity logarithmic in reward class cardinality.