Bayesian method corrects for model selection multiplicity in regression.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Max-rank improves multiple testing in conformal prediction.
A method for inferring ground-truth signals from degraded sensor data.
Closed formulas for η-corrections in the once-punctured torus identified.
RAMEN corrects observational data biases for multiple environments.
We consider off-policy policy evaluation when the trajectory data are generated by multiple behavior policies. Recent work has shown the key role played by the state or state-action stationary distribution corrections in the infinite horizon context for off-policy policy evaluation. We propose estimated mixture policy …
Approach generates multiple correct predictions from single supervision.
Industrial recommender systems deal with extremely large action spaces -- many millions of items to recommend. Moreover, they need to serve billions of users, who are unique at any point in time, making a complex user state space. Luckily, huge quantities of logged implicit feedback (e.g., user clicks, dwell time) are …
We determine the sample complexity of pure exploration bandit problems with multiple good answers. We derive a lower bound using a new game equilibrium argument. We show how continuity and convexity properties of single-answer problems ensures that the Track-and-Stop algorithm has asymptotically optimal sample complexi…
The problem of finding itemsets that are statistically significantly enriched in a class of transactions is complicated by the need to correct for multiple hypothesis testing. Pruning untestable hypotheses was recently proposed as a strategy for this task of significant itemset mining. It was shown to lead to greater s…
Ogburn et al. (2019, arXiv:1910.05438) discuss "The Blessings of Multiple Causes" (Wang and Blei, 2018, arXiv:1805.06826). Many of their remarks are interesting. But they also claim that the paper has "foundational errors" and that its "premise is...incorrect." These claims are not substantiated. There are no foundatio…
We address the problem of learning vector representations for entities and relations in Knowledge Graphs (KGs) for Knowledge Base Completion (KBC). This problem has received significant attention in the past few years and multiple methods have been proposed. Most of the existing methods in the literature use a predefin…
Paper shows unique decomposition of 3-manifolds and multiplicative property of Reidemeister torsion.
We address the problem of non-parametric multiple model comparison: given candidate models, decide whether each candidate is as good as the best one(s) or worse than it. We propose two statistical tests, each controlling a different notion of decision errors. The first test, building on the post selection inference…
In many applications we want to find the number of clusters in a dataset. A common approach is to use the penalized k-means algorithm with an additive penalty term linear in the number of clusters. An open problem is estimating the value of the coefficient of the penalty term. Since estimating the value of the coeffici…
This paper certifies cluster assignments from sum-of-norms clustering algorithms.
In machine learning the best performance on a certain task is achieved by fully supervised methods when perfect ground truth labels are available. However, labels are often noisy, especially in remote sensing where manually curated public datasets are rare. We study the multi-modal cadaster map alignment problem for wh…
Leveraging weak or noisy supervision for building effective machine learning models has long been an important research problem. Its importance has further increased recently due to the growing need for large-scale datasets to train deep learning models. Weak or noisy supervision could originate from multiple sources i…
CRC improves multivariate forecasting accuracy without risking performance degradation.
New MCMC method corrects bias without extra cost.
Tumors often contain multiple subpopulations of cancerous cells defined by distinct somatic mutations. We describe a new method, PhyloWGS, that can be applied to WGS data from one or more tumor samples to reconstruct complete genotypes of these subpopulations based on variant allele frequencies (VAFs) of point mutation…
The study corrects measurement error in evaluating health effects of multiple pollutants.
This work explores test-time scaling strategies for LLMs, improving sample efficiency and expressiveness.
Proposes a new algorithm to estimate invariant subspaces across multilayer networks.
Most distortion correction methods focus on simple forms of distortion, such as radial or linear distortions. These works undistort images either based on measurements in the presence of a calibration grid, or use multiple views to find point correspondences and predict distortion parameters. When possible distortions …
Training-free method improves text-to-image generation quality.
Logit correction improves model performance by correcting spurious correlations.
In this paper, we extend the -CNMF to two dimensions and derive exact multiplicative updates for its factors. The new updates generalize and correct the nonnegative matrix factor deconvolution previously proposed by Schmidt and Mørup. We show by simulation that the updates lead to a monotonically decreasing -dive…
Confidence intervals improve evaluation of binary prediction rules in data mining.
A new method improves super learner validation efficiency.
In this letter, we generalize the convolutional NMF by taking the -divergence as the contrast function and present the correct multiplicative updates for its factors in closed form. The new updates unify the -NMF and the convolutional NMF. We state why almost all of the existing updates are inexact and approximat…
Proposes a method to compute valid lower confidence bounds for multiple models selected based on their performance.
This note corrects a mistake in the original book in the evolution equations of total curvature for the curve-shrinking flow in an ambient Ricci Flow. The resulting upper bound for the evolution of total curvature is an exponential bound in time. The change involves the multiplicative constant. Here we show that it dep…
Distributed learning is an effective way to analyze big data. In distributed regression, a typical approach is to divide the big data into multiple blocks, apply a base regression algorithm on each of them, and then simply average the output functions learnt from these blocks. Since the average process will decrease th…
Unified approach for multicalibration in weakly supervised learning.
The paper develops methods to assess and correct model uncertainties in graphical models.
QC-ST and CoCo methods correct batch effects in metabolomics data.
This work uses adversarial learning to detect and correct feature shifts in various datasets.
Algorithm identifies and corrects noisy labels using Gaussian process regression.
Crowdlab uses classifiers to estimate consensus labels and annotator quality.
Corrected graph convolutions improve node classification on graphs.
New method corrects bias in stochastic gradient samplers.
End-to-end KBQA system learns from multiple reasoning paths without labeled paths.
In 1978 Brakke introduced the mean curvature flow in the setting of geometric measure theory. There exist multiple variants of the original definition. Here we prove that most of them are indeed equal. One central point is to correct the proof of Brakke's §3.5, where he develops an estimate for the evolution of the mea…
Geometric approach finds correspondences between different conditions.
New method learns to answer questions from correct demonstrations without assuming bounded complexity of the demonstrator.
After being trained, classifiers must often operate on data that has been corrupted by noise. In this paper, we consider the impact of such noise on the features of binary classifiers. Inspired by tools for classifier robustness, we introduce the same classification probability (SCP) to measure the resulting distortion…
Suppose that multiple experts (or learning algorithms) provide us with alternative Bayesian network (BN) structures over a domain, and that we are interested in combining them into a single consensus BN structure. Specifically, we are interested in that the consensus BN structure only represents independences all the g…