Simple bounds show most cross-sectional predictability findings are likely true.
problem Determining the validity of cross-sectional return predictability findings.
method Developed simple and intuitive bounds on the false discovery rate (FDR).
result Bounds show the FDR is small, indicating most findings are likely true.
Paper generalizes PU classification for class prior shift and asymmetric error scenarios.
problem Bottlenecks in binary classification from PU data due to test marginal distribution and equal error penalties.
method Analysis of Bayes optimal classifier, risk minimization framework, and density ratio estimation framework.
result PU classification under class prior shift is equivalent to PU classification with asymmetric error.
New findings suggest weight maps from classifiers may not reliably indicate neural signals.
problem The reliability of interpreting weight maps from classifiers in neuroimaging studies.
method Used semi-simulated ECoG data to investigate signal-to-noise ratio and sparsity effects.
result Not all cases produce false positives and high-weight features are unlikely to be FP.
Proposes tau-FPL for efficient false-positive rate control in linear time.
problem Learning classifiers with strict false-positive rate constraints.
method Scoring-thresholding approach with linear time efficiency.
result Superior performance over existing approaches in false-positive rate control.
The study examines how class imbalance affects precision-recall curves.
problem Understanding how precision changes with class imbalance ratios.
method Analyzes the relationship between precision, class imbalance ratio, and true/false positive rates.
result Predicts changes in precision-recall curves and other measures with class imbalance ratios.
Machine learning detects subhalos in lensed images with high accuracy and low false positives.
problem Detecting substructure in strongly lensed images.
method Developed a neural network for image segmentation to locate and mass estimate subhalos.
result The network can detect subhalos with masses m≳108.5M⊙ and measure the subhalo mass function. This paper optimizes value investing with predictive modeling.
problem Empirical optimization of systematic value investing.
method Predictive modeling using financial metrics and statistical methods.
result Improved portfolio performance compared to traditional strategies.
Two kernel Stein tests control decision errors in non-parametric model comparison.
problem Non-parametric multiple model comparison.
method Two statistical tests controlling false positive and false discovery rates.
result The first test has a higher true positive rate than the second under appropriate conditions.
Paper controls false positives in high-dimensional models using a novel approach.
problem Controlling false positives in high-dimensional models with the Lasso.
method Recast SQRT-Lasso as a false positive control method, extend to all GLMs, use fast Lasso solvers.
result Shows novel false positive control using random weighted self-normalized sums in finite samples.
Reduces false positives in lung nodule detection by using unlabeled data.
problem Lack of labeled data for training supervised algorithms in medical imaging.
method Uses pseudo-negative labels from unlabeled data to refine a pulmonary nodule detection network.
result False positive rate reduced from 0.4864 to 0.1266 while maintaining sensitivity.
The paper reviews techniques for detecting errors in semantic segmentation models.
problem Detecting false positives and false negatives in semantic segmentation models.
method Uncertainty quantification techniques applied to semantic segmentation.
result Techniques for detecting false positives and false negatives are proposed and discussed.
New method reduces unfairness in binary classification.
problem Achieving similar false positive and negative rates across two populations.
method Penalizes unfairness to achieve balanced false positive and negative rates.
result Empirically validated approach improves fairness and accuracy.
New method reduces false positives in weakly supervised pixel-level localization.
problem Reduces false positives in weakly supervised pixel-level localization.
method Proposes a deep learning method using conditional entropy to constrain the localizer.
result Significant improvements in image-level classification and pixel-level localization.
IHT reduces false positives and negatives in GWAS analysis.
problem Corrupted model selection in GWAS analysis.
method Iterative Hard Thresholding (IHT) algorithm for model selection.
result IHT reduces false positives and negatives while maintaining computational efficiency.
A new method for multiple testing reduces false discoveries while maximizing power.
problem Maximizing statistical power while controlling false discoveries in multiple testing scenarios.
method Adaptive sampling approach inspired by multi-armed bandits to minimize sample size.
result The method achieves sample complexity close to information theoretic lower bounds and outperforms uniform sampling.
New method calibrates false detection rates in sequential change detection.
problem Challenges in setting time-invariant thresholds for false positives.
method Simulation-based approach to time-varying thresholds.
result Accurately targets desired expected runtime while keeping false positive rate constant.
Bayesian approach confirms no return predictability for 1926-2004 data, weak evidence for 1953-2021.
problem Investigating return predictability using Bayesian methods.
method Developed a new shrinkage type prior for a model parameter in a VAR system, compared to other estimation methods.
result Bayesian approach outperforms reduced-bias estimator in terms of size and power.
Data augmentation improves keyword spotting accuracy in noisy conditions.
problem Maintaining low false reject rates in far-field KWS with playback interference.
method Artificially corrupted training data with mixed music and TV audio.
result 30-45% reduction in false reject rates under audio playback.
The Lasso path mixes true and false positives, leading to false discoveries.
problem The Lasso method's performance in sparse settings with low correlations.
method Approximate message passing (AMP) theory and adaptive selection of Lasso regularizing parameter.
result The Lasso path mixes true and false positives, leading to false discoveries.
Study controls error rates of binary classifiers using hypothesis testing.
problem Traditional binary classifiers have uncontrolled error rates.
method Combines binary classification with statistical hypothesis testing.
result Trained classifiers can be made to meet target error rate thresholds.
PatternLocal improves XAI for non-linear models by suppressing suppressor variables.
problem Suppressor variables cause false-positive feature attributions in non-linear models.
method PatternLocal uses locally linear surrogate models and transforms weights into a generative representation.
result PatternLocal reduces false-positive attributions and provides more reliable explanations.
An adjusted NN algorithm reduces false negatives in imbalanced data.
problem Learning from imbalanced data, focusing on reducing false negatives.
method Introduces a reweighted distance scheme to modify Voronoi regions and decision boundaries.
result The method yields the best performance, especially when combined with sampling methods.
A statistical test controls false positives in anomaly localization using diffusion models.
problem Uncertainty and bias in generative models for anomaly localization.
method Selective inference to quantify significance and control false positives.
result The method effectively controls false positive detection rates.
New method improves feature selection by integrating stability paths.
problem Improving feature selection with tighter false positive control.
method Integrating stability paths to strengthen theoretical bounds on E(FP).
result Significantly more true positives with same E(FP) control.
New representations on surfaces with positive cross ratios.
problem Understanding representations of surfaces with specific geometric properties.
method Using geodesic currents and Anosov representations, proving systolic inequalities.
result Systolic inequalities hold for all positively ratioed representations.
Positive representations on surfaces have positive cross-ratios and satisfy a collar lemma.
problem Characterizing representations of surface groups with positive properties.
method Proving a collar lemma and showing positivity of cross-ratios for Θ-positive representations. result Closed subsets of representation varieties are characterized by Θ-positive representations. Method detects batch heterogeneity in genomic data.
problem Batch effects confound genomic diagnostics.
method Bayesian model evidence clustering.
result Detects batch effects without known labels.
Evolutionary algorithm improves DNN watermarking with fewer false positives.
problem Protecting deep learning models from piracy and proving ownership.
method Evolutionary algorithm for generating and optimizing trigger patterns.
result Reduces false positive rates in DNN watermarking.
AnyThreat detects insider threats with minimal false positives.
problem High false positives in detecting insider threats.
method Opportunistic knowledge discovery system with four components: feature engineering, oversampling, class decomposition, and classification.
result Detects 87.5% of malicious insider threats with minimal false positives.
Nonparametric IPSS selects features with false discovery control.
problem Feature selection in high-dimensional data with theoretical false discovery control.
method Integrated Path Stability Selection (IPSS) applied to nonparametric feature importance scores.
result IPSS accurately controls false discovery rate and detects more true positives than existing methods.
Study compares shallow and deep learning for MS lesion segmentation.
problem Automated segmentation of white matter lesions in early-stage MS patients.
method Training and testing shallow and deep learning architectures on 32 patients.
result Combining shallow and deep architectures improves lesion-wise metrics.
Algorithm reconstructs triangle-free networks from data, certifying correctness.
problem Reconstructing triangle-free dynamic networks from observational data.
method Developed an algorithm for triangle-free networks, providing guarantees on correctness.
result Algorithm either certifies correctness or outputs a sparser graph with no false positives.
Framework uses human feedback to safely set OOD detection thresholds, reducing false positives.
problem Challenges in setting OOD detection thresholds for safety-critical applications.
method Mathematically grounded framework leveraging expert feedback to dynamically update thresholds.
result Guaranteed to meet FPR constraint while minimizing human feedback, maintaining FPR at most 5%.
New algorithm for adaptive experimental design in scientific settings.
problem Identifying true positives while controlling false discoveries in adaptive experimental design.
method Provably sample efficient adaptive algorithm for FDR control.
result First provably sample efficient adaptive algorithm for adaptive experimental design.
The paper introduces a method to incorporate feedback into tree-based anomaly detection to reduce false positives.
problem Difficulty in human analysts examining high-ranking anomalies due to false positives.
method Incorporates simple binary feedback into tree-based anomaly detectors, focusing on the Isolation Forest algorithm.
result Significantly improves the performance of the Isolation Forest algorithm by reducing false positives.
Develops a new criterion for subgroup fairness in algorithmic decision support.
problem Identifying fair recommendations in algorithms despite group-level differences.
method IJDI criterion and IJDI-Scan approach to detect and mitigate disparities.
result Identifies significant disparities in recommendations across subpopulations.
Deep Learning predicts e-commerce activity from Italian enterprise websites.
problem Predicting e-commerce activity from Italian enterprise websites.
method Developed a sophisticated processing pipeline using Convolutional Neural Networks and Word Embeddings.
result Deep Learning outperforms traditional Machine Learning methods for text classification.
Improved online changepoint detection for autocorrelated data.
problem Changepoint detection in autocorrelated data with false positives or delays.
method Generalized Likelihood Ratio (GLR) statistic for AR(p) processes, online focus algorithm.
result AR(p)-focus algorithm achieves high detection power in correlated data.
Proposes cost-sensitive feature selection for SVMs.
problem Asymmetric misclassification costs in feature selection.
method Mathematical optimization-based approach for SVMs.
result Substantial reduction in feature count with desired error rates.
The paper optimizes A/B tests by balancing lift and cost in large-scale settings.
problem Balancing lift and cost in A/B tests for large-scale experimentation.
method Empirical Bayes approach using a greedy knapsack algorithm to rank experiments based on lift-to-cost ratio, incorporating local false discovery rate (lfdr).
result The proposed method maximizes expected profit while controlling false discovery rate, demonstrating superior performance in large-scale settings.
Deep learning improves seizure detection in EEGs.
problem Challenges in automated seizure detection in EEGs due to low signal-to-noise ratio and confusion with artifacts.
method Evaluation of hybrid deep structures including Convolutional Neural Networks and Long Short-Term Memory Networks on the TUH EEG Seizure Corpus.
result 30% sensitivity at 7 false alarms per 24 hours using a novel recurrent convolutional architecture.
Dual Teaching improves semi-supervised learning for practical applications.
problem Difficulty in applying semi-supervised wrapper methods to real-world data.
method Dual Teaching uses two external classifiers to estimate and adjust for false positives and negatives, training a base learner from partially labeled data.
result Dual Teaching effectively trains a base learner from partially labeled data as effectively as a fully-labeled-data-trained classifier.
3D G-CNNs reduce false positives in lung nodule detection.
problem Reducing false positives in pulmonary nodule detection.
method Used 3D roto-translation group convolutions (G-Convs) instead of traditional convolutions.
result 3D G-CNNs achieved FROC scores close to those of a CNN trained on ten times more data.
New neural network outperforms existing methods in scene matching.
problem Automated scene matching with high accuracy and low false positives.
method Convolutional hashing using a new loss function and training scheme.
result Significantly higher true positive rate and 100-fold reduction in false positives.
New algorithm balances user reward and statistical inference by mixing TS with UR based on difference size.
problem Combining statistical inference with user reward in adaptive experiments.
method TS-PostDiff algorithm that uses UR when differences are small and TS when large.
result TS-PostDiff reduces false positives and increases statistical power for small differences, while maximizing reward for large ones.
Novel online graph-based method detects changes in high-dimensional data.
problem Challenges in detecting changes in high-dimensional data.
method Graph-based similarity measure derived from graph-spanning ratio.
result High detection power and controlled false alarm rate for high-dimensional data.
Benchmarking recursive collapse claims with a new framework under false-positive control.
problem Evaluating recursive systems for failure patterns and warning claims.
method Developed Loopzero framework for testing recursive failures, specified claim boundaries in Lean, evaluated under FP constraint, and compared with standard detectors.
result No standard detectors or Loopzero's pre-registered quantile detector achieved the required operating point under the false-positive contract.
New method improves false-/true-positive-rate estimation in fraud detection with noisy labels.
problem Estimating FPR/TPR in fraud detection with class-conditional label noise.
method Directly cleaning model's validation data to de-correlate cleaning error with model scores.
result Improves accuracy of FPR/TPR estimates, especially in asymmetric label noise scenarios.