Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

59117176234 · May 202619922001200920182026
48 results for false positive ratio

Simple bounds show most cross-sectional predictability findings are likely true.

problem Determining the validity of cross-sectional return predictability findings.
method Developed simple and intuitive bounds on the false discovery rate (FDR).
result Bounds show the FDR is small, indicating most findings are likely true.

Paper generalizes PU classification for class prior shift and asymmetric error scenarios.

problem Bottlenecks in binary classification from PU data due to test marginal distribution and equal error penalties.
method Analysis of Bayes optimal classifier, risk minimization framework, and density ratio estimation framework.
result PU classification under class prior shift is equivalent to PU classification with asymmetric error.

New findings suggest weight maps from classifiers may not reliably indicate neural signals.

problem The reliability of interpreting weight maps from classifiers in neuroimaging studies.
method Used semi-simulated ECoG data to investigate signal-to-noise ratio and sparsity effects.
result Not all cases produce false positives and high-weight features are unlikely to be FP.

Machine learning detects subhalos in lensed images with high accuracy and low false positives.

problem Detecting substructure in strongly lensed images.
method Developed a neural network for image segmentation to locate and mass estimate subhalos.
result The network can detect subhalos with masses m108.5Mm\gtrsim 10^{8.5} M_{\odot} and measure the subhalo mass function.

This paper optimizes value investing with predictive modeling.

problem Empirical optimization of systematic value investing.
method Predictive modeling using financial metrics and statistical methods.
result Improved portfolio performance compared to traditional strategies.

Paper controls false positives in high-dimensional models using a novel approach.

problem Controlling false positives in high-dimensional models with the Lasso.
method Recast SQRT-Lasso as a false positive control method, extend to all GLMs, use fast Lasso solvers.
result Shows novel false positive control using random weighted self-normalized sums in finite samples.

Reduces false positives in lung nodule detection by using unlabeled data.

problem Lack of labeled data for training supervised algorithms in medical imaging.
method Uses pseudo-negative labels from unlabeled data to refine a pulmonary nodule detection network.
result False positive rate reduced from 0.4864 to 0.1266 while maintaining sensitivity.

The paper reviews techniques for detecting errors in semantic segmentation models.

problem Detecting false positives and false negatives in semantic segmentation models.
method Uncertainty quantification techniques applied to semantic segmentation.
result Techniques for detecting false positives and false negatives are proposed and discussed.

New method reduces false positives in weakly supervised pixel-level localization.

problem Reduces false positives in weakly supervised pixel-level localization.
method Proposes a deep learning method using conditional entropy to constrain the localizer.
result Significant improvements in image-level classification and pixel-level localization.

A new method for multiple testing reduces false discoveries while maximizing power.

problem Maximizing statistical power while controlling false discoveries in multiple testing scenarios.
method Adaptive sampling approach inspired by multi-armed bandits to minimize sample size.
result The method achieves sample complexity close to information theoretic lower bounds and outperforms uniform sampling.

New method calibrates false detection rates in sequential change detection.

problem Challenges in setting time-invariant thresholds for false positives.
method Simulation-based approach to time-varying thresholds.
result Accurately targets desired expected runtime while keeping false positive rate constant.

Bayesian approach confirms no return predictability for 1926-2004 data, weak evidence for 1953-2021.

problem Investigating return predictability using Bayesian methods.
method Developed a new shrinkage type prior for a model parameter in a VAR system, compared to other estimation methods.
result Bayesian approach outperforms reduced-bias estimator in terms of size and power.

PatternLocal improves XAI for non-linear models by suppressing suppressor variables.

problem Suppressor variables cause false-positive feature attributions in non-linear models.
method PatternLocal uses locally linear surrogate models and transforms weights into a generative representation.
result PatternLocal reduces false-positive attributions and provides more reliable explanations.

An adjusted NN algorithm reduces false negatives in imbalanced data.

problem Learning from imbalanced data, focusing on reducing false negatives.
method Introduces a reweighted distance scheme to modify Voronoi regions and decision boundaries.
result The method yields the best performance, especially when combined with sampling methods.

A statistical test controls false positives in anomaly localization using diffusion models.

problem Uncertainty and bias in generative models for anomaly localization.
method Selective inference to quantify significance and control false positives.
result The method effectively controls false positive detection rates.

Positive representations on surfaces have positive cross-ratios and satisfy a collar lemma.

problem Characterizing representations of surface groups with positive properties.
method Proving a collar lemma and showing positivity of cross-ratios for ΘΘ-positive representations.
result Closed subsets of representation varieties are characterized by ΘΘ-positive representations.

AnyThreat detects insider threats with minimal false positives.

problem High false positives in detecting insider threats.
method Opportunistic knowledge discovery system with four components: feature engineering, oversampling, class decomposition, and classification.
result Detects 87.5% of malicious insider threats with minimal false positives.

Nonparametric IPSS selects features with false discovery control.

problem Feature selection in high-dimensional data with theoretical false discovery control.
method Integrated Path Stability Selection (IPSS) applied to nonparametric feature importance scores.
result IPSS accurately controls false discovery rate and detects more true positives than existing methods.

Study compares shallow and deep learning for MS lesion segmentation.

problem Automated segmentation of white matter lesions in early-stage MS patients.
method Training and testing shallow and deep learning architectures on 32 patients.
result Combining shallow and deep architectures improves lesion-wise metrics.

Algorithm reconstructs triangle-free networks from data, certifying correctness.

problem Reconstructing triangle-free dynamic networks from observational data.
method Developed an algorithm for triangle-free networks, providing guarantees on correctness.
result Algorithm either certifies correctness or outputs a sparser graph with no false positives.

Framework uses human feedback to safely set OOD detection thresholds, reducing false positives.

problem Challenges in setting OOD detection thresholds for safety-critical applications.
method Mathematically grounded framework leveraging expert feedback to dynamically update thresholds.
result Guaranteed to meet FPR constraint while minimizing human feedback, maintaining FPR at most 5%.

New algorithm for adaptive experimental design in scientific settings.

problem Identifying true positives while controlling false discoveries in adaptive experimental design.
method Provably sample efficient adaptive algorithm for FDR control.
result First provably sample efficient adaptive algorithm for adaptive experimental design.

The paper introduces a method to incorporate feedback into tree-based anomaly detection to reduce false positives.

problem Difficulty in human analysts examining high-ranking anomalies due to false positives.
method Incorporates simple binary feedback into tree-based anomaly detectors, focusing on the Isolation Forest algorithm.
result Significantly improves the performance of the Isolation Forest algorithm by reducing false positives.

Develops a new criterion for subgroup fairness in algorithmic decision support.

problem Identifying fair recommendations in algorithms despite group-level differences.
method IJDI criterion and IJDI-Scan approach to detect and mitigate disparities.
result Identifies significant disparities in recommendations across subpopulations.

Deep Learning predicts e-commerce activity from Italian enterprise websites.

problem Predicting e-commerce activity from Italian enterprise websites.
method Developed a sophisticated processing pipeline using Convolutional Neural Networks and Word Embeddings.
result Deep Learning outperforms traditional Machine Learning methods for text classification.

Improved online changepoint detection for autocorrelated data.

problem Changepoint detection in autocorrelated data with false positives or delays.
method Generalized Likelihood Ratio (GLR) statistic for AR(p) processes, online focus algorithm.
result AR(p)-focus algorithm achieves high detection power in correlated data.

The paper optimizes A/B tests by balancing lift and cost in large-scale settings.

problem Balancing lift and cost in A/B tests for large-scale experimentation.
method Empirical Bayes approach using a greedy knapsack algorithm to rank experiments based on lift-to-cost ratio, incorporating local false discovery rate (lfdr).
result The proposed method maximizes expected profit while controlling false discovery rate, demonstrating superior performance in large-scale settings.

Deep learning improves seizure detection in EEGs.

problem Challenges in automated seizure detection in EEGs due to low signal-to-noise ratio and confusion with artifacts.
method Evaluation of hybrid deep structures including Convolutional Neural Networks and Long Short-Term Memory Networks on the TUH EEG Seizure Corpus.
result 30% sensitivity at 7 false alarms per 24 hours using a novel recurrent convolutional architecture.

Dual Teaching improves semi-supervised learning for practical applications.

problem Difficulty in applying semi-supervised wrapper methods to real-world data.
method Dual Teaching uses two external classifiers to estimate and adjust for false positives and negatives, training a base learner from partially labeled data.
result Dual Teaching effectively trains a base learner from partially labeled data as effectively as a fully-labeled-data-trained classifier.

New algorithm balances user reward and statistical inference by mixing TS with UR based on difference size.

problem Combining statistical inference with user reward in adaptive experiments.
method TS-PostDiff algorithm that uses UR when differences are small and TS when large.
result TS-PostDiff reduces false positives and increases statistical power for small differences, while maximizing reward for large ones.

Benchmarking recursive collapse claims with a new framework under false-positive control.

problem Evaluating recursive systems for failure patterns and warning claims.
method Developed Loopzero framework for testing recursive failures, specified claim boundaries in Lean, evaluated under FP constraint, and compared with standard detectors.
result No standard detectors or Loopzero's pre-registered quantile detector achieved the required operating point under the false-positive contract.

New method improves false-/true-positive-rate estimation in fraud detection with noisy labels.

problem Estimating FPR/TPR in fraud detection with class-conditional label noise.
method Directly cleaning model's validation data to de-correlate cleaning error with model scores.
result Improves accuracy of FPR/TPR estimates, especially in asymmetric label noise scenarios.