A new metric, Weighted Regret, unifies FDR and power evaluation in online multiple testing.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Optimizes quickest detection of drift in Brownian motion with false negatives.
New indefinite false theta functions match homological blocks for a specific 3-manifold.
Study controls error rates of binary classifiers using hypothesis testing.
In adversarial imitation learning, a discriminator is trained to differentiate agent episodes from expert demonstrations representing the desired behavior. However, as the trained policy learns to be more successful, the negative examples (the ones produced by the agent) become increasingly similar to expert ones. Desp…
In recent years, deep learning methods have outperformed other methods in image recognition. This has fostered imagination of potential application of deep learning technology including safety relevant applications like the interpretation of medical images or autonomous driving. The passage from assistance of a human d…
Develops a new criterion for subgroup fairness in algorithmic decision support.
Proposes cost-sensitive feature selection for SVMs.
State-of-the-art approaches for Knowledge Base Completion (KBC) exploit deep neural networks trained with both false and true assertions: positive assertions are explicitly taken from the knowledge base, whereas negative ones are generated by random sampling of entities. In this paper, we argue that random sampling is …
In this paper, we consider voxel selection for functional Magnetic Resonance Imaging (fMRI) brain data with the aim of finding a more complete set of probably correlated discriminative voxels, thus improving interpretation of the discovered potential biomarkers. The main difficulty in doing this is an extremely high di…
Algorithm reconstructs triangle-free networks from data, certifying correctness.
Generating large quantities of quality labeled data in medical imaging is very time consuming and expensive. The performance of supervised algorithms for various tasks on imaging has improved drastically over the years, however the availability of data to train these algorithms have become one of the main bottlenecks f…
We present a separation property for the gaps in the length spectrum of a compact Riemannian manifold with negative curvature. In arbitrary small neighborhoods of the metric for some suitable topology, we show that there are negatively curved metrics with a length spectrum exponentially separated from below. This prope…
Bayesian model improves categorization of explosions from sparse data.
We present a new approach for mitigating unfairness in learned classifiers. In particular, we focus on binary classification tasks over individuals from two populations, where, as our criterion for fairness, we wish to achieve similar false positive rates in both populations, and similar false negative rates in both po…
A genome-wide association study (GWAS) correlates marker variation with trait variation in a sample of individuals. Each study subject is genotyped at a multitude of SNPs (single nucleotide polymorphisms) spanning the genome. Here we assume that subjects are unrelated and collected at random and that trait values are n…
Study compares two methods to extend invariants, finding incompatibility for Brieskorn spheres.
This paper has been withdrawn due to an error in the proof. The paper implicitly assumed that every homotopically trivial knot is related to the unknot by negative crossings, which is false (as was pointed out to the author by Peter Ozsvath).
New algorithm balances user reward and statistical inference by mixing TS with UR based on difference size.
New techniques prove quantum modularity for various functions.
Information systems have widely been the target of malware attacks. Traditional signature-based malicious program detection algorithms can only detect known malware and are prone to evasion techniques such as binary obfuscation, while behavior-based approaches highly rely on the malware training samples and incur prohi…
We consider the problem of estimating the set of all inputs that leads a system to some particular behavior. The system is modeled by an expensive-to-evaluate function, such as a computer experiment, and we are interested in its excursion set, i.e. the set of points where the function takes values above or below some p…
Paper proves edge-connectivity equals minimum degree for graphs with non-negative curvature.
We introduce the State Classification Problem (SCP) for hybrid systems, and present Neural State Classification (NSC) as an efficient solution technique. SCP generalizes the model checking problem as it entails classifying each state of a hybrid automaton as either positive or negative, depending on whether or not …
We propose a novel algorithm for learning fair representations that can simultaneously mitigate two notions of disparity among different demographic subgroups in the classification setting. Two key components underpinning the design of our algorithm are balanced error rate and conditional alignment of representations. …
Paper introduces a statistical framework for watermarking LLM-generated text.
Boosting theory extended to handle cost-sensitive and multi-objective losses.
The paper tackles online learning with two types of losses and shows it's impossible without certain assumptions.
Noisy Pooled PCR tests large groups more efficiently.
Paper introduces negative margin loss for better few-shot classification accuracy.
PAC-Wrap provides provable guarantees for semi-supervised anomaly detection.
AdaDetectGPT improves text authorship detection with statistical guarantees.
Paper estimates FPR of Bayes classifier using soft labels.
MI attacks often mislabel non-training samples, making them impractical.
Environmental acoustic sensing involves the retrieval and processing of audio signals to better understand our surroundings. While large-scale acoustic data make manual analysis infeasible, they provide a suitable playground for machine learning approaches. Most existing machine learning techniques developed for enviro…
In regression settings where explanatory variables have very low correlations and there are relatively few effects, each of large magnitude, we expect the Lasso to find the important variables with few errors, if any. This paper shows that in a regime of linear sparsity---meaning that the fraction of variables with a n…
In this paper, we propose a Dual Focal Loss (DFL) function, as a replacement for the standard cross entropy (CE) function to achieve a better treatment of the unbalanced classes in a dataset. Our DFL method is an improvement on the recently reported Focal Loss (FL) cross-entropy function, which proposes a scaling metho…
New method identifies causes in time series with latent variables.
Detects harmful shifts without labels for model performance.
We study the Thompson sampling algorithm in an adversarial setting, specifically, for adversarial bit prediction. We characterize the bit sequences with the smallest and largest expected regret. Among sequences of length with zeros, the sequences of largest regret consist of alternating zeros and …
Bottlenecks of binary classification from positive and unlabeled data (PU classification) are the requirements that given unlabeled patterns are drawn from the test marginal distribution, and the penalty of the false positive error is identical to the false negative error. However, such requirements are often not fulfi…
The paper corrects bias in synthetic data for imbalanced learning.
The paper examines how machine learning tools in justice settings can unfairly affect different racial groups.
The paper shows how demographic data can lead to biased predictions, proposing 'Affirmative Information' as a solution.
Semi-supervised wrapper methods are concerned with building effective supervised classifiers from partially labeled data. Though previous works have succeeded in some fields, it is still difficult to apply semi-supervised wrapper methods to practice because the assumptions those methods rely on tend to be unrealistic i…
Study benchmarks label noise detection methods, identifying best practices.
We study the interplay between sequential decision making and avoiding discrimination against protected groups, when examples arrive online and do not follow distributional assumptions. We consider the most basic extension of classical online learning: "Given a class of predictors that are individually non-discriminato…
The generalized Chen's conjecture on biharmonic submanifolds asserts that any biharmonic submanifold of a non-positively curved manifold is minimal (see e.g., [CMO1], [MO], [BMO1], [BMO2], [BMO3], [Ba1], [Ba2], [Ou1], [Ou2], [IIU]). In this paper, we prove that this conjecture is false by constructing foliations of pro…