We present an efficient algorithm for simultaneously training sparse generalized linear models across many related problems, which may arise from bootstrapping, cross-validation and nonparametric permutation testing. Our approach leverages the redundancies across problems to obtain significant computational improvement…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The problem of multiple hypothesis testing arises when there are more than one hypothesis to be tested simultaneously for statistical significance. This is a very common situation in many data mining applications. For instance, assessing simultaneously the significance of all frequent itemsets of a single dataset entai…
Due to the increasing availability of high-dimensional empirical applications in many research disciplines, valid simultaneous inference becomes more and more important. For instance, high-dimensional settings might arise in economic studies due to very rich data sets with many potential covariates or in the analysis o…
Unified framework for large-scale hypothesis testing with confounders.
Financial econometrics has become an increasingly popular research field. In this paper we review a few parametric and nonparametric models and methods used in this area. After introducing several widely used continuous-time and discrete-time models, we study in detail dependence structures of discrete samples, includi…
Statistical tests that compare classification algorithms are univariate and use a single performance measure, e.g., misclassification error, measure, AUC, and so on. In multivariate tests, comparison is done using multiple measures simultaneously. For example, error is the sum of false positives and false negatives…
New classifiers converge under large data, simplifying complex models.
Max-rank improves multiple testing in conformal prediction.
This article studies local and global inference for smoothing spline estimation in a unified asymptotic framework. We first introduce a new technical tool called functional Bahadur representation, which significantly generalizes the traditional Bahadur representation in parametric models, that is, Bahadur [Ann. Inst. S…
Pareto Testing optimizes model performance under multiple constraints.
New method identifies structural parameters without assuming uncorrelated errors.
Paper proposes a new framework for hypothesis testing in imaging.
We propose two nonparametric statistical tests of goodness of fit for conditional distributions: given a conditional probability density function and a joint sample, decide whether the sample is drawn from for some density . Our tests, formulated with a Stein operator, can be applied to any…
Two statistical tasks are shown to have equivalent sample complexity.
Efficiently estimates marginal posteriors for complex simulations.
The paper provides high-probability bounds on false discovery proportions in conformal inference.
New attacks show ML models can be compromised even when targeting one concept.
QC-ST and CoCo methods correct batch effects in metabolomics data.
We consider a process , which is observed on a finite time interval , at discrete times This process is an Itô semimartingale with stochastic volatility . Assuming that has jumps on , we derive tests to decide whether the volatility process has jumps occurring simultan…
A distributed bootstrap method for high-dimensional data reduces communication rounds efficiently.
Unified data representation learning improves non-parametric two-sample testing.
A new witness two-sample test improves data efficiency and power.
There is a significant literature on methods for incorporating knowledge into multiple testing procedures so as to improve their power and precision. Some common forms of prior knowledge include (a) beliefs about which hypotheses are null, modeled by non-uniform prior weights; (b) differing importances of hypotheses, m…
Current statistical inference problems in areas like astronomy, genomics, and marketing routinely involve the simultaneous testing of thousands -- even millions -- of null hypotheses. For high-dimensional multivariate distributions, these hypotheses may concern a wide range of parameters, with complex and unknown depen…
New PAC-Bayes bound controls multiple error types simultaneously.
Free adversarial training reduces the generalization gap compared to vanilla method.
Consider the online testing of a stream of hypotheses where a real--time decision must be made before the next data point arrives. The error rate is required to be controlled at {all} decision points. Conventional \emph{simultaneous testing rules} are no longer applicable due to the more stringent error constraints and…
Develops model selection for bandits balancing adversarial and stochastic guarantees.
New exact tests detect changepoints in binary and count data, especially when normal approximations fail.
Natural and social multivariate systems are commonly studied through sets of simultaneous and time-spaced measurements of the observables that drive their dynamics, i.e., through sets of time series. Typically, this is done via hypothesis testing: the statistical properties of the empirical time series are tested again…
Proposes a method to calibrate data for more accurate linear correlation testing.
Paper models cloud outages for cyber insurance stress-testing.
The use of deep networks to extract embeddings for speaker recognition has proven successfully. However, such embeddings are susceptible to performance degradation due to the mismatches among the training, enrollment, and test conditions. In this work, we propose an adversarial speaker verification (ASV) scheme to lear…
System tackles indeterminacies in automated audio captioning.
Implicit models can match or exceed explicit models with more test-time compute.
A goodness-of-fit test for DCSBM improves scalability and power for large sparse networks.
We present a generalization of the Simultaneous Long-Short (SLS) trading strategy described in recent control literature wherein we allow for different parameters across the short and long sides of the controller; we refer to this new strategy as Generalized SLS (GSLS). Furthermore, we investigate the conditions under …
Framework calibrates ML models for risk control in various tasks.
Under the Fundamental Review of the Trading Book (FRTB) capital charges for the trading book are based on the coherent expected shortfall (ES) risk measure, which show greater sensitivity to tail risk. In this paper it is argued that backtesting of expected shortfall - or the trading book model from which it is calcula…
A PID-based feedback-control system improves multiple KPIs in RTB display advertising.
Regression-via-Classification (RvC) is the process of converting a regression problem to a classification one. Current approaches for RvC use ad-hoc discretization strategies and are suboptimal. We propose a neural regression tree model for RvC. In this model, we employ a joint optimization framework where we learn opt…
Paper debiases multiple word embedding biases simultaneously.
A method for rank verification in multivariate Gaussian data, improving on existing approaches.
This paper studies the problem of learning with augmented classes (LAC), where augmented classes unobserved in the training data might emerge in the testing phase. Previous studies generally attempt to discover augmented classes by exploiting geometric properties, achieving inspiring empirical performance yet lacking t…
Transformer RL optimizes A/B testing for time series experiments.
Topic modeling based on latent Dirichlet allocation (LDA) has been a framework of choice to perform scene recognition and annotation. Recently, a new type of topic model called the Document Neural Autoregressive Distribution Estimator (DocNADE) was proposed and demonstrated state-of-the-art performance for document mod…
The continually increasing number of complex datasets each year necessitates ever improving machine learning methods for robust and accurate categorization of these data. This paper introduces Random Multimodel Deep Learning (RMDL): a new ensemble, deep learning approach for classification. Deep learning models have ac…
Adaptive auditing improves AI robustness testing with anytime-valid guarantees.