A communication-efficient method controls FDR in network settings.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Extends knockoff filter for composite null hypotheses in variable selection.
New method controls false discoveries in online testing with deadlines.
Paper develops a framework to derive lower bounds on FDR and FNR in multiple testing.
Unified framework for large-scale hypothesis testing with confounders.
Online method selects candidates from data streams, ensuring irreversible decisions.
Transformer learns representations from time series data for money laundering detection.
This paper defines systematic value investing as an empirical optimization problem. Predictive modeling is introduced as a systematic value investing methodology with dynamic and optimization features. A predictive modeling process is demonstrated using financial metrics from Gray & Carlisle and Buffett & Clark. A 31-y…
New test ensures quality of shared data in machine learning.
We propose a linear-time, single-pass, top-down algorithm for multiple testing on directed acyclic graphs (DAGs), where nodes represent hypotheses and edges specify a partial ordering in which hypotheses must be tested. The procedure is guaranteed to reject a sub-DAG with bounded false discovery rate (FDR) while satisf…
New method improves reliability of selecting individuals based on predicted treatment effects.
Multiple hypothesis testing is a core problem in statistical inference and arises in almost every scientific field. Given a set of null hypotheses , Benjamini and Hochberg introduced the false discovery rate (FDR), which is the expected proportion of false positives among rejected nu…
Paper proposes a privacy-preserving method to control false discoveries.
As datasets grow richer, an important challenge is to leverage the full features in the data to maximize the number of useful discoveries while controlling for false positives. We address this problem in the context of multiple hypotheses testing, where for each hypothesis, we observe a p-value along with a set of feat…
PH-CS selects test inputs with reliability guarantees, adapting FDR to data.
We provide the first differentially private algorithms for controlling the false discovery rate (FDR) in multiple hypothesis testing, with essentially no loss in power under certain conditions. Our general approach is to adapt a well-known variant of the Benjamini-Hochberg procedure (BHq), making each step differential…
The paper provides high-probability bounds on false discovery proportions in conformal inference.
One important partition of algorithms for controlling the false discovery rate (FDR) in multiple testing is into offline and online algorithms. The first generally achieve significantly higher power of discovery, while the latter allow making decisions sequentially as well as adaptively formulating hypotheses based on …
Kalshi prediction markets forecast cryptocurrency volatility through monetary policy and inflation signals.
RL-Exec uses reinforcement learning to optimize BTC-USD liquidation, outperforming traditional methods.
ROOFS helps researchers select robust biomarker features from complex data.
We study learning problems involving arbitrary classes of functions , distributions and targets . Because proper learning procedures, i.e., procedures that are only allowed to select functions in , tend to perform poorly unless the problem satisfies some additional structural property (e.g., that is co…
A study shows that a fine-tuned model's directional accuracy in financial forecasting is largely due to chance, not skill.
The need for parameter estimation with massive datasets has reinvigorated interest in stochastic optimization and iterative estimation procedures. Stochastic approximations are at the forefront of this recent development as they yield procedures that are simple, general, and fast. However, standard stochastic approxima…
We describe a procedure which verifies that a group given by generators and relators is word-hyperbolic. This procedure always works with a group which is word-hyperbolic, provided there is sufficient memory and time devoted to the problem. If the group is not word-hyperbolic, the procedure continues indefinitely. We a…
Myopic procedures are shown to be asymptotically optimal in ranking and selection problems.
Several authors have pointed out the connection between Barbilian's metric introduced in 1934 and the recent study of Apollonian metrics. We provide examples of various distances that can be obtained by Barbilian's metrization procedure and we discuss the relation between this metrization procedure and important Rieman…
Biclustering, the process of simultaneously clustering the rows and columns of a data matrix, is a popular and effective tool for finding structure in a high-dimensional dataset. Many biclustering procedures appear to work well in practice, but most do not have associated consistency guarantees. To address this shortco…
Chemical plants are complex and dynamical systems consisting of many components for manipulation and sensing, whose state transitions depend on various factors such as time, disturbance, and operation procedures. For the purpose of supporting human operators of chemical plants, we are developing an AI system that can s…
A regularized risk minimization procedure for regression function estimation is introduced that achieves near optimal accuracy and confidence under general conditions, including heavy-tailed predictor and response variables. The procedure is based on median-of-means tournaments, introduced by the authors in [8]. It is …
ARK improves knockoffs robustness to feature distribution misspecification.
A new procedure, called DDa-procedure, is developed to solve the problem of classifying d-dimensional objects into q >= 2 classes. The procedure is completely nonparametric; it uses q-dimensional depth plots and a very efficient algorithm for discrimination analysis in the depth space [0,1]^q. Specifically, the depth i…
Investigation of the market graph attracts a growing attention in market network analysis. One of the important problem connected with market graph is to identify it from observations. Traditional way for the market graph identification is to use a simple procedure based on statistical estimations of Pearson correlatio…
A new method calibrates forecasts without sacrificing expertise.
Procedure groups nonparametric regression curves automatically.
Procgen Benchmark uses procedurally generated games to test reinforcement learning.
We model the quantities appearing in Internal Revenue Service (IRS) tax guidance for calculating the health insurance premium tax credit created by the Patient Protection and Affordable Care Act, also called Obamacare. We ask the question of whether there is a procedure, computable by hand, which can calculate the appr…
We present a new, fully generative model for constructing astronomical catalogs from optical telescope image sets. Each pixel intensity is treated as a random variable with parameters that depend on the latent properties of stars and galaxies. These latent properties are themselves modeled as random. We compare two pro…
Let $\cF$ be a set of classification procedures with values in . Given a loss function, we want to construct a procedure which mimics at the best possible rate the best procedure in $\cF$. This fastest rate is called optimal rate of aggregation. Considering a continuous scale of loss functions with various …
We develop a mixture procedure for multi-sensor systems to monitor data streams for a change-point that causes a gradual degradation to a subset of the streams. Observations are assumed to be initially normal random variables with known constant means and variances. After the change-point, observations in the subset wi…
The paper describes a method to infer the signal-to-noise ratio in portfolio optimization.
In this article we give our contribution to the problem of segmentation with plug-in procedures. We give general sufficient conditions under which plug in procedure are efficient. We also give an algorithm that satisfy these conditions. We give an application of the used algorithm to hyperspectral images segmentation. …
EarlyStopping package helps prevent overfitting in iterative learning procedures.
Variable selection plays an important role in the high-dimensional data analysis. However the high-dimensional data often induces the strongly correlated variables problem. In this paper, we propose Elastic Net procedure for partially linear models and prove the group effect of its estimate. By a simulation study, we s…
Two novel procedures track quantiles efficiently using an oracle.
Boundary properties of hyperbolic groups are invariant under a maximization procedure.
In the present paper we discuss the cabling procedure for the colored HOMFLY polynomial. We describe how it can be used and how one can find all the quantities such as projectors and -matrices, which are needed in this procedure. The constructed matrix forms of the projectors and the fundamental $\mathcal{…
Proposes cost-sensitive feature selection for SVMs.