Responds to critiques on tests for causal parameter confidence intervals.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Bagging stabilizes models without distributional assumptions.
An assumption-free automatic check of medical images for potentially overseen anomalies would be a valuable assistance for a radiologist. Deep learning and especially Variational Auto-Encoders (VAEs) have shown great potential in the unsupervised learning of data distributions. In principle, this allows for such a chec…
Bounds and sensitivity analysis for causal effects with MNAR confounders.
For many causal effect parameters of interest, doubly robust machine learning (DRML) estimators are the state-of-the-art, incorporating the good prediction performance of machine learning; the decreased bias of doubly robust estimators; and the analytic tractability and bias reduction of sample splitting wi…
We characterize distributional equivalence in latent-variable models with cycles.
We propose a novel method for clustering data which is grounded in information-theoretic principles and requires no parametric assumptions. Previous attempts to use information theory to define clusters in an assumption-free way are based on maximizing mutual information between data and cluster labels. We demonstrate …
Often machine learning methods are applied and results reported in cases where there is little to no information concerning accuracy of the output. Simply because a computer program returns a result does not insure its validity. If decisions are to be made based on such results it is important to have some notion of th…
A method to assess sensitivity to unmeasured confounding with sharp bounds.
New method bounds hardware noise without assumptions.
New algorithm reduces offline RL sample complexity for MDPs.
Random feature mapping (RFM) is a popular method for speeding up kernel methods at the cost of losing a little accuracy. We study kernel ridge regression with random feature mapping (RFM-KRR) and establish novel out-of-sample error upper and lower bounds. While out-of-sample bounds for RFM-KRR have been established by …
NMDR estimates complex mixtures of distributions efficiently.
Choppy optimizes ranked list truncation using Transformer architecture.
Privacy concern has been increasingly important in many machine learning (ML) problems. We study empirical risk minimization (ERM) problems under secure multi-party computation (MPC) frameworks. Main technical tools for MPC have been developed based on cryptography. One of limitations in current cryptographically priva…
We give a polynomial-time algorithm for learning neural networks with one layer of sigmoids feeding into any Lipschitz, monotone activation function (e.g., sigmoid or ReLU). We make no assumptions on the structure of the network, and the algorithm succeeds with respect to {\em any} distribution on the unit ball in …
C-PP-COAD detects anomalies with limited real data, reducing dependency on real calibration data.
When estimating finite mixture models, it is common to make assumptions on the mixture components, such as parametric assumptions. In this work, we make no distributional assumptions on the mixture components and instead assume that observations from the mixture model are grouped, such that observations in the same gro…
Paper presents a machine learning method to improve significance tests for misspecified linear models.
Adaptive PI by reweighting nonconformity scores improves model uncertainty reflection.
RQR improves prediction intervals for skewed data.
Given only positive (P) and unlabeled (U) data, PU learning can train a binary classifier without any negative data. It has two building blocks: PU class-prior estimation (CPE) and PU classification; the latter has been well studied while the former has received less attention. Hitherto, the distributional-assumption-f…
Proposes new methods for inference in GLMs without assuming model correctness.
Study agnostic feature-based dynamic pricing models with linear policies and noisy valuations.
We study the dynamic assortment planning problem, where for each arriving customer, the seller offers an assortment of substitutable products and customer makes the purchase among offered products according to an uncapacitated multinomial logit (MNL) model. Since all the utility parameters of MNL are unknown, the selle…
New methods provide stable ranking without assumptions on data distributions.
Spofe bridges statistical rigor and interpretability in feature extraction from tabular data.
Efficiently calculates privacy guarantees for 2020 Census data.
Study limits of testing algorithms without assumptions, finding key performance bounds.