Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

1223 · May 202519922001200920172026
23 results for biostatistics

A new imputation method MissARF uses adversarial random forests for fast and accurate missing value imputation.

problem Handling missing values in biostatistical analyses.
method Adversarial Random Forests (ARF) for density estimation and data synthesis.
result MissARF performs comparably to state-of-the-art methods in imputation quality and runtime.

A new VAE model identifies and estimates treatment effects with limited overlap.

problem Identifying and estimating treatment effects when subjects with certain features belong to a single treatment group.
method Developed a latent variable model to estimate a prognostic score, which is sufficient for treatment effects. The model is a new type of VAE called β-Intact-VAE.
result The model identifies individualized treatment effects and provides TE error bounds.

Bayesian Beta regression for proportions in high dimensions with theoretical guarantees.

problem Modeling bounded continuous responses in high-dimensional settings with theoretical guarantees.
method Proposes a Bayesian approach using a tempered posterior with Horseshoe prior for shrinkage and variable selection.
result Demonstrates improved estimation accuracy and model interpretability in high-dimensional scenarios.

New method for estimating and optimizing MDPs without stationarity.

problem Challenges in offline contextual MDP estimation without stationarity.
method Introduces a new adaptive estimation and cost optimization approach for contextual MDPs.
result First robust, theoretically backed method for offline contextual MDP estimation.

Study adaptive clinical trial methods for identifying patient subpopulations with treatment benefit.

problem Adaptive identification of patient subpopulations with treatment benefit in clinical trials.
method Proposes AdaGGI and AdaGCPI meta-algorithms for subpopulation construction.
result Empirical investigation of AdaGGI and AdaGCPI performance across various simulation scenarios.

New deep Cox mixture model improves survival analysis performance.

problem Challenges in survival analysis due to censoring and healthcare applications.
method Learning mixtures of Cox regressions with deep neural networks for hazard ratios and non-parametric baseline hazard.
result Our approach outperforms classical and modern survival analysis methods, especially in minority demographics.

New methods for selecting variables in complex biomedical data.

problem Selecting important variables in multivariate, functional, and complex biomedical data.
method Optimization-based variable selection methods for various regression models.
result Outperforms state-of-the-art methods in accuracy and speed.

In the recent years more and more high-dimensional data sets, where the number of parameters pp is high compared to the number of observations nn or even larger, are available for applied researchers. Boosting algorithms represent one of the major advances in machine learning and statistics in recent years and are su…

2017-02-10abs ↗pdf ↗

The paper provides rigorous guarantees for m-out-of-n bootstrap estimators of sample quantiles.

problem Lack of parameter-free guarantees for robust inference with heavy-tailed data.
method Central limit theorem and Edgeworth expansion for m-out-of-n bootstrap estimators of sample quantiles.
result Established rigorous guarantees for the soundness of m-out-of-n bootstrap estimators of sample quantiles.

The paper uses graph learning to detect valid instruments in high-dimensional data for house pricing.

problem Endogeneity bias and invalid instrument validation in high-dimensional data.
method Merge variable selection algorithms and probabilistic graphs to estimate house prices and causal structure.
result Efficient data-driven instrument selection and invalid instrument purge in high-dimensional data.

Improved IV estimates by weighting on compliance reduces noise in treatment effect estimation.

problem Noisy IV estimates in settings with non-random treatment receipt.
method Weighting observations by estimated compliance, leveraging machine learning for compliance estimation.
result Compliance weighting reduces IV variance, improving precision of treatment effect estimates.