Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

2795588361,115 · Jun 202019922001200920172026
48 results for RCT data

Method uses pseudo-samples to improve RCT data in ride-hailing pricing studies.

problem Small and biased RCT data leads to significant bias when generalizing to broader user base.
method Pseudo-sample matching to expand and match RCT data with observational data.
result 0.41% improvement in profit through pseudo-sample matching.

CVIB uses information theory to learn counterfactuals from MNAR data without RCTs.

problem Debiasing learning from missing-not-at-random (MNAR) data in recommendation systems.
method CVIB, a variational information bottleneck, separates task-aware mutual information into factual and counterfactual parts.
result CVIB significantly enhances both shallow and deep models in recommendation systems.

The paper benchmarks OS with RCTs, accounting for right-censoring.

problem Benchmarking observational studies with experimental data under censoring.
method Two cases: independent and dependent censoring. Censoring-doubly-robust signal for CATE.
result Effectiveness of censoring-aware tests verified via experiments and real data.

Paper proposes a novel method to assess treatment effect estimators using cross-validation.

problem Lack of ground truth to objectively assess treatment effect estimators in RCTs.
method Cross-validation-like methodology combining noisy difference-of-means estimate and aggregation across RCTs.
result Aggressive downweighting or truncation of large values reduces variance and improves treatment effect estimation.

Our paper improves uplift model evaluation on randomized controlled trials (RCT) data.

problem Variance in uplift evaluation metrics makes their signals arbitrary and unreliable.
method Theoretical analysis and statistical adjustment of the outcome to reduce variance.
result Variance reduction methods improve uplift evaluation metrics on RCT data.

A new method improves treatment effect inferences in RCTs by adjusting for covariates and heteroskedasticity.

problem Improving treatment effect inferences in RCTs with efficient and powerful methods.
method Weighted Prognostic Covariate Adjustment Method (Weighted PROCOVA) for heteroskedasticity.
result The method reduces variance, maintains Type I error rate, and increases test power for treatment effect.

Machine learning boosts RCT efficiency by controlling type I error and improving statistical power.

problem Improving statistical efficiency in RCTs with complex covariate adjustments.
method Machine learning-assisted adjustment under Rosenbaum's framework for exact tests.
result The proposed method robustly controls type I error and significantly boosts statistical efficiency.

New estimator improves policy evaluation in resource allocation RCTs.

problem Difficulty in evaluating policies optimizing limited resource allocation through RCTs.
method Proposes a novel estimator involving retrospective reshuffling of participants across experimental arms.
result The new estimator provides more accurate policy evaluations than common methods.

Transformer learns shipping costs more accurately than traditional methods.

problem Inaccurate shipping cost estimates lead to poor financial decisions.
method Proposes Rate Card Transformer (RCT) using self-attention to encode shipping information.
result Cost predictions made by RCT have 28.82% less error compared to GBDT models.

Prognostic scores improve logistic regression analysis in RCTs with binary outcomes.

problem Non-collapsibility in logistic regression analysis of RCTs with binary endpoints.
method Prognostic score adjustment using AI predictions to address non-collapsibility.
result Prognostic score adjustment increases power or reduces sample size for estimating conditional odds ratios.

New study finds targeting based on treatment effects outperforms risk-based targeting in social interventions.

problem Lack of accurate treatment effect estimates for machine learning-based targeting in social domains.
method Empirical assessment of targeting strategies using data from 5 real-world RCTs in various domains.
result Treatment effect-based targeting outperforms risk-based targeting, even with biased estimates.

In this paper we present tools for applied researchers that re-purpose off-the-shelf methods from the computer-science field of machine learning to create a "discovery engine" for data from randomized controlled trials (RCTs). The applied problem we seek to solve is that economists invest vast resources into carrying o…

2017-07-05abs ↗pdf ↗

New methods improve causal inference generalization using trial and observational data.

problem Limited trial data makes generalizing causal inferences to target populations statistically infeasible.
method Develops algorithms that combine trial and observational data to estimate complex nuisance functions.
result Improves generalization of causal inferences when the additional observational study is high-quality.

New findings show single-treatment effects are unidentifiable in factorial experiments.

problem Identifying the effect of a single intervention in factorial experiments.
method Formalized sufficient conditions for the identifiability of single-treatment effects and developed nonparametric sharp bounds.
result Researchers must justify assumptions for extrapolating single-treatment effects.

Deep neural networks (DNNs) are incredibly brittle due to adversarial examples. To robustify DNNs, adversarial training was proposed, which requires large-scale but well-labeled data. However, it is quite expensive to annotate large-scale data well. To compensate for this shortage, several seminal works are utilizing l…

2019-11-20abs ↗pdf ↗

Paper introduces EnCounteR for estimating causal effects using encouragement data.

problem Challenges in estimating causal effects due to incomplete randomization and limited encouragement data.
method Introduces a generalized IV estimator, EnCounteR, leveraging both observational and encouragement data.
result Demonstrates superior performance of EnCounteR over existing methods.

This study quantifies uncertainty in comparing treatments using RCTs with before-and-after measures.

problem Uncertainty in comparing treatments using RCTs with before-and-after measures.
method New statistical modeling principle called ETZ enables counterfactual uncertainty quantification (CUQ) in RCTs with Before-and-After Repeated Measures.
result CUQ typically has lower variability than factual uncertainty quantification and can be achieved in RCTs.

Paper introduces multi-scale methods to improve CATE estimation from EO data.

problem Challenges in balancing fine-grained and contextual information in EO-based causal inference.
method Multi-Scale Representation Concatenation, combining Vision Transformer and Causal Forests.
result Multi-scale approach captures effect heterogeneity better than single-scale models.

Novel approach to compute hazard ratios from observational studies using SCMs and backdoor adjustment.

problem Identifying causal relationships from observational data using hazard ratios.
method Backdoor adjustment through structural causal models (SCMs) and do-calculus.
result Novel approach for computing hazard ratios from observational studies.

Causal graph aids observational study insights in aSAH patients.

problem Lack of clear objectives and tools for identifying necessary adjustments in observational studies.
method Uses causal directed acyclic graphs (DAGs) to provide insights mid-study and identify necessary data enhancements.
result Midway insights and necessary data enhancements identified for meaningful causal questions.

New method uses latent variables to estimate treatment effects from single-arm trials.

problem Estimating treatment effects from single-arm trials due to lack of external control groups.
method Latent-variable modeling with amortized variational inference for patient matching and direct effect estimation.
result Improved performance in direct treatment effect estimation and effect estimation via patient matching compared to previous methods.

Machine learning experiments often mislead due to unmet assumptions.

problem Machine learning experiments with pooled data may not meet necessary assumptions for unbiased causal effect estimation.
method Analysis of assumptions required for unbiased causal effect estimation in machine learning experiments.
result Practical applications of A/B-tests with machine learning models may not yield unbiased estimates of causal effect.

CausalSim corrects bias in trace-driven simulations for more accurate results.

problem Bias in trace-driven simulations due to system conditions during trace collection.
method CausalSim learns a causal model of system dynamics and latent factors from an RCT to remove bias from trace data.
result CausalSim reduces simulation errors by 53% and 61% compared to baselines, providing more accurate insights.

New method estimates optimal personalized treatment rules from mixed data sources.

problem Combining RCT and observational data for personalized treatment rules.
method Doubly robust estimator for value function, maximizing within pre-specified ITR class.
result Consistent and asymptotically normal optimal value estimator with N1/3N^{-1/3} rate of convergence.

Study compares Cox model and RSF for predicting patient survival, finding RSF superior in certain scenarios.

problem Comparing predictive accuracy of Cox proportional hazards model and Random Survival Forest for patient-specific survival probabilities.
method Conducted a comprehensive comparison study using simulation scenarios and real-world datasets.
result RSF outperforms Cox model in nonproportional hazards settings and with treatment-covariate interactions.

The paper proposes a method to find subgroups with significant treatment effects in noisy data.

problem Estimating the causal effects of interventions on noisy outcomes.
method A machine-learning method specifically optimized for finding subgroups with significant effects, designed to maximize the probability of obtaining a statistically significant positive treatment effect.
result The proposed method yields higher power in detecting subgroups affected by the treatment compared to standard tree-based tools.

CausalBench aims to advance causal learning research with a transparent platform.

problem Lack of unified benchmark datasets, algorithms, metrics, and evaluation interfaces for causal learning.
method Introduces CausalBench, a flexible benchmark framework for causal analysis and machine learning.
result Promotes scientific collaboration, reproducibility, and awareness in causal learning research.

New method uses observational data to improve trial design efficiency.

problem Scarce randomized controlled trials; inefficiency of using observational data.
method Active Residual Learning, R-Design framework, R-EPIG criterion.
result Efficiently estimating residuals to correct observational bias improves trial design.

Q-Learner estimates ratio-based treatment effects without imposing parametric structures.

problem Estimating treatment effects as ratios in non-linear settings.
method Decomposes ratio-CATE into two classification tasks, using doubly robust augmentations.
result Q-Learner outperforms other methods in low-conversion and observational data settings.

New estimator improves ATT estimation efficiency with external controls.

problem Reduced efficiency when incorporating external controls into ATT estimation.
method Proposes a novel doubly robust estimator for ATT that maintains higher efficiency than standard approaches.
result Demonstrates improved efficiency of the new estimator compared to standard approaches, even under model misspecification.

G-computation improves clinical trial power with machine learning.

problem Balancing prognostic factors in randomized trials to prevent near-confounders.
method G-computation with penalized models (Lasso, Elasticnet) and algorithm-based methods (neural network, SVM, super learner).
result G-computation with Elasticnet and splines reduces variance and increases power in RCTs.

A new method flips class values to address class and treatment imbalance in uplift modeling and HTE.

problem Class and treatment imbalance in imbalanced RCT data.
method Class flipping approach to address imbalance without distorting predictions.
result The method does not distort predicted effects and does not require calibration.

New method minimizes decision errors in large treatment spaces.

problem Improving decision-making in large treatment spaces with biased observational data.
method Loss minimizes classification error of actions in large action space.
result Proves improved decision-making performance in large combinatorial action spaces.

Synthetic control method improves policy evaluation in high-dimensional settings.

problem Evaluating the impact of new policies in large-scale applications.
method Two-phase approach: nearest neighbor matching followed by supervised learning.
result The method successfully improves estimate accuracy in large-scale experiments.