Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

0.7%1.4%2.1%2.8% · Apr 202619922001200920172026
48 results for Doubly Debiased LASSO

Study improves statistical inference for CATEs using Lasso and DML.

problem Estimating and inferring CATEs in high-dimensional settings.
method Doubly robust estimator, Lasso regularization, debiased Lasso, DML.
result TDL (triple/debiased Lasso) achieves n\sqrt{n}-consistency and confidence intervals.

DWTS uses observational data to improve clinical trial efficiency.

problem Lack of definitive conclusions from randomized clinical trials due to insufficient patient cohorts and confounding biases.
method DWTS combines observational data with randomized clinical trials using Doubly Debiased LASSO (DDL) to identify reliable covariates.
result DWTS reduces cumulative regret in clinical trials compared to standard methods.

This paper introduces metrics for welfare analysis in dynamic models. We develop estimation and inference for these parameters even in the presence of a high-dimensional state space. Examples of welfare metrics include average welfare, average marginal welfare effects, and welfare decompositions into direct and indirec…

2019-08-24abs ↗pdf ↗

Ridge regression is revisited with debiasing and thresholding, offering advantages over Lasso.

problem High-dimensional data challenges classical ridge regression's sparsity detection and bias issues.
method Debiasing and thresholding ridge regression, introducing a wild bootstrap for confidence regions and hypothesis testing, and a hybrid bootstrap for prediction intervals.
result Debiased and thresholded ridge regression can offer similar performance to thresholded Lasso and may be preferable in some settings.

Proposes online debiasing to correct bias in adaptive data collection for high-dimensional linear regression.

problem Bias in adaptive data collection for high-dimensional linear regression.
method Online debiasing procedure for LASSO and other estimators.
result Optimal debiasing of LASSO estimator in specific sparsity regime.

Two approaches to directly estimating Riesz representer are shown to be numerically equivalent under certain conditions.

problem Estimating Riesz representer in semiparametric statistics.
method Two distinct optimization problems solved by automatic debiased machine learning and sieve methods for conditional moment models.
result Numerical equivalence of estimators under specific regularization schemes, but not for others.

Paper proposes methods to reduce bias and variance in recommender systems.

problem Bias in recommender systems due to users' preferences.
method Proposes a principled approach to reduce bias and variance in DR methods, and a novel semi-parametric collaborative learning approach.
result The proposed methods outperform existing debiasing methods in both theory and experiments.

Novel characterization of augmented balancing weights combining outcome and weighting models.

problem Improving estimation accuracy in machine learning models with balancing weights.
method Characterization of augmented balancing weights as linear models, extending to ridge and lasso regression.
result Equivalence and closed-form expressions for specific model choices, providing insights into performance.

Bayesian method corrects bias in treatment effect estimation.

problem Estimating treatment effects from observational data with high-dimensional nuisance parameters.
method Bayesian debiasing, targeted modeling, sample splitting.
result Marginal posterior for ATE satisfies Bernstein-von Mises theorem under correct nuisance model specification.

We devise a one-shot approach to distributed sparse regression in the high-dimensional setting. The key idea is to average "debiased" or "desparsified" lasso estimators. We show the approach converges at the same rate as the lasso as long as the dataset is not split across too many machines. We also extend the approach…

2015-03-14abs ↗pdf ↗

Contextual multi-armed bandit algorithms are widely used in sequential decision tasks such as news article recommendation systems, web page ad placement algorithms, and mobile health. Most of the existing algorithms have regret proportional to a polynomial function of the context dimension, dd. In many applications ho…

2019-07-26abs ↗pdf ↗

The paper analyzes LASSO penalization for high-dimensional Beta regression models.

problem Theoretical analysis of LASSO in high-dimensional Beta regression.
method Non-convexity handling through a neighborhood framework, debiasing for confidence intervals, proximal gradient algorithm.
result Non-asymptotic bound on 1\ell_1-error of stationary points.

Corrects mismatch in consistency of nuisance estimators for doubly robust methods.

problem Mismatch in consistency of nuisance estimators in doubly robust methods.
method Calibrated debiased machine learning (calibrated DML) with isotonic regression adjustment.
result Calibrated DML yields doubly robust asymptotic normality with slower convergence of nuisance estimators.

New algorithm for computing Wasserstein barycenters with guarantees.

problem Computing Wasserstein barycenters with varying regularization strengths.
method Damped Sinkhorn iterations followed by exact maximization/minimization steps.
result First non-asymptotic convergence guarantees for approximating Wasserstein barycenters.

Theory establishes optimal rates for estimating linear functionals without structural assumptions.

problem Estimating linear functionals of unknown nuisance components without structural assumptions.
method Structure-agnostic framework, doubly robust estimators, first-order debiasing.
result Characterization of minimax optimal rates and regimes for double robustness.

We consider the problem of distributed multi-task learning, where each machine learns a separate, but related, task. Specifically, each machine learns a linear predictor in high-dimensional space,where all tasks share the same small support. We present a communication-efficient estimator based on the debiased lasso and…

2015-10-02abs ↗pdf ↗

This paper investigates robust and efficient DR/RDR estimators for WATEs.

problem Lack of systematic investigation into robustness and efficiency conditions for WATE estimation.
method Proposes three RDR estimators using semiparametric efficient influence function and double/debiased machine learning.
result Demonstrates the practical relevance of the methods in medical and social sciences.

Computing partition functions, the normalizing constants of probability distributions, is often hard. Variants of importance sampling give unbiased estimates of a normalizer Z, however, unbiased estimates of the reciprocal 1/Z are harder to obtain. Unbiased estimates of 1/Z allow Markov chain Monte Carlo sampling of "d…

2016-10-15abs ↗pdf ↗

Develops a new method for uncertainty quantification in high-dimensional learning.

problem Challenges in uncertainty quantification in high-dimensional regression or learning problems.
method Data-driven approach for UQ that corrects bias terms from training data.
result Non-asymptotic confidence intervals that avoid overestimating uncertainty.

New method for estimating treatment effects without complex propensity models.

problem Estimating treatment effects in dynamic treatment regimes.
method Recursive Riesz representer estimation for de-biasing corrections.
result Directly estimates de-biasing corrections without auxiliary models.

This paper introduces a novel online inference method for high-dimensional GLMs.

problem Real-time analysis of sequentially collected data in high-dimensional settings.
method Adaptive stochastic gradient descent with online debiasing for dynamic objective functions.
result Established the asymptotic normality of the Adaptive Debiased Lasso (ADL) estimator.

ADML combines debiased learning with data-driven model selection for efficient inference.

problem Debiased machine learning estimators can be unstable and biased in nonparametric models.
method Data-driven model selection techniques combined with debiased machine learning.
result ADML estimators yield superefficient inference for pathwise differentiable parameters.

Develops a method to estimate average hazard under non-proportional hazards without relying on proportional hazards assumption.

problem Estimation of treatment effects when hazards are non-proportional, leading to unstable hazard ratios.
method Semiparametric, doubly robust framework for covariate-adjusted average hazard estimation.
result Valid sqrt{n} inference with small bias and near-nominal confidence-interval coverage across proportional and non-proportional hazards settings.

DoubleGen addresses bias in generative modeling of counterfactuals.

problem Bias in generative models for counterfactual outcomes.
method Doubly robust framework that modifies generative modeling training objectives to mitigate confounding and misspecification biases.
result Successfully addresses confounding bias even if only one auxiliary model is correct.

DebiNet uses over-parameterized neural networks to improve linear model performance and debiasing.

problem Improving linear model performance and debiasing in high-dimensional settings.
method Incorporates over-parameterized neural networks into semi-parametric models to estimate parameters consistently.
result DebiNet offers valid inference and accurate prediction by leveraging neural networks' universal approximation and linear model's interpretability.

The Lasso method is analyzed for high-dimensional regression with Gaussian designs, leading to new insights on its performance.

problem Analyzing the Lasso method for high-dimensional regression with Gaussian designs.
method Generalizing the Lasso characterization to Gaussian correlated designs with non-singular covariance structure.
result Establishing non-asymptotic bounds on the distance between the distribution of various quantities in the two models.

A new method for feature selection robust to noise and design variability.

problem Feature selection in high-dimensional regression under sampling variability and measurement error.
method Injects controlled additive noise into the design matrix, fits a base selector, and aggregates selection frequencies.
result Improved robustness compared to Stability Selection and standard base selectors.

Estimates causal effects using machine learning for binary treatment and mediator.

problem Estimating direct and indirect quantile treatment effects under selection-on-observables.
method Double/debiased machine learning estimators based on efficient score functions.
result Uniform consistency and asymptotic normality of effect estimators.

Improves robustness of propensity score estimators in challenging settings.

problem Limited overlap, small sample sizes, or unbalanced data.
method Extends calibration techniques for propensity score models, focusing on sample-splitting schemes.
result Calibration reduces variance and bias in inverse probability weighting and double/debiased machine learning frameworks.

Semiparametric method removes bias in functional bilevel gradient estimation.

problem First-order bias in plug-in hypergradient when lower-level problem is nonparametric.
method Semiparametric debiasing theory based on efficient influence function leads to cross-fitted orthogonal hypergradient estimator.
result Asymptotic normality and uniform control over outer parameter established for the estimator.

Performing statistical inference in high-dimension is an outstanding challenge. A major source of difficulty is the absence of precise information on the distribution of high-dimensional estimators. Here, we consider linear regression in the high-dimensional regime pnp\gg n. In this context, we would like to perform in…

2015-08-11abs ↗pdf ↗

High-dimensional inference for sparse spectral precision matrices

problem Inference on the spectral precision matrix at a fixed frequency
method Full likelihood-based inference using neighboring discrete Fourier transforms
result Simultaneous control of regularization, finite-sample truncation, and smoothing biases

Method estimates dynamic treatment effects using machine learning and g-estimation.

problem Estimating treatment effects over time with multiple treatments and potential future outcomes.
method Double/debiased machine learning framework for dynamic treatment effects, extending Neyman orthogonal cross-fitted gg-estimation.
result Provides finite sample guarantees and allows for non-linear effect heterogeneity and high-dimensional parameterizations.