Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

12.5%25.0%37.5%50.0% · Dec 199319922001200920182026
48 results for double averaging

A new algorithm speeds up multi-agent reinforcement learning.

problem Complex interactions between agents in multi-agent reinforcement learning.
method Double averaging scheme for decentralized convex-concave saddle-point problems.
result The algorithm converges to the optimal solution at a global geometric rate.

The paper investigates how calibrating propensity scores improves DML estimates of average treatment effects.

problem Improving the accuracy of DML estimates in finite samples.
method Propensity score calibration within the Double/debiased machine learning framework.
result Calibrating propensity scores reduces the root mean squared error of DML estimates of average treatment effects in finite samples.

Researchers develop a method to measure treatment effects in settings with shared states.

problem Measuring treatment effects in settings with shared states like prices, recommendations, or social signals.
method Double machine learning (DML) theorem with conditions for efficient inference under shared-state interference.
result Efficient estimation of average direct effect (ADE) and global average treatment effect (GATE) in various models.

Extends double linear policy with time-varying weights and proves robust positive expectation.

problem Ensuring robustness in policy optimization with time-varying parameters.
method Employed a novel elementary symmetric polynomials characterization approach to prove robust positive expectation (RPE). Derived explicit expressions for expected cumulative gain-loss and variance.
result Proved the robust positive expectation property holds for the extended double linear policy.

High-dimensional models can outperform simpler ones in causal inference.

problem Estimating average treatment effects with many covariates.
method High-dimensional linear regression and synthetic control with many control units.
result Adding more control units can improve imputation performance even when pre-treatment fit is perfect.

Unified theory explains and mitigates double descent in data reconstruction.

problem Understanding and mitigating double descent in reduced order modeling.
method Data-Noise Averaging theory, sufficient criteria, detailed risk curve prediction, regularization mechanisms.
result Detailed risk curves predicted at reduced computational cost, instability traced to individual sensors.

Chernozhukov, Chetverikov, Demirer, Duflo, Hansen, and Newey (2016) provide a generic double/de-biased machine learning (DML) approach for obtaining valid inferential statements about focal parameters, using Neyman-orthogonal scores and cross-fitting, in settings where nuisance parameters are estimated using a new gene…

2017-01-30abs ↗pdf ↗

Double Q-learning has the same mean-squared error as Q-learning under certain conditions.

problem Comparing the mean-squared error of Double Q-learning and Q-learning.
method Theoretical analysis based on Lyapunov equations for both tabular and linear function approximation settings.
result The asymptotic mean-squared error of Double Q-learning is exactly equal to that of Q-learning under specific conditions.

The paper analyzes the effectiveness of principal component regression with varying numbers of features.

problem The study examines the performance of principal component regression with different numbers of selected features.
method The analysis considers the least squares linear regression over uncorrelated Gaussian features selected in decreasing variance order. It also analyzes the prediction error in the average-case setting as the number of features and samples grow.
result The prediction error shows a 'double descent' shape as the number of features increases, and conditions are established for achieving minimum risk in the interpolating regime.

The paper extends Hodge-de Rham and Lichnérowicz Laplacians to double forms and proves vanishing theorems.

problem Extending Laplacians to double forms and proving vanishing theorems.
method Introduced a new product on double forms to establish index-free formulas for curvature terms in Weitzenböck formulas for ΔΔ, Δ~\widetildeΔ, and ΔLΔ_L. Proved vanishing theorems for ΔΔ and ΔLΔ_L on symmetric double forms.
result Vanishing theorems for the Hodge-de Rham Laplacian and ΔLΔ_L on symmetric double forms.

New methods combine machine learning with doubly robust estimators for better treatment effect estimation.

problem Estimating average treatment effects from observational data.
method Doubly robust methods using machine learning techniques.
result Machine learning improves the performance of doubly robust estimators.

We develop methods to approximate derivatives for causal inference problems using data.

problem Estimating causal effects from data when distributions are not known.
method Constructive algorithm approximating Gateaux derivatives via finite differencing.
result Derives conditions for finite-difference approximations to preserve statistical benefits.

FIDDLE uses deep learning to estimate ATE from complex data.

problem Estimating ATE from high-dimensional, correlated covariates with sparse nonlinear effects.
method Factor-augmented deep learning for propensity and outcome models.
result FIDDLE consistently estimates ATE under model misspecification and is semiparametrically efficient.

The paper analyzes MACD using operator theory.

problem Understanding the mathematical foundation of MACD.
method Developed a functional-analytic framework interpreting MACD as a phase-corrected, smoothed derivative operator.
result MACD is structurally equivalent to a band-pass filter and can be expressed as a finite difference of delayed and doubly averaged signals.

Efficient method for pricing Bermudan moving average options using GPR-GHQ.

problem High-dimensional pricing of Bermudan moving average options in energy markets.
method Gaussian Process Regression and Gauss-Hermite quadrature.
result GPR-GHQ method efficiently handles long windows and high dimensionality.

Proposes efficient estimators for weighted cumulative treatment effects in observational studies.

problem Inconsistent and inefficient estimators due to model misspecification and lack of overlap.
method Double/debiased machine learning for weighted cumulative causal effects.
result Proposed estimators are consistent, asymptotically linear, and reach semiparametric efficiency bounds.

New rule reduces exploration regret to logarithmic, improving bad episode handling.

problem Improving exploration regret in average reward MDPs.
method Replacing Doubling Trick with Vanishing Multiplicative rule in EVI-based algorithms.
result Regret is logarithmic under the new rule, significantly better than linear.

A semiparametric test evaluates instrument validity and complier characteristics.

problem Evaluating the validity of instruments and complier characteristics.
method Semiparametric test, doubly robust moment, machine learning update.
result Validates instrument validity and complier characteristics.

New algorithm for solving minimax problems over distributions converges to Nash equilibrium.

problem Solving minimax problems over probability distributions.
method Symmetric Mean-field Langevin Dynamics (MFL-AG and MFL-ABR) with weighted averaging and best response dynamics.
result Converges to mixed Nash equilibrium with average-iterate and last-iterate convergence.

New algorithm reduces reinforcement learning regret to sqrt(T) without strong dynamics assumptions.

problem Infinite-horizon average-reward reinforcement learning with linear MDPs.
method Approximate by discounted-reward MDPs and apply optimistic value iteration.
result Achieves O(sqrt(T)) regret with polynomial complexity.

Co-DQL improves traffic signal control using multi-agent reinforcement learning.

problem Optimizing signal timing for large-scale traffic control.
method Cooperative double Q-learning (Co-DQL) with mean field approximation and reward allocation.
result Co-DQL reduces average waiting time for vehicles in the road system.

We present an experimental and simulated model of a multi-agent stock market driven by a double auction order matching mechanism. Studying the effect of cumulative information on the performance of traders, we find a non monotonic relationship of net returns of traders as a function of information levels, both in the e…

2006-10-04abs ↗pdf ↗

This paper investigates robust and efficient DR/RDR estimators for WATEs.

problem Lack of systematic investigation into robustness and efficiency conditions for WATE estimation.
method Proposes three RDR estimators using semiparametric efficient influence function and double/debiased machine learning.
result Demonstrates the practical relevance of the methods in medical and social sciences.

SHIFT improves robustness in estimating dose-response functions with heavy-tailed contamination.

problem Outliers bias estimates of average dose-response functions in heavy-tailed data.
method SHIFT combines cross-fit nuisance orthogonalization, Welsch-loss, and defensive OLS refit.
result SHIFT reduces RMSE from 1.03 to 0.33 on localized contamination test.

Post-hoc transforms can reverse model performance trends, especially in noisy settings.

problem Post-hoc transforms can reverse model performance trends, especially in noisy settings.
method Empirical study and analysis of post-hoc transforms like temperature scaling, ensembling, and SWA.
result Post-hoc reversal can prevent double descent and mitigate mismatches between test loss and test error.

Proposes a method to estimate causal effects of continuous treatments using instrumental variables.

problem Estimating causal effects of continuous treatments in the presence of unmeasured confounders.
method Introduces a novel framework using instrumental variables and a uniform regular weighting function to identify and estimate average dose-response functions.
result Establishes the asymptotic properties of the proposed methods for estimating average dose-response functions.

Most modern supervised statistical/machine learning (ML) methods are explicitly designed to solve prediction problems very well. Achieving this goal does not imply that these methods automatically deliver good estimators of causal parameters. Examples of such parameters include individual regression coefficients, avera…

2016-07-30abs ↗pdf ↗

Deep networks generalize well even when they fit training data perfectly, thanks to overparametrization.

problem Understanding generalization in overparametrized deep networks.
method Random features regression, asymptotic analysis, ensemble averaging.
result Bias remains constant beyond the interpolation threshold, while variance components decay with overparametrization.

Estimates causal effects from a single time-series without further assumptions.

problem Estimating causal effects from a single time-series without additional assumptions.
method Proposes a general class of averages of conditional causal parameters, estimated using a targeted maximum likelihood estimator (TMLE).
result Asymptotic consistency and normality of the TMLE for estimating causal parameters.

Paper estimates non-causal graphical models using covariance extension and transportation distance.

problem Estimating non-causal graphical models with smoothing relations.
method Proposes a covariance extension problem and uses transportation distance to minimize error with white noise.
result Solution is a double-sided autoregressive non-causal graphical model.

Multi-label classification is a type of supervised learning where an instance may belong to multiple labels simultaneously. Predicting each label independently has been criticized for not exploiting any correlation between labels. In this paper we propose a novel approach, Nearest Labelset using Double Distances (NLDD)…

2017-02-15abs ↗pdf ↗

Proposes a scalable method for counterfactual prediction using machine learning.

problem De-bias causal estimators with high-dimensional data in observational studies.
method Uses entropy balancing to learn weights minimizing Jensen-Shannon divergence, leading to robust counterfactual predictions.
result Consistent causal estimation if either propensity score or outcome model is correctly specified.

New estimator improves ATT estimation efficiency with external controls.

problem Reduced efficiency when incorporating external controls into ATT estimation.
method Proposes a novel doubly robust estimator for ATT that maintains higher efficiency than standard approaches.
result Demonstrates improved efficiency of the new estimator compared to standard approaches, even under model misspecification.