Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

0.5%0.9%1.4%1.9% · Apr 202619922001200920172026
48 results for cross-fitted workflow

DR-FRL learns functional states from irregular histories for causal inference.

problem Causal inference with irregularly sampled longitudinal data.
method DR-FRL workflow combining functional and temporal encoders, nuisance heads, and EIF-targeted validation.
result DR-FRL can improve causal inference when pseudo-outcomes are heavy-tailed or measurement is informative.

Improved estimators for causal inference using cross-fitting and undersmoothing.

problem Estimating expected conditional covariance in causal inference.
method Double cross-fit doubly robust (DCDR) estimators with undersmoothing for non-smooth nuisance functions.
result DCDR estimators achieve n\sqrt{n}-consistency and asymptotic normality under minimal conditions.

Paper develops efficient DML estimators for multiway clustered data without cross-fitting.

problem Efficient inference in models with multiway clustered dependence.
method Neyman-orthogonal moment conditions combined with localisation-based empirical process approach.
result Valid inference achieved without cross-fitting, showing debiased GMM estimators are asymptotically linear and normal.

Develops a test for conditional local independence of counting processes.

problem Testing the hypothesis of conditional local independence among continuous time stochastic processes.
method Introduces a new functional parameter, the Local Covariance Measure (LCM), and proposes a test called (X)-LCT using nonparametric estimators and sample splitting or cross-fitting.
result The (X)-LCT test can be controlled uniformly with modest rates, and it works well without restrictive parametric assumptions.

Proposes a method to make statistical inferences robust in spatially dependent settings with missing at random labels.

problem Statistical inference challenges with missing at random labels and spatial dependence.
method Doubly robust estimator with cross-fit nuisances and jackknife spatial HAC variance correction.
result Asymptotically valid confidence intervals with improved finite-sample calibration.

Serverless cloud computing speeds up double machine learning model estimation.

problem Efficiently estimating double machine learning models with minimal cloud resource management.
method Serverless computing with AWS Lambda for repeated cross-fitting.
result Demonstrates significant reduction in estimation times and costs.

New method stabilizes machine learning predictions across random seeds.

problem Machine learning predictions vary across random seeds, causing instability.
method Introduces adaptive cross-bagging to eliminate seed dependence.
result Adaptive cross-bagging achieves targeted stability in debiased machine learning.

Proposes a new estimator for weak instrumental variables in panel data models.

problem Weak instrumental variables due to ignored nonlinearities in panel data.
method Triangular simultaneous equation model with a nonlinear reduced form equation and a control function approach using Super Learner.
result The proposed SLCF estimator is consistent and asymptotically normal, achieving a parametric rate of convergence.

Paper presents a workflow for reliable unsupervised learning in science.

problem Lack of standardization in unsupervised learning workflows for reproducible scientific discoveries.
method Structured workflow including data preparation, modeling, validation, and communication.
result Illustrates the importance of validation in unsupervised learning.

This paper analyzes machine learning workflows in climate modeling.

problem Challenges in integrating machine learning with climate modeling.
method Analysis of case studies focusing on design patterns and workflow structure.
result Synthesis of workflow design patterns across diverse projects in ML-enabled climate modeling.

Python package automates causal parameter estimation using Riesz regression.

problem Efficient estimation of causal and structural parameters.
method Automatic DML and generalized Riesz regression framework.
result Automatic construction of balancing link functions for generalized Riesz regression.

Generative AI agents improve ERP systems by automating complex financial tasks.

problem Static, rule-based workflows limit adaptability and intelligence in ERP systems.
method Introducing Generative Business Process AI Agents (GBPAs) that integrate generative AI with business process modeling and multi-agent orchestration.
result GBPAs achieve up to 40% reduction in processing time and 94% drop in error rate.

Researchers found that avoiding synthetic data generation prevents model collapse in machine learning.

problem Model collapse in machine learning where models degenerate over generations.
method Comparing discard and augment workflows, focusing on Linear Regression.
result Theoretical evidence shows that for Linear Regression, test risk is bounded by π²/6 of original data alone.

Study efficient inference for network quantile causal effects with partial interference.

problem Estimating network causal effects on outcome quantiles with partial interference.
method Developed a nonparametric efficiency theory and a nonparametrically efficient estimator using a three-way cross-fitting procedure.
result Proposed estimator is consistent, asymptotically normal, and allows flexible estimation of nuisance functions.

A new method for releasing AI workflows to avoid premature incorrect results.

problem Statistical challenges in releasing AI workflows with adaptive scoring.
method Wrapper that calibrates and accumulates evidence from high-scoring failures.
result Reduces premature incorrect release while still releasing on moderate evidence.

CPCR mitigates bias in PCR for overparameterized models.

problem Bias in Principal Component Regression (PCR) for overparameterized models.
method Calibrated Principal Component Regression (CPCR) learns a low-variance prior in the PC subspace and calibrates the model in the original feature space.
result CPCR outperforms standard PCR in overparameterized settings, improving prediction across multiple problems.

Deep learning model improves seismic rock property estimation.

problem Estimating reservoir rock properties from seismic reflection data.
method Proposes a deep learning-based seismic inversion workflow that models seismic traces spatiotemporally.
result Achieves best performance on SEAM dataset with r2r^{2} coefficient of 79.77\%

Deep learning speeds up pressure prediction in carbon storage reservoirs.

problem Accurately forecasting reservoir pressure in geologic carbon storage projects with sparse well data.
method Combining InSAR surface displacement data with deep learning and data assimilation techniques.
result Workflow can predict reservoir pressure with high efficiency and uncertainty quantification.

MEC improves efficiency and robustness in semi-supervised inference.

problem Efficient inference with limited labeled data and robust uncertainty quantification.
method Machine-Learning-Assisted Generalized Entropy Calibration (MEC) using cross-fitted, calibration-weighted PPI.
result MEC achieves semiparametric efficiency bounds under weaker assumptions and provides near-nominal coverage.

Paper proposes an active learning method for surgical workflow recognition using long-range temporal dependency.

problem Challenges in automatic surgical workflow recognition due to lack of large-scale labelled datasets.
method NL-RCNet with non-local block for capturing long-range temporal dependency and intra-clip dependency score for selection.
result Our approach outperforms state-of-the-art methods by selecting only 50% of samples for training.

This thesis builds a real-time VaR calculation workflow for crypto derivatives.

problem Managing risk in volatile cryptocurrency markets.
method Applied EMWA, GARCH, and HAR models to forecast volatility; used delta-gamma-theta approach and Cornish-Fisher expansion.
result Real-time VaR estimates with millisecond calculation latencies.

Study uses non-invasive sensors to assess microbial contamination in green salads.

problem Assessing microbial contamination in ready-to-eat green salads.
method Unified spectra analysis workflow with data normalization, feature selection, and partial least squares regression.
result Unified workflow provides efficient estimation of microbial contamination and shelf life.

Data application developers and data scientists spend an inordinate amount of time iterating on machine learning (ML) workflows -- by modifying the data pre-processing, model training, and post-processing steps -- via trial-and-error to achieve the desired model performance. Existing work on accelerating machine learni…

2018-08-03abs ↗pdf ↗

GSR optimizes tasks in scientific workflows, improving performance across diverse applications.

problem Uncertainty in task selection and evaluation in scientific workflow optimization.
method Generate-Select-Refine (GSR) framework that alternates between task generation and optimization.
result GSR outperforms existing LLM-based optimizers in various scientific applications.

Study evaluates deep learning methods for dermatology, finding they perform poorly under non-ideal conditions.

problem Lack of robustness of deep learning methods in dermatology under real-world conditions.
method Simulated non-ideal conditions on user-submitted dermatology images.
result Deep learning methods show significant drop in accuracy and prediction changes under non-ideal conditions.

Proposes a method for inference in high-dimensional classification with non-differentiable surrogate losses.

problem Lack of inference procedures for identifying driving factors in high-dimensional classification with non-differentiable surrogate losses.
method Kernel-smoothed decorrelated score and cross-fitted version for hypothesis tests and interval estimators.
result Valid and superior inference methods for high-dimensional classification with non-differentiable surrogate losses.

Neyman's framework evaluates personalized treatment rules using experiments.

problem Evaluating the efficacy of individualized treatment rules derived by machine learning.
method Neyman's repeated sampling framework applied to cross-fitted ITRs.
result Ex-post evaluation of ITRs can be more efficient than random assignment.

Proposes a new estimator for causal mediation with continuous treatments.

problem Estimation of direct and indirect effects with continuous treatments.
method Kernel smoothing approach with cross-fitting for non-parametric estimation.
result Multiply robust and asymptotically normal estimator for continuous treatments.

Foundation models alter medical data science workflow, challenging veridical data science principles.

problem Foundation models disrupt traditional data science practices in medicine.
method Critically examined the medical foundation model lifecycle and its deviation from veridical data science principles.
result Foundation models challenge veridical data science principles of predictability, computability, and stability.