Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

4.5%9.1%13.6%18.2% · Feb 202619922001200920172026
48 results for statistical adjustment

Machine learning boosts RCT efficiency by controlling type I error and improving statistical power.

problem Improving statistical efficiency in RCTs with complex covariate adjustments.
method Machine learning-assisted adjustment under Rosenbaum's framework for exact tests.
result The proposed method robustly controls type I error and significantly boosts statistical efficiency.

Improves trial efficiency by adjusting for historical prognostic scores.

problem Reducing statistical uncertainty in randomized trial estimates.
method Linear covariate adjustment using a prognostic model trained on historical data.
result Prognostic covariate adjustment achieves minimum variance and reduces mean-squared error.

A new method improves treatment effect inferences in RCTs by adjusting for covariates and heteroskedasticity.

problem Improving treatment effect inferences in RCTs with efficient and powerful methods.
method Weighted Prognostic Covariate Adjustment Method (Weighted PROCOVA) for heteroskedasticity.
result The method reduces variance, maintains Type I error rate, and increases test power for treatment effect.

Quantum tech speeds up financial risk assessment.

problem Improving credit valuation adjustments using quantum mechanics.
method Developed a quantum algorithm using Bayesian quantum amplitude estimation and engineered likelihood functions.
result Significant speedup in quantum computations for CVA over classical methods.

New methods for calculating credit valuation adjustment with reduced noise and faster computation.

problem High statistical noise in computing sensitivities of CVA due to non-differentiable default intensities.
method Ad hoc analytical estimators to overcome non-differentiability and finite differences.
result Low statistical noise and fast computation of sensitivities to market quotes.

Paper tackles long-tailed labels in classification problems.

problem Imbalanced or long-tailed label distribution in real-world classification problems.
method Logit adjustment applied post-hoc or during training to encourage a large relative margin between rare and dominant labels.
result Unified and generalised techniques for coping with long-tailed labels, improving generalisation and performance.

Develops a new model to better estimate cryptocurrency and stock volatility.

problem Misrepresentation of volatility and co-movement in traditional models.
method Introduces liquidity-sensitive multivariate volatility framework with novel liquidity measures.
result Liquidity-adjusted models yield more stable and interpretable risk structures.

This paper improves MDS visualization by adjusting Wasserstein distances for heavy-tailed data.

problem Enhancing Multidimensional Scaling (MDS) for better pattern recognition with heavy-tailed distributions.
method Introduces Max-D-SW, a metric adjustment of Max-Sliced Wasserstein distance that aggregates over orthonormal bases.
result Max-D-SW provides a clear numerical advantage in MDS outcomes, especially for heavy-tailed distributions.

This paper unifies two types of statistical methods for estimating treatment effects.

problem Isolating online A/B-tests and off-policy evaluation.
method Establishes formal equivalence between online Difference-in-Means and off-policy Inverse Propensity Scoring methods.
result Standard online methods are mathematically equivalent to off-policy methods with optimal control variates.

The paper proposes a test to assess rater accuracy while accounting for rater covariates.

problem Assessing the accuracy of raters in medical imaging and forensic studies.
method Covariate-adjusted homogeneity test to determine differences in accuracy among multiple rater groups.
result The proposed test identifies statistically significant differences among five participant groups in a face recognition study.

The US Census Bureau corrupts data to protect privacy, but we show how to clean and analyze it effectively.

problem Analyzing Census data with intentional corruption to maintain privacy.
method Formulated a semiparametric model, proposed data cleaning, estimation, and inference procedures.
result Demonstrated that data cleaning can maintain precision and provided theoretical and empirical support.

Quantile regression is a tool for learning conditional distributions. In this paper we study quantile regression in the setting where a protected attribute is unavailable when fitting the model. This can lead to "unfair'' quantile estimators for which the effective quantiles are very different for the subpopulations de…

2019-07-19abs ↗pdf ↗

Paper develops a framework to discover bioprocessing regulatory mechanisms using symbolic and statistical learning.

problem Challenges in modeling complex intracellular regulation, stochastic system behavior, and limited experimental data.
method Symbolic and statistical learning framework based on stochastic differential equations and Bayesian learning.
result Improved sample efficiency and robust model selection compared to state-of-the-art approaches.

Bayesian neural networks use temperature adjustments to improve predictive performance.

problem Lack of theoretical generalization guarantees for Bayesian neural networks.
method Temperature adjustments to balance likelihood and prior regularization.
result Improved predictive performance through temperature adjustments.

The paper proposes a test to determine the number of latent classes in ordinal categorical data.

problem Determining the correct number of latent classes in latent class models with ordinal categorical data.
method The test statistic centers the largest singular value of a normalized residual matrix by a simple sample-size adjustment.
result The test statistic converges to zero under the null hypothesis and exceeds a fixed positive constant under an under-fitted alternative.

The distribution of health care payments to insurance plans has substantial consequences for social policy. Risk adjustment formulas predict spending in health insurance markets in order to provide fair benefits and health care coverage for all enrollees, regardless of their health status. Unfortunately, current risk a…

2019-01-28abs ↗pdf ↗

Paper proposes faster adaptation to distribution shifts in online settings.

problem Violation of exchangeability assumption in evolving data environments.
method Online conformal inference with retrospective adjustment.
result Faster adaptation to distributional shifts demonstrated through numerical studies.

We develop methods to approximate derivatives for causal inference problems using data.

problem Estimating causal effects from data when distributions are not known.
method Constructive algorithm approximating Gateaux derivatives via finite differencing.
result Derives conditions for finite-difference approximations to preserve statistical benefits.

A goodness-of-fit test for DCSBM improves scalability and power for large sparse networks.

problem Testing goodness-of-fit for degree-corrected stochastic block models (DCSBM) in large sparse networks.
method Proposes an adjusted chi-square test statistic for multinomial distributions, adjusted for degree-corrected networks, and applies it to compressed adjacency matrices.
result The test statistic converges in distribution under null, and is consistent in recovering the number of communities.

Benchmarking deep learning models for financial time series, focusing on risk-adjusted performance.

problem Optimizing risk-adjusted performance in financial time series prediction.
method Evaluation of various deep learning architectures including linear models, RNNs, transformers, state space models, and sequence representation approaches.
result Hybrid models like VSN with LSTM and xLSTM achieve the highest overall Sharpe ratio and superior downside adjusted characteristics.

Investments with best performance are not associated with best Sharpe ratios.

problem The relationship between performance and risk-adjusted return (Sharpe ratio) is counterintuitive for heavy-tailed distributions.
method Synthetic and real data analysis of returns distributions.
result The best-performing investments are not the best in terms of Sharpe ratio, and vice versa.

The study uses pre-trained neural networks to adjust for confounding in non-tabular data.

problem Neglecting non-tabular data sources can lead to biased ATE estimates.
method Leverages latent features from pre-trained neural networks to adjust for confounding.
result Neural networks can achieve fast convergence rates for ATE estimation with latent features.

EASE estimator improves probabilistic value estimation efficiency.

problem Efficiently estimating probabilistic values like Shapley and semivalues.
method Developed an Efficiency-Aware Surrogate-adjusted Estimator (EASE) that minimizes first-order mean squared error.
result EASE consistently outperforms existing estimators for various probabilistic values.

Develops an empirical likelihood framework for random forests and ensembles.

problem Quantifying the statistical uncertainty of random forests and ensembles.
method Empirical likelihood framework exploiting the incomplete UU-statistic structure of ensemble predictions.
result Modified empirical likelihood statistic achieves accurate coverage and practical reliability.

Machine learning improves trial analysis precision by adjusting for prognostic variables.

problem Improving precision in randomized trial analyses using covariate adjustment.
method Targeted machine learning estimation (TMLE) with adaptive pre-specification.
result Maximized empirical efficiency through cross-validated variance minimization.

Paper uses bond pricing and convexity adjustments to explain herd immunity paradox.

problem Early onset of herd immunity contradicts R value estimates from early stage growth.
method Utilizes Vasicek's bond pricing formula and de Finetti's Theorem approach.
result Reduces modeling discrepancy to simple convexity formulas.

It is common practice in using regression type models for inferring causal effects, that inferring the correct causal relationship requires extra covariates are included or ``adjusted for''. Without performing this adjustment erroneous causal effects can be inferred. Given this phenomenon it is common practice to inclu…

2019-06-17abs ↗pdf ↗

In this paper one studies the distribution of log-returns (tick-by-tick) in the Lisbon stock market and shows that it is well adjusted by the solution of the equation, {dpxdx=βqpxq(βqβq)pxq\frac{dp_{x}}{d| x|}=-β_{q^{\prime }}p_{x}^{q^{\prime}}-(β_{q}-β_{q^{\prime}}) p_{x}^{q}}, which corresponds to a generalization of the differential …

2004-03-24abs ↗pdf ↗

We constructed an analog electrical circuit which generates fluctuations in which probability density function has power law tails. In the circuit fluctuations with an arbitrary exponent of the power law can be obtained by adjusting the resistance. With this low cost circuit the random fluctuations which have the simil…

2001-04-18abs ↗pdf ↗

Develops model-free methods for event history analysis and efficient covariate adjustment.

problem Estimating treatment effects while accounting for confounding and understanding event history.
method Model-free prediction techniques, Local Covariance Measure (LCM), Debiased Outcome-adapted Propensity Estimator (DOPE), Aalen Covariance Measure (ACM).
result Demonstrates the effectiveness and robustness of the proposed methods in various settings.

Enhanced ROOT-SGD optimizes stochastic optimization with diminishing stepsizes.

problem Improving statistical efficiency in stochastic optimization.
method Integrates a diminishing stepsize strategy into ROOT-SGD.
result Achieves optimal convergence rates with improved stability and precision.

New model corrects bias in crowdsourced ratings for diverse items.

problem Bias and noise in crowdsourced ratings for training data.
method Bayesian rating model with item-level effects for difficulty, discriminativeness, and guessability.
result New model avoids bias in training data, improving model goodness of fit.