Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

86172258344 · Jun 202019922001200920182026
48 results for variance testing

Proposes a modified Morgan-Pitman test for evaluating variances in machine learning models.

problem Limited ability to account for sampling variability in model selection.
method Enhances the classic Morgan-Pitman test for robustness in non-linear models with heavy-tailed distributions or outliers.
result Demonstrates the test's effectiveness and practical utility in model evaluation and selection.

Adversarial training leads to large generalization gap, decomposed into bias and variance.

problem Understanding the large generalization gap in adversarially trained models.
method Bias-Variance decomposition of test risk as a function of adversarial perturbation radius.
result Bias increases monotonically with adversarial perturbation radius and is dominant in test risk.

Develops abstention procedure for nonparametric regression via variance testing.

problem Prediction with selective abstention in error-critical machine learning.
method Nonparametric heteroskedastic regression via testing hypothesis on conditional variance.
result Non-asymptotic risk bounds and convergence regimes for the estimator.

Modern neural networks show no bias-variance tradeoff with increased parameters.

problem The traditional bias-variance tradeoff does not hold in over-parameterized neural networks.
method Empirical measurements and theoretical analysis of bias and variance in modern neural networks.
result Bias and variance can decrease as the number of parameters grows in over-parameterized neural networks.

This work uses ANOVA to understand how different factors contribute to test error in machine learning models.

problem Understanding why overparametrized models generalize well despite potentially fitting noise.
method Analysis of variance (ANOVA) to decompose test error into components of variance.
result The interaction between training samples and initialization can dominate variance, and there are phase transitions in variance behavior.

Study tests GMVP weights in high-dimensional settings, comparing sample and shrinkage estimators.

problem Testing GMVP weights in high-dimensional settings with varying sample size and asset count.
method Developed two tests based on sample and shrinkage estimators of GMVP weights.
result Shrinkage estimator test performs well even for high asset counts.

Adaptive sampling method reduces variance in stochastic optimization.

problem Reducing variance in stochastic optimization with limited gradient computations.
method Adaptive increase in sample size based on inner product test.
result Algorithm converges globally on nonconvex functions and linearly on strongly convex functions.

Introduces TPV to analyze model robustness without labels.

problem Analyzing post-training robustness of machine learning models.
method Parameter perturbations and test prediction variance (TPV) as a unifying framework.
result TPV connects various perturbations under a single lens, providing insights into model stability.

Paper explains why Dropout and BN lead to worse performance when combined and proposes solutions.

problem Worse performance when Dropout and BN are combined.
method Theoretical analysis and experiments on various networks to identify variance shift and propose solutions.
result Dropout shifts variance of a specific neural unit, while BN maintains accumulated variance, leading to unstable predictions.

Empirical study finds variance swap rate is affine in spot variance for S&P500 data.

problem Investigating the relationship between variance swap rate and spot variance.
method Empirical analysis using S&P500 data from 2006-2018, testing different models.
result Affine relationship between variance swap rate and spot variance is supported.

Develops new e-processes and confidence sequences for Gaussian means with unknown variance.

problem Constructing valid t-tests and confidence sequences for Gaussian means with unknown variance.
method Explores generalized nonintegrable martingales and extended Ville's inequality, developing two new e-processes and confidence sequences.
result Analyzes the width of resulting confidence sequences with a polynomial dependence on error probability, proving it to be unavoidable and even better than classical fixed-sample t-tests.

Paper proposes a new importance sampling method for reducing variance.

problem Reducing variance in importance sampling when training and testing data come from different distributions.
method A new variant of importance sampling that reduces variance by orders of magnitude.
result The new estimator can improve estimates of treatment effectiveness using limited data.

Improved ANOVA test under differential privacy with higher statistical power.

problem Carrying out ANOVA tests while maintaining privacy.
method Developed a new test statistic \(F_1\) and a method to compute its reference distribution.
result Our test \(F_1\) achieves a significant improvement in statistical power compared to previous methods.

Faster convergence of kernel mean embeddings using variance information.

problem Speeding up the convergence rate of kernel mean embeddings.
method Leveraging variance information in reproducing kernel Hilbert space and estimating variance from data.
result Efficiently estimate variance information from data to achieve distribution-agnostic convergence bounds.

Deep networks generalize well even when they fit training data perfectly, thanks to overparametrization.

problem Understanding generalization in overparametrized deep networks.
method Random features regression, asymptotic analysis, ensemble averaging.
result Bias remains constant beyond the interpolation threshold, while variance components decay with overparametrization.

This paper proposes a new AED framework for multi-metric experiments with fixed budget.

problem Statistical power challenges in testing multiple metrics simultaneously.
method Two-phase structure: adaptive exploration followed by validation. SHRVar algorithm with relative-variance-based sampling.
result Achieves provable error probability that decreases exponentially.

A/B testing improves marketing decisions by selecting effective stratification variables.

problem Improving the sensitivity of A/B testing through stratified sampling.
method Designing an algorithm to select a subset of stratification variables for variance reduction.
result The subset selection method outperforms other variance reduction techniques in A/B testing.

New insights into bias and variance in over-parameterized models.

problem Understanding bias and variance in over-parameterized models.
method Analytic expressions derived from statistical physics for two minimal models.
result Over-parameterized models can overfit even in noiseless conditions.

Study proposes memory-efficient backpropagation for linear layers in neural networks.

problem Significant memory usage in backpropagation through linear layers in neural networks.
method Randomized matrix multiplications to reduce memory usage with a moderate decrease in test accuracy.
result Demonstrated benefits of the proposed method on fine-tuning pre-trained models.

Scaling laws in linear regression explain model performance improvements with size and data.

problem Disagreement between empirical neural scaling laws and conventional wisdom on variance error.
method Infinite dimensional linear regression setup, one-pass SGD, Gaussian prior, power-law spectrum.
result Variance error is dominated by other errors, disappearing from the bound due to SGD's implicit regularization.

A simple method treats heteroscedastic variance variatively, improving model calibration and sample quality.

problem Brittle optimization impacts model likelihoods for mean and variance estimation.
method Proposes a variational approach to heteroscedastic variance, improving predictive mean and variance calibration.
result The proposed method significantly improves parameter calibration and sample quality for regression and VAEs.

USNRT uses tree-structured learning to improve uncertainty quantification of variance networks.

problem Improving uncertainty quantification of variance networks.
method Tree-structured local neural network model that partitions feature space into regions for training region-specific neural networks to predict mean and variance.
result USNRT shows superior performance in estimating uncertainty with variances on UCI datasets compared to recent methods.

This paper analyzes the posterior variance of Gaussian processes and derives a new bound.

problem Lack of suitable analysis of posterior variance for finite and infinite training data.
method Derives a novel bound for posterior variance requiring only local information.
result Proves sufficient conditions for the convergence of posterior variance to zero and demonstrates improved average learning bound.

Proposes incorporating noise sources in machine learning evaluation for more reliable conclusions.

problem Inadequate handling of nondeterminism in machine learning research leads to unreliable results.
method Uses linear mixed effects models (LMEMs) and generalized likelihood ratio tests (GLRT) to analyze performance evaluation scores and assess performance differences.
result Demonstrates how to incorporate various sources of noise and data properties into statistical significance testing and reliability analysis.

New algorithm detects changes in high-dimensional data with mean and variance.

problem Challenges in detecting changes in high-dimensional data with mean and variance.
method Complete graph-based approach to detect changes of mean and variance from low to high-dimensional online data.
result The proposed method outperforms existing methods in terms of detection power.

Detects underspecification in pre-trained models using local ensembles.

problem Underspecification in pre-trained models where many predictors are consistent with training data.
method Uses local second-order information to approximate prediction variance across an ensemble of models.
result Capable of detecting underspecification in pre-trained models on test data.

The paper examines skill estimation and variance under model misspecification in IRT.

problem Underestimation and overestimation of skills when non-compensatory model is misspecified as compensatory.
method Theoretical approach to analyze underestimation and overestimation of skills and variance.
result Overestimation of skills occurs around the origin and asymptotic variance differs under model misspecification.

Paper tests for time-varying entropy in stock prices, finding periods of inefficiency.

problem Testing for time-varying entropy in stock price dynamics.
method Unbiased approximation of Shannon entropy variance, optimal rolling window selection, hypothesis testing.
result Existence of periods of market inefficiency for meme stocks.

Method selects number of communities in weighted networks.

problem Selecting the number of communities in weighted networks.
method Proposes a novel weighted DCSBM and uses a sequential testing framework with spectral clustering and matrix scaling.
result Method is consistent in estimating the true number of communities under mild conditions.

Normalization effects on deep neural networks impact output variance and test accuracy.

problem The impact of normalization on deep neural networks' statistical behavior and test accuracy.
method Asymptotic expansion analysis of neural network's output for different γiγ_i values.
result Equal γiγ_i values (one) provide the best statistical behavior and test accuracy.

Unified method for MMD variance estimation improves accuracy and computational efficiency.

problem Variance estimation for MMD in nonparametric testing.
method Unified finite-sample characterization of MMD variance through U-statistic and Hoeffding decomposition; exact acceleration method for univariate case.
result Unified estimators improve accuracy and computational efficiency for MMD variance.

New Riemannian optimization improves variance estimation in mixed models.

problem Challenges in estimating variance parameters in linear mixed models due to constraints.
method Formulated as an optimization problem on a Riemannian manifold, using Riemannian gradient and Hessian.
result Yields higher quality variance parameter estimates compared to existing methods.

Method clusters molecular systems based on dynamics or structure similarity.

problem Clustering molecular systems based on dynamics or structure similarity.
method Ward's minimum variance clustering using Jensen-Shannon divergence.
result Method avoids overfitting in supervised learning.