Proposes a new cross-validation method to estimate model performance.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Learn2Evaluate uses learning curves to estimate high-dimensional prediction performance.
Estimates neural architecture performance speedily.
This paper introduces a new property of estimators of the strength of statistical association, which helps characterize how well an estimator will perform in scenarios where dependencies between continuous and discrete random variables need to be rank ordered. The new property, termed the estimator response curve, is e…
Study compares different covariance estimation methods for portfolio allocation.
Empirical median performs well in estimating location with varying scales.
Self-distillation optimally improves model performance in spiked covariance models.
New estimators outperform maximum likelihood without hyper-parameter estimation.
PromptEval estimates LLM performance across many prompts, improving reproducibility.
Estimates policy performance in small-data settings without sacrificing data.
Estimating properties of discrete distributions is a fundamental problem in statistical learning. We design the first unified, linear-time, competitive, property estimator that for a wide class of properties and for all underlying distributions uses just samples to achieve the performance attained by the empirical…
We study the distributions of the LASSO, SCAD, and thresholding estimators, in finite samples and in the large-sample limit. The asymptotic distributions are derived for both the case where the estimators are tuned to perform consistent model selection and for the case where the estimators are tuned to perform conserva…
Study evaluates model selection methods for time series forecasting.
We propose a supervised anomaly detection method based on neural density estimators, where the negative log likelihood is used for the anomaly score. Density estimators have been widely used for unsupervised anomaly detection. By the recent advance of deep learning, the density estimation performance has been greatly i…
The positivity assumption, or the experimental treatment assignment (ETA) assumption, is important for identifiability in causal inference. Even if the positivity assumption holds, practical violations of this assumption may jeopardize the finite sample performance of the causal estimator. One of the consequences of pr…
New method reduces variance in subpopulation model performance estimates.
Random variables of the generalized Pareto distribution, can be transformed to that of the Pareto distribution. Explicit expressions exist for the maximum likelihood estimators of the parameters of the Pareto distribution. The performance of the estimation of the shape parameter of generalized Pareto distributed using …
Improved LDA using a nonlinear covariance estimator for better performance.
Conditional forecasts improve performative prediction accuracy.
Improved MoM estimator enhances classical shadows protocol for quantum measurements.
The Kalman filter and Heston model are used to estimate asset prices and trading performance.
Estimates model performance under distribution shift using domain-invariant predictors.
A fast bootstrap method estimates cross-validation standard error.
New GLS estimator handles high-dimensional data with autocorrelated errors.
Meta-learners improve causal effect estimation in small samples.
The Rasch model is widely used for item response analysis in applications ranging from recommender systems to psychology, education, and finance. While a number of estimators have been proposed for the Rasch model over the last decades, the available analytical performance guarantees are mostly asymptotic. This paper p…
Recently, a framework for application-oriented optimal experiment design has been introduced. In this context, the distance of the estimated system from the true one is measured in terms of a particular end-performance metric. This treatment leads to superior unknown system estimates to classical experiment designs bas…
EBQL reduces bias in Q-learning for improved performance.
This paper studies the partial estimation of Gaussian graphical models from high-dimensional empirical observations. We derive a convex formulation for this problem using -regularized maximum-likelihood estimation, which can be solved via a block coordinate descent algorithm. Statistical estimation performance …
Extends covariance estimation with multiple targets for better performance.
Non-convex regularizers usually improve the performance of sparse estimation in practice. To prove this fact, we study the conditions of sparse estimations for the sharp concave regularizers which are a general family of non-convex regularizers including many existing regularizers. For the global solutions of the regul…
ProEval efficiently estimates AI performance and discovers failures using pre-trained Gaussian Processes.
Paper extends Chernoff sampling for active testing and parameter estimation, improving neural network and regression models.
Proposes EM for sparse horseshoe estimation.
New algorithms estimate Hessians using random directions for faster stochastic optimization.
Paper presents deep learning and ML for automated student performance estimation.
New weighted Lasso estimates improve logistic regression performance with measurement error.
This paper explores how entropic regularization improves Wasserstein estimators' performance.
We consider a distributed parameter estimation problem, in which multiple terminals send messages related to their local observations using limited rates to a fusion center who will obtain an estimate of a parameter related to observations of all terminals. It is well known that if the transmission rates are in the Sle…
Study proposes efficient estimators for matrix-valued linear regression under sparsity assumptions.
The estimation of class prevalence, i.e., the fraction of a population that belongs to a certain class, is a very useful tool in data analytics and learning, and finds applications in many domains such as sentiment analysis, epidemiology, etc. For example, in sentiment analysis, the objective is often not to estimate w…
New method uses MMD estimators to enforce model invariance with missing data.
OPERA blends multiple OPE estimators to evaluate new policies offline.
Unified DICE estimators as regularized Lagrangians for improved off-policy evaluation.
Well begun is half done. In the crowdfunding market, the early fundraising performance of the project is a concerned issue for both creators and platforms. However, estimating the early fundraising performance before the project published is very challenging and still under-explored. To that end, in this paper, we pres…
We study the design of portfolios under a minimum risk criterion. The performance of the optimized portfolio relies on the accuracy of the estimated covariance matrix of the portfolio asset returns. For large portfolios, the number of available market returns is often of similar order to the number of assets, so that t…
The ability to perform offline A/B-testing and off-policy learning using logged contextual bandit feedback is highly desirable in a broad range of applications, including recommender systems, search engines, ad placement, and personalized health care. Both offline A/B-testing and off-policy learning require a counterfa…
Mixtures-of-Experts (MoE) are conditional mixture models that have shown their performance in modeling heterogeneity in data in many statistical learning approaches for prediction, including regression and classification, as well as for clustering. Their estimation in high-dimensional problems is still however challeng…