Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,738 papers · 148 categories

Trend · papers per month

3907801,1691,559 · Jun 202019922001200920172026
48 results for Total Survey Error Model

Machine learning improves official statistics but needs rigorous validation.

problem Lack of methodological robustness in machine learning for official statistics.
method Total Machine Learning Error (TMLE) framework to validate ML models.
result TMLE addresses representativeness and measurement errors in ML models.

In this note we survey recent results on the extrinsic geometry of the Jacobian locus inside Ag\mathsf{A}_g. We describe the second fundamental form of the Torelli map as a multiplication map, recall the relation between totally geodesic subvarieties and Hodge loci and survey various results related to totally geodesic…

2018-09-17abs ↗pdf ↗

Study shows non-systematic bias in customer satisfaction surveys limits data value.

problem Non-systematic bias in customer satisfaction surveys limits data value.
method Used real customer satisfaction survey data of a large retail bank to show the irreducible error and suggest thoughtful survey design methods.
result A thoughtful survey design can reduce non-systematic error in customer satisfaction surveys.

Germany's tax admin costs likely exceed 20% of total revenue, requiring system improvement.

problem High tax administrative costs in Germany and other jurisdictions.
method Statistical data, surveys, and a novel approach to measure total administrative cost as a percentage of total tax revenue.
result Germany's 2021 tax administrative costs likely exceeded 20% of total tax revenue.

The study examines methods to correct measurement error in nutritional epidemiology studies.

problem Measurement error in nutritional studies leads to biased and underconfident estimates.
method The article reviews various bias-correction models for exposure variables in nutritional epidemiology.
result Bias-correction methods are essential for accurate inference in nutritional studies.

Method completes mixed matrix from complex surveys with heterogeneous missingness.

problem Recovering a mixed dataframe matrix from complex survey sampling with different missingness patterns.
method Two-stage procedure: logistic regression for missingness modeling, and weighted log-likelihood maximization with low-rank constraint.
result The proposed method achieves sublinear convergence and shows superior performance compared to existing methods.

The paper tests the credibility of public and private surveys using linear regression and differential privacy.

problem Ensuring the validity of data analysis results from sample surveys using linear regression.
method Designing an algorithm to test the credibility of surveys and extending it to handle LDP.
result The algorithm achieves optimal estimation error bound for 1\ell_1 linear regression and reduces sample complexity.

SDRF estimates complex survey designs for conditional distributions.

problem Estimating conditional distributions under complex survey designs.
method Survey-calibrated distributional random forest (SDRF) with pseudo-population bootstrap and MMD split criterion.
result Established design consistency and model consistency for survey designs.

With a growing interest in using non-representative samples to train prediction models for numerous outcomes it is necessary to account for the sampling design that gives rise to the data in order to assess the generalized predictive utility of a proposed prediction rule. After learning a prediction rule based on a non…

2017-11-13abs ↗pdf ↗

Survey examines deep neural networks' ability to approximate functions.

problem Approximation of target functions by deep neural networks.
method Examination of feed-forward and residual architectures, focusing on optimization problems in regression and classification.
result Deep neural networks can approximate functions effectively, especially with ReLU activation functions.

Dynamic tracking error framework shows similar performance but varying volatility across different constraints.

problem Differences in governance parameters between Total Portfolio Approach and Strategic Asset Allocation.
method Portfolio simulations using U.S. equity and bond data from 2000 to 2026, spanning 2004 to 2026.
result Realized tracking error volatility varies 12-fold across different constraints, with costs highest during crises.

New method improves false-/true-positive-rate estimation in fraud detection with noisy labels.

problem Estimating FPR/TPR in fraud detection with class-conditional label noise.
method Directly cleaning model's validation data to de-correlate cleaning error with model scores.
result Improves accuracy of FPR/TPR estimates, especially in asymmetric label noise scenarios.

Error estimates found between SGD with momentum and Langevin diffusion.

problem Quantifying the difference between SGD with momentum and Langevin diffusion.
method Established error estimates using 1-Wasserstein and total variation distances.
result Quantitative error estimates between SGD with momentum and underdamped Langevin diffusion.

The data torrent unleashed by current and upcoming astronomical surveys demands scalable analysis methods. Many machine learning approaches scale well, but separating the instrument measurement from the physical effects of interest, dealing with variable errors, and deriving parameter uncertainties is often an after-th…

2017-07-14abs ↗pdf ↗

Deep learning models outperform MICE in large survey imputation but with hyperparameter tuning.

problem Comparing deep learning and MICE for missing data imputation in large surveys.
method Extensive simulation studies comparing four machine learning-based MI methods: MICE with classification trees, MICE with random forests, generative adversarial imputation networks, and multiple imputation using denoising autoencoders.
result MICE with classification trees consistently outperforms deep learning methods in terms of bias, mean squared error, and coverage.

New online learning algorithm combines PA and TER for binary classification.

problem Binary classification with non-separable data and data imbalance.
method Online Passive-Aggressive (PA) and Total-Error-Rate (TER) learning combined into PATER algorithm.
result PATER algorithms outperform existing online learning algorithms in efficiency and effectiveness.

We introduce anti-invariant Riemannian submersions from Sasakian manifolds onto Riemannian manifolds. We survey main results of anti-invariant Riemannian submersions defined on Sasakian manifolds. We investigate necessary and sufficient condition for an anti-invariant Riemannian submersion to be totally geodesic and ha…

2013-02-20abs ↗pdf ↗

We survey some LpL^{p}-vanishing results for solutions of Bochner or Simons type equations with refined Kato inequalities, under spectral assumptions on the relevant Schrödinger operators. New aspects are included in the picture. In particular, an abstract version of a structure theorem for stable minimal hypersurfaces…

2010-11-24abs ↗pdf ↗

Approximate message passing algorithm enjoyed considerable attention in the last decade. In this paper we introduce a variant of the AMP algorithm that takes into account glassy nature of the system under consideration. We coin this algorithm as the approximate survey propagation (ASP) and derive it for a class of low-…

2018-07-03abs ↗pdf ↗

Bayesian framework improves robustness in nonlinear regression models.

problem Measurement error, model misspecification, and distributional misspecification in regression analyses.
method Joint Dirichlet process prior on latent covariate-response distribution, updating with posterior pseudo-samples.
result Improved stability and consistency in estimators under increasing measurement error.

Langevin dynamics fails to produce accurate samples even with small score function errors.

problem Robustness of Langevin dynamics to score function errors.
method Analysis of Langevin dynamics and score function errors.
result Langevin dynamics produces a distribution far from the target distribution in TV distance even with small L2L^2 errors in the score function.

We survey what is known about minimal surfaces in R3\bold R^3 that are complete, embedded, and have finite total curvature. The only classically known examples of such surfaces were the plane and the catenoid. The discovery by Costa, early in the last decade, of a new example that proved to be embedded sparked a great…

1995-08-09abs ↗pdf ↗

StatEcoNet models species distribution using neural networks to correct observation errors.

problem Correcting observation errors in wildlife surveys for accurate species distribution modeling.
method StatEcoNet integrates a graphical generative model with neural networks to address SDM challenges.
result StatEcoNet outperforms traditional methods on simulated and real datasets.

Example shows learnable distributions not privately learnable.

problem Learnable distributions under non-private conditions not transferable to differential privacy.
method Example of a distribution class learnable up to constant error in total variation distance but not under differential privacy.
result Contradicts conjecture of Ashtiani on learnability under differential privacy.

PPI uses survey sampling methods for inference, bridging ML and statistics.

problem Combining machine learning predictions with small labeled data for valid inference.
method Equivalence of PPI estimators to survey sampling methods.
result PPI estimators are algebraically equivalent to survey sampling methods.

Estimates TV distance between autoregressive models under different access models.

problem Estimating the total variation distance between two autoregressive distributions.
method Three access models: sample access, logit access, and noisy logit access; provides query complexity for each.
result Improved query complexity for estimating TV distance in autoregressive models.

DFM models are analyzed for generating distributions with provable convergence.

problem Training DFM models to generate distributions that match true data.
method Theoretical analysis decomposes error into approximation and estimation errors.
result DFM models converge to true data distribution as training set size increases.

AI-assisted interviews allow respondents to describe experiences naturally, but mapping those accounts into structured survey variables is fallible.

problem Mapping AI-assisted interview responses into structured survey variables is fallible.
method Adaptive Matrix Validation (AMV) is proposed, which involves mapping responses into tabular data and using a small set of structured questions for statistical adjustment.
result The estimator calibrates mapped values using validation answers from other respondents and corrects remaining error with validation answers observed for the target respondent.

PICZL improves photometric redshifts for AGN in all-sky surveys.

problem Challenges in accurately computing photo-z for AGN due to interplay of SMBH and host galaxy emissions.
method PICZL uses an ensemble of CNNs with cross-channel integration of image and catalog data, leveraging Gaussian mixture models.
result PICZL achieves a photo-z variance of 4.5% and outlier fraction of 5.6% on a validation sample of 8098 AGN, outperforming previous methods.

In the first half of this article, we survey the new quasi-local and total angular momentum and center of mass defined in [9] and summarize the important properties of these definitions. To compute these conserved quantities involves solving a nonlinear PDE system (the optimal isometric embedding equation), which is ra…

2014-09-17abs ↗pdf ↗

The purpose of this paper is to both survey and offer some new results on the non-triviality of the characteristic classes of Riemannian foliations. We give examples where the primary Pontrjagin classes are all linearly independent. The independence of the secondary classes is also discussed, along with their total var…

2008-06-22abs ↗pdf ↗

We consider two active binary-classification problems with atypical objectives. In the first, active search, our goal is to actively uncover as many members of a given class as possible. In the second, active surveying, our goal is to actively query points to ultimately predict the proportion of a given class. Numerous…

2012-06-27abs ↗pdf ↗

We introduce anti-invariant Riemannian submersions from cosymplectic manifolds onto Riemannian manifolds. We survey main results of anti-invariant Riemannian submersions defined on cosymplectic manifolds. We investigate necessary and sufficient condition for an anti-invariant Riemannian submersion to be totally geodesi…

2013-02-20abs ↗pdf ↗