Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

99199298397 · Jun 202019922001200920182026
48 results for statistical quality assessment

Framework for automatically assessing and correcting data quality issues without domain knowledge.

problem Ensuring data quality in datasets across various domains.
method Hybrid approach combining statistical and machine learning methods.
result Effective detection and correction of missing values, duplicates, and typographical errors.

Paper proposes a method to estimate intra-observer variability in echocardiography quality assessment.

problem Intra-observer variability in echocardiography quality assessment impacts deep neural network reliability.
method Modeling intra-observer variability as aleatoric uncertainty in a regression problem.
result The proposed method reduces error from 0.11 to 0.09, improving test accuracy by 5.7%.

PQMass assesses generative model quality using chi-squared tests.

problem Assessing the quality of generative models without density assumptions.
method Divides sample space into regions, applies chi-squared tests to p-values.
result Effectively assesses generative model quality, novelty, and diversity.

AI systems need reliable testing to ensure safety and trustworthiness.

problem Current AI Act lacks functional trustworthiness for AI systems.
method Define technical application distribution, set risk-based performance, and conduct statistically valid testing.
result Reliable functional trustworthiness is essential for AI systems.

The bootstrap provides a simple and powerful means of assessing the quality of estimators. However, in settings involving large datasets, the computation of bootstrap-based quantities can be prohibitively demanding. As an alternative, we present the Bag of Little Bootstraps (BLB), a new procedure which incorporates fea…

2012-06-27abs ↗pdf ↗

AgraSSt assesses graph generators using Stein operators and kernel discrepancies.

problem Assessing the quality of graph generators that are implicit or not in explicit form.
method AgraSSt uses Stein operators and kernel discrepancies to assess graph generators, providing interpretable criticisms.
result Theoretical guarantees and empirical validation for various graph models.

Paper controls false discovery rate in crowdsourced annotator quality.

problem Crowdsourced annotators may have position bias affecting label quality.
method Statistical framework with knockoff filters and Inverse Scale Space dynamics.
result Controls false discovery rate without prior knowledge of biased annotators.

The paper evaluates and improves uncertainty estimates in neural networks for safety-critical applications.

problem Quantifying uncertainty in neural networks for safety-critical systems.
method Proposes a statistical test for evaluating uncertainty realism in neural networks and transfers a classification architecture to image-to-image tasks.
result The variational U-Net architecture significantly improves uncertainty realism in image-to-image tasks compared to a plain model.

MRI image quality affects statistical and predictive analysis of brain morphology.

problem Impact of MRI image quality on statistical and predictive analysis of brain morphology.
method Systematic testing of image quality on univariate statistics and machine learning classification using three large datasets.
result Low-quality MRI data significantly affects detecting significant sex/gender differences in smaller samples, but not in larger ones.

The paper assesses quality measures for machine learning models using cross-validation.

problem Evaluating the accuracy and robustness of quality measures for machine learning models.
method Cross-validation approach to estimate prediction error and quantify explained variation. Confidence bounds and local quality measures derived from residuals.
result The reliability and robustness of quality measures are assessed through numerical examples and confidence bounds.

Statistical inference on graphs is a burgeoning field in the applied and theoretical statistics communities, as well as throughout the wider world of science, engineering, business, etc. In many applications, we are faced with the reality of errorfully observed graphs. That is, the existence of an edge between two vert…

2012-11-15abs ↗pdf ↗

Study improves data quality assessment for structural monitoring data.

problem Ensuring reliability of structural health monitoring data.
method Probabilistic data quality assessment using a conditional diffusion model.
result Significantly improves accuracy of data quality assessment.

This paper compares automatic metrics for re-speaking quality assessment.

problem Estimating the quality of re-speaking results is challenging.
method Comparing and adapting automatic evaluation metrics (BLEU, EBLEU, NIST, METEOR, etc.) to re-speaking quality.
result Automatic metrics are compared to human-derived NER metric for re-speaking quality.

This paper evaluates features for assessing digital ophthalmoscopy image quality.

problem Accurate teleophthalmology requires high-quality ophthalmoscopic imagery.
method Statistical metrics, gradient-based metrics, and wavelet transform coefficient derived indicators were tested using machine learning.
result Suitability of features for image quality assessment confirmed, though on a small data set.

Automatically computes reference ranges for UK Biobank cardiac data.

problem Improving healthcare by discovering patterns in large-scale population data.
method Fully automatic pipeline for 3D cardiac MR image analysis.
result Statistically significant agreement between manual and automatic indexes.

A framework assesses the quality of crowdsourced weather data.

problem Quality control and assessment of crowdsourced weather data from third-party stations.
method Proposes a simple, scalable, and interpretable AI/Stats/ML framework to assess TPAWS data.
result Demonstrates the performance of the framework using synthetic and real data.

AutoMOS uses neural nets to assess speech quality without human raters.

problem Assessing the quality of synthesized speech using human raters is time-consuming and costly.
method Developed a deep recurrent neural network that uses raw waveforms as input to predict mean opinion scores (MOS).
result AutoMOS models provide utterance-level MOS estimates only slightly inferior to human raters and can be averaged for multiple utterances.

The study introduces a holdout-based framework to assess synthetic data fidelity and privacy.

problem Evaluating the quality and privacy of synthetic data solutions for mixed-type tabular data.
method Holdout-based empirical assessment framework measuring fidelity and privacy risk.
result Synthetic data samples are as close to the training as to the holdout data, indicating generalization and independence from individual records.

Researchers develop tests to assess quality of GAN-generated images.

problem Lack of objective means to evaluate domain-relevant quality of GAN-generated images.
method Designed stochastic context models (SCMs) and statistical classifiers to detect high-order spatial arrangements in GAN-generated images.
result GANs can generate images that appear accurate visually but lack specific high-order spatial arrangements.

Generative models assess quality on time-series data using ITS and FITD.

problem Lack of consensus for quality assessment of class-conditional generative models on time-series data.
method Introduced InceptionTime Score (ITS) and Frechet InceptionTime Distance (FITD) to evaluate generative models.
result ITS and FITD combined with TSTR can accurately assess generative model performance on time-series data.

New methods reduce bias in estimating optimality gaps for risk-averse stochastic programs.

problem Optimality gap estimation bias in risk-averse stochastic programs.
method Two independent samples, each estimating a different component of the optimality gap.
result Our method reduces bias in estimating optimality gaps for risk-averse problems.

The paper defines and assesses the quality of datasets using a novel expected diameter metric.

problem Lack of rigorous methods to assess data quality.
method Formal definition of data quality, expected diameter metric, Fourier analysis, algebraic methods, probabilistic analysis.
result The expected diameter metric provides theoretical guarantees and practical solutions for data quality assessment.

Paper introduces metrics to evaluate missing data imputation without ground truth.

problem Handling missing data in time series without ground truth.
method Introduces Wasserstein distance (WD) and Jensen-Shannon divergence (JSD) as metrics to evaluate imputation quality.
result WD and JSD are effective metrics for assessing missing data imputation quality.

The paper introduces a method to assess machine translation quality with confidence intervals.

problem Evaluating the uncertainty and quality of machine translation.
method Utilizes conformal predictive distributions to produce prediction intervals with guaranteed coverage.
result The method outperforms a baseline on six language pairs in terms of coverage and sharpness.

The paper improves itemset quality assessment by incorporating background knowledge.

problem Assessing the quality of discovered itemsets is challenging due to many patterns being explainable by background knowledge.
method The authors introduce a maximum entropy approach to efficiently infuse additional background knowledge such as row margins, lazarus counts, and bounds of ones.
result More sophisticated models that incorporate background knowledge fit the data better and improve frequency prediction of itemsets.

Dual quality assessment method tackles adversarial robustness issues across various metrics.

problem Varying robustness levels and bias in adversarial attacks and defenses.
method Model agnostic dual quality assessment method, including robustness levels.
result Current networks and defenses are vulnerable at all robustness levels, highlighting the need for a dual approach.

A popular tool for unsupervised modelling and mining multi-aspect data is tensor decomposition. In an exploratory setting, where and no labels or ground truth are available how can we automatically decide how many components to extract? How can we assess the quality of our results, so that a domain expert can factor th…

2015-03-11abs ↗pdf ↗

This paper uses Monte Carlo simulation to value quality options in agricultural futures contracts.

problem Valuation of quality options in agricultural futures to prevent manipulation and improve hedging performance.
method Monte Carlo simulation with antithetic variables for efficiency.
result Demonstrates a method to estimate the value of quality options in agricultural futures contracts.

Adapts deep learning models trained on simulated images for use with real images.

problem Difficulty in training deep neural networks on large amounts of experimental data.
method Adversarial domain adaptation method to mitigate domain shift between simulated and experimental image data.
result Adversarial domain adaptation successfully mitigates domain shift and improves numerical observer performance.

The study analyzes how probabilistic forecasts improve battery trading strategies in electricity markets.

problem Improvements in statistical forecast quality do not directly translate to economic value in battery trading strategies.
method The study frames battery optimization as a stochastic program based on fully probabilistic forecasts and examines decision quality under different uncertainty models.
result The study identifies two critical flaws in quantile-based trading strategies and provides theoretical justification and empirical evidence.

New framework assesses deep learning models for spatio-temporal data with missing data.

problem Challenges in assessing deep learning models for spatio-temporal data with missing and heterogeneous data.
method Residual correlation analysis framework using spatio-temporal graphs and asymptotically distribution-free summary statistics.
result Identification and localization of regions where predictive performance can be improved.

A deep learning framework assesses physical rehabilitation exercises.

problem Lack of versatile, robust, and practical assessment methods for rehabilitation exercises.
method Deep learning framework with metrics, scoring functions, and neural networks.
result First implementation of deep neural networks for rehabilitation performance assessment.

This paper evaluates forecast quality in electricity markets beyond traditional accuracy measures.

problem Traditional accuracy measures fail to reflect the economic value of electricity price forecasts.
method Investigates four quality dimensions: accuracy, dispersion, association, and extremum identification.
result Dispersion- and association-based measures better capture forecast economic value.

The bootstrap provides a simple and powerful means of assessing the quality of estimators. However, in settings involving large datasets---which are increasingly prevalent---the computation of bootstrap-based quantities can be prohibitively demanding computationally. While variants such as subsampling and the mm out o…

2011-12-21abs ↗pdf ↗

Study subjective perception of low light restored images and develop an unsupervised QA model.

problem Lack of subjective QA for low light restored images and challenges in collecting human opinion scores.
method Create a dataset, conduct subjective QA study, develop self-supervised contrastive learning technique to extract features.
result Unsupervised NR QA model achieves state-of-the-art performance for low light restored images.

A test assesses how well data fits a target density function.

problem Measuring how well data fits a target density function without assuming a specific form.
method Stein's method using Reproducing Kernel Hilbert Space functions to construct a divergence measure, estimated via V-statistic.
result The proposed test accurately assesses goodness of fit for various data types and contexts.