Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,982 papers · 148 categories

Trend · papers per month

166332498664 · Jun 202019922001200920172026
48 results for evaluation accuracy

The paper evaluates forecast accuracy of realized volatility measures in large cross-sections.

problem Forecast evaluation of realized volatility measures in large cross-sections of financial data.
method Equal predictive accuracy testing procedures, LASSO shrinkage, measurement error correction, cross-sectional jump component measures.
result The augmented HAR model outperforms the standard HAR model in forecasting realized volatility.

New method evaluates text-to-image synthesis for realism, variety, and semantic accuracy.

problem Lack of metrics revealing semantic accuracy in text-to-image synthesis.
method Uses Inception network representations and t-SNE visualization for semantic evaluation.
result Classification accuracy of generated images to real images' visual concepts correlates with semantic accuracy.

LaRT models LLMs' response accuracy and CoT length to evaluate reasoning ability and speed.

problem Valid evaluation of Large Language Models (LLMs) via response accuracy and chain-of-thought length.
method Introduces Latency-Response Theory (LaRT) to jointly model response accuracy and CoT length using latent ability and latent speed.
result LaRT yields higher estimation accuracy and shorter confidence intervals for latent traits compared to IRT.

This paper evaluates forecast quality in electricity markets beyond traditional accuracy measures.

problem Traditional accuracy measures fail to reflect the economic value of electricity price forecasts.
method Investigates four quality dimensions: accuracy, dispersion, association, and extremum identification.
result Dispersion- and association-based measures better capture forecast economic value.

Current OOD benchmarks overestimate model robustness to spurious correlations.

problem Spurious correlations degrade OOD performance, but benchmarks show the opposite.
method Analyze OOD datasets for spurious correlations and derive conditions for robustness.
result Current OOD benchmarks are misspecified and overestimate model robustness.

This paper evaluates and improves metrics for identifying important features in machine learning models.

problem Evaluation metrics for explainable AI are limited by multicollinearity and model accuracy.
method Proposes Expected Accuracy Interval (EAI) to predict model accuracy with multicollinearity.
result EAI is a useful metric for identifying important features in models with multicollinearity.

New estimator uses clustering to improve off-policy evaluation accuracy.

problem Improving off-policy evaluation accuracy when logging and evaluation policies differ.
method Proposes an estimator that shares information across similar contexts using clustering.
result Clustering contexts improves estimation accuracy, especially in deficient information settings.

New method evaluates feature interactions using orthogonal variance decomposition.

problem Feature selection fails to account for interactions between features.
method Orthogonal variance decomposition to evaluate feature subsets considering interactions.
result Our method accurately identifies relevant features and improves model accuracy.

We introduce a new metric to evaluate corruption robustness of ML classifiers.

problem Evaluating corruption robustness of machine learning classifiers.
method We propose a test data augmentation method using minimal class separation distance to derive a robustness distance ε and a metric MSCR.
result The MSCR metric allows interpretable comparison of classifier robustness on different datasets.

Mixed-size training improves CNN accuracy and speed.

problem Training CNNs on fixed image sizes limits their adaptability to various image sizes.
method Mixed-size training: training on multiple image sizes at once.
result Models trained with mixed-size images achieve higher accuracy and faster inference.

New method uses model comparison signals to improve LLM evaluation accuracy.

problem Limited benchmark sizes and model stochasticity in evaluating LLMs' mathematical reasoning.
method Combines standard labeled outcomes with model comparison signals to design a statistically efficient evaluation framework.
result Semiparametric estimator achieves the semiparametric efficiency bound and substantially improves ranking accuracy.

Evaluates explanations of LTR models using decision paths and compares their accuracy.

problem Challenges in evaluating local explanations of LTR models due to lack of ground truth feature importance scores.
method Focuses on tree-based LTR models, extracts ground truth feature importance scores using decision paths, and compares them with explanation techniques.
result Explanation accuracy varies depending on the model and data point.

Pairwise ranking aligns subjective clinical evaluations with objective indicators.

problem Aligning subjective clinical evaluations with objective indicators for improved diagnosis.
method Pairwise ranking methods to align subjective evaluations with objective indicators.
result The resulting score improves classification accuracy and provides a nuanced severity assessment.

This paper reviews PU learning evaluation methods and provides practical recommendations.

problem Evaluating PU learning methods when only positive and unlabelled data are available.
method Critical review of 51 articles proposing PU classifiers and alternative predictive accuracy measures.
result Practical recommendations for improving PU learning evaluation.

Probabilistic modeling is cyclical: we specify a model, infer its posterior, and evaluate its performance. Evaluation drives the cycle, as we revise our model based on how it performs. This requires a metric. Traditionally, predictive accuracy prevails. Yet, predictive accuracy does not tell the whole story. We propose…

2016-05-24abs ↗pdf ↗

Study evaluates deep learning methods for dermatology, finding they perform poorly under non-ideal conditions.

problem Lack of robustness of deep learning methods in dermatology under real-world conditions.
method Simulated non-ideal conditions on user-submitted dermatology images.
result Deep learning methods show significant drop in accuracy and prediction changes under non-ideal conditions.

Suitability filter detects model performance degradation in real-world deployment.

problem Ensuring model reliability in safety-critical domains without access to ground truth labels.
method Uses suitability signals to evaluate classifier performance on unlabeled user data.
result The suitability filter reliably detects performance deviations due to covariate shift.

This paper evaluates algorithms for classification and outlier detection accuracies in temporal data. We focus on algorithms that train and classify rapidly and can be used for systems that need to incorporate new data regularly. Hence, we compare the accuracy of six fast algorithms using a range of well-known time-ser…

2018-05-02abs ↗pdf ↗

ISAAC audits deep models for drug-target interactions, revealing structural differences.

problem Deep models for DTI often use irrelevant features, making them hard to evaluate.
method ISAAC uses intervention-based structural auditing to evaluate model sensitivity.
result ISAAC reveals significant structural differences in DTI models' reasoning.

This research evaluates neural network robustness through loss visualization and a new metric.

problem Neural networks' robustness property is insufficiently investigated compared to adversarial attacks and defenses.
method Loss visualization and a new robustness metric to evaluate model stability.
result The proposed robustness metric provides a more reliable evaluation of model stability, uniformed across different models and settings.

Proposes an IRT-based ensemble method to improve machine learning accuracy.

problem Improving the accuracy of machine learning models, especially for hard-to-classify instances.
method Introduces Item Response Theory (IRT) to evaluate sample difficulty and classifier ability, creating three models with different assumptions.
result The proposed IRT ensemble model outperforms other methods on 19 datasets.

New method estimates model performance bounds without ground truth labels.

problem Evaluation of weakly supervised models without direct access to ground truth labels.
method Formulates model evaluation as a partial identification problem and uses Fréchet bounds for performance estimation.
result Derives accurate and computationally efficient bounds for key metrics like accuracy, precision, recall, and F1-score.

This paper evaluates deferring systems using causal inference.

problem Evaluating the impact of deferring systems on model accuracy is challenging.
method The paper uses a causal inference framework to evaluate deferring systems, distinguishing between scenarios with and without access to human predictions.
result The approach allows identifying causal effects of deferring strategies on predictive accuracy.

New metrics CWSA and CWSA+ improve model evaluation under confidence thresholds.

problem Lack of metrics capturing model reliability under confidence thresholds.
method Introducing CWSA and CWSA+ metrics that reward confident accuracy and penalize overconfident mistakes.
result CWSA and CWSA+ outperform classical metrics in trust-sensitive tests.

Proposes a new framework for evaluating diagnostic models with multiple co-primary endpoints.

problem Overoptimistic assessments of predictive performance in automated medical testing devices.
method Multiple testing framework for diagnostic accuracy studies with co-primary endpoints, using a parametric simultaneous test procedure and Bayesian approach to determine optimal number of models.
result Our approach leads to a better final diagnostic model and increased statistical power.

This work benchmarks neural embeddings for link prediction in evolving knowledge graphs.

problem Evaluating the robustness of neural embeddings in changing knowledge graphs.
method Proposes an open-source evaluation pipeline using relation-centric connectivity measures.
result Demonstrates the importance of simulating embedding accuracy for frequently updated knowledge graphs.

New feature mapping approach improves recommendation accuracy and explainability.

problem Balancing recommendation accuracy and explainability using metadata.
method Maps uninterpretable features to interpretable aspect features, minimizing both prediction and interpretation losses.
result Strong performance in recommendation and explainability, eliminating metadata need.

CJE calibrates cheap LLM judges against an oracle, achieving high accuracy at a fraction of the cost.

problem Inexpensive LLM judges can produce biased rankings, leading to unreliable outcomes.
method CJE uses a small oracle to calibrate cheap scores, then evaluates at scale with valid uncertainty.
result CJE achieves 99% pairwise ranking accuracy at 14x lower cost compared to a 16x oracle/judge cost ratio.

This work introduces adversarial sparsity to measure robustness beyond adversarial accuracy.

problem Evaluating robustness to adversarial attacks beyond just accuracy.
method Adversarial sparsity, which quantifies the difficulty of finding perturbations.
result Sparsity provides valuable insights into neural networks and suggests improvements in robustness.

Interpretable companion model for black-box classifiers.

problem Dilemma between interpretable and black-box models.
method Trains a companion model from data and black-box model predictions, optimizing a combination of accuracy and complexity.
result Companion model provides interpretable predictions with a slight accuracy loss for user choice.

A new framework evaluates model performance on single input points, revealing insights into data and model structure.

problem Traditional evaluation methods in machine learning are insufficient for understanding model performance and data structure.
method Developed a pointwise framework to measure model performance on individual input points, analyzing the relationship between pointwise and average performance.
result Profiles of data points reveal different types of correlations between pointwise and average performance, challenging existing models of learning.