Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

126251377502 · Jun 202019922001200920182026
48 results for prediction auditing

Audited Conformal Prediction improves conditional coverage in pretrained models under distribution shift.

problem Uncertainty quantification for pretrained models under unknown distribution shift
method Leverages a small labeled dataset to train an audit model for marginal coverage, integrates outputs into conformal prediction framework
result Significantly higher conditional coverage than existing approaches

RUE audits machine learning predictions for reliability.

problem Ensuring trust in machine learning models for high-stakes applications.
method Resampling uncertainty estimation (RUE) algorithm to audit model reliability after training.
result RUE more effectively detects inaccurate predictions than existing tools.

Develops a method to audit indirect feature influence in complex models.

problem Auditing indirect feature influence in complex, black-box models.
method Disentangled influence audits using disentangled representations.
result Can detect proxy features and show which ones affect model outcomes most.

Study examines auditing fairness in evolving models, identifying strategic updates that preserve audit properties.

problem Auditing fairness in machine learning models that adapt to changing environments.
method Characterizes strategic updates that preserve audit properties, proposes a generic PAC auditing framework.
result Establishes distribution-free auditing bounds for statistical parity using the SP dimension.

Paper tackles transparency and auditability of machine learning in credit scoring.

problem Missed potential in using modern machine learning for credit scoring due to lack of transparency.
method Develops a framework for making black box machine learning models transparent, auditable, and explainable.
result Comparable interpretability can be achieved with machine learning while maintaining predictive power.

ISAAC audits deep models for drug-target interactions, revealing structural differences.

problem Deep models for DTI often use irrelevant features, making them hard to evaluate.
method ISAAC uses intervention-based structural auditing to evaluate model sensitivity.
result ISAAC reveals significant structural differences in DTI models' reasoning.

Develops a technique to audit text-generation models trained on personal data.

problem Enforce data-protection regulations like GDPR and detect unauthorized data usage.
method Black-box auditing method that queries a model to detect if a user's data was used for training.
result Successfully audits well-generalized models without overfitting to training data.

Community moderation drifts towards majority, study finds.

problem How to ensure crowd-sourced moderation systems trust and reward accurate evaluations.
method Consensus-based auditing with a two-stage algorithm that weights contributors by the stability of their past residuals.
result Minority contributors' evaluations drift towards the majority, and their participation share falls on controversial topics.

Signed compression progress on a sealed audit is goodhart-resistant.

problem Intrinsic motivation for agents to improve their world models by compressing experience.
method Rewarding agents for the signed decrease of a fixed sealed-audit loss.
result Cumulative reward telescopes exactly to endpoint audit improvement, preventing infinite reward push while true audit performance stagnates.

AI-driven framework improves enterprise financial audits and risk identification.

problem Manual auditing is inefficient and limited by data complexity and evolving fraud tactics.
method Machine learning algorithms (SVM, RF, KNN) applied to a dataset of audit project counts, violations, and fraud instances.
result Random Forest achieves best performance with F1-score of 0.9012, identifying fraud and compliance anomalies.

Wrapper improves black-box model auditability and decision trustworthiness.

problem Lack of transparency and auditability in machine learning models used in complex applications.
method Integrates uncertainty measures into black-box models to enhance auditability and decision trustworthiness.
result Improves trust in machine learning models by providing actionable mechanisms to reject uncertain predictions.

Study reveals AI skin cancer classifiers underperform for darker skin phototypes, advocating for fairness auditing.

problem AI bias in dermatology, particularly for darker skin phototypes.
method Predictive Representativity (PR) framework, evaluating classifiers on HAM10000 and BOSQUE Test sets.
result Substantial performance disparities by skin phototype, highlighting AI bias.

RLFA estimates misstated monetary fraction with weighted sampling without replacement.

problem Estimating misstated monetary fraction with given accuracy and confidence.
method Developed new confidence sequences for weighted average of unknown values using randomized weighted sampling and side information.
result Adaptive methods improve accuracy of estimates based on side information's predictive power.

JAWS audits predictive uncertainty under covariate shift using jackknife+ weighted methods.

problem Auditing predictive uncertainty under data distribution shifts.
method JAW and JAWA methods for distribution-free uncertainty quantification.
result JAW relaxes the jackknife+'s assumption of data exchangeability for covariate shift.

EL framework certifies and flags bias in ML models without distributional assumptions.

problem Systematic performance disparities across sensitive subpopulations in ML models.
method Empirical likelihood-based approach for non-parametric fairness auditing.
result EL framework outperforms bootstrap methods in certification and subpopulation discovery.

Study shows auditing fairness of personalized interventions is impossible due to unknown ground truths.

problem Auditing fairness of personalized interventions in social services, education, and healthcare.
method Point-identification of quantities under monotone treatment response assumption, providing sensitivity analysis for violations.
result Proves impossibility of auditing fairness using standard metrics and provides methods for auditing using partially-identified ROC and xROC curves.

AWARE-FX uses AI to audit foreign-exchange risk disclosures in corporate reports.

problem Weakly structured foreign-exchange risk disclosures in corporate reports.
method Combines lexicon, logic, encoders, and aggregation methods to convert text into traceable measures.
result FinBERT outperforms in most comparisons, improving F1 scores by up to 0.077.

Signed Evidence Flow (SEF) combines fitted prediction with signed feature attributions to measure evidence conflict and stability.

problem Modern data analysis lacks mechanisms to show the clarity, conflict, or stability of evidence behind predictions.
method Signed Evidence Flow (SEF) combines fitted prediction with signed feature attributions.
result SEF measures conflict and stability, and shows that conflict can improve loss prediction beyond confidence.

Neural networks help auditors efficiently assess financial statements by learning underlying data patterns.

problem Efficiently auditing large volumes of financial statements and journal entries.
method Vector Quantised-Variational Autoencoder (VQ-VAE) neural networks.
result VQ-VAE neural networks can learn a quantized representation of accounting data, uncovering latent factors and providing a representative audit sample.

The paper addresses fairness in online learning by extending auditing schemes and presenting efficient algorithms.

problem Ensuring fairness in online learning while maximizing predictive accuracy.
method Extending auditing schemes to handle multiple auditors and presenting oracle-efficient algorithms.
result Presented algorithms achieve upper bounds on regret and fairness violations, improving on existing bounds.

AI models failed to profitably predict cryptocurrency extrema on Binance Spot.

problem Tackling the profitability of candle-based machine learning models for short-term cryptocurrency trading.
method Scripted fixed-seed model runs and deterministic simulators with human supervision.
result Strongest evidence found negative, with models underperforming buy-and-hold strategies.

Audit fees change based on company and economic factors during auditor switching.

problem Understanding how audit fees change when auditors switch firms.
method Examined the impact of auditor switching on audit fees, considering company characteristics and economic data.
result The direction and magnitude of audit fee changes during switching depend on economic stability and company characteristics.

Data-trained predictive models see widespread use, but for the most part they are used as black boxes which output a prediction or score. It is therefore hard to acquire a deeper understanding of model behavior, and in particular how different features influence the model prediction. This is important when interpreting…

2016-02-23abs ↗pdf ↗

New model predicts dynamic tax evasion with audits and imitation.

problem Static treatment of tax compliance and evasion in Bertotti and Modanese model.
method Piecewise Deterministic Markov Processes (PDMPs) for audits and imitation mechanisms.
result Model shows persistent fluctuations and stationary distribution, not extreme equilibrium.

This study improves audit sampling by using sequential procedures with statistical guarantees.

problem Improving audit efficiency and reliability with statistical methods.
method Formulated as a sequential testing problem, defining null and alternative hypotheses, stopping and decision rules, and exact boundary conditions.
result Exact design yields ex ante control of decision error probabilities, and simulation-based implementation approximates this design.

RESHAPE explains financial statement anomalies by aggregating explanations from AENNs.

problem Detecting and explaining accounting anomalies in financial audits is challenging.
method Proposes RESHAPE to explain model output on an aggregated attribute-level.
result RESHAPE provides more comprehensible explanations compared to existing methods.

Automates summarizing federal grant audits with machine learning.

problem Manual analysis of large federal grant audits is time-consuming and error-prone.
method Sentence clustering, k-means, proximity to centroids, human input for refinement.
result Automated summaries are comparable to human-generated ones using ROUGE metric.

Improved canary crafting for one-run privacy auditing reduces leakage estimates.

problem Detecting canaries in one-run privacy auditing to estimate leakage effectively.
method Optimizes canaries for detectability and diversity, using a greedy initialization and bilevel optimization.
result Achieves stronger leakage estimates at lower computational cost.

OpenAlpha validates decentralized capital strategies using game theory and market aggregation.

problem Decentralized capital management's lack of trust-minimised, adaptive deployment.
method Game-theoretic validation, adversarial auditing, market-based belief aggregation.
result Confidence scores from validation phases inform capital allocation rules.

Fairness audits fail under missing protected labels, especially at zero access.

problem Understanding the reliability of fairness audits with incomplete protected-label data.
method Introduced a seed-calibrated stress test to separate missingness effects from seed-to-seed movement.
result Missing protected labels do not significantly alter fairness mitigation methods, but they can lead to harmful intersectional outcomes.

The paper tackles intersectional fairness in machine learning, proposing methods to assess and mitigate bias.

problem Achieving optimal predictive performance is not enough; fairness with respect to multiple sensitive attributes is crucial.
method The paper introduces a comprehensive framework for auditing and achieving intersectional fairness in classification problems, including metrics, estimation methods, and post-processing techniques.
result The proposed methods can robustly estimate and mitigate intersectional bias in classification models without compromising predictive performance.

New methods target conditional demographic parity using optimal transport distances.

problem Auditing and enforcing conditional demographic parity (CDP) in models with complex conditioning variables.
method Developed novel measures of conditional demographic disparity (CDD) based on optimal transport distances and regularization-based approaches.
result Validated methods airbit{} and airlp{} effectively target CDP in real-world datasets with continuous model outputs.

A new metric GNQ audits LLMs for privacy risks during training.

problem Auditing LLMs for privacy risks during training is computationally hard.
method Gradient Uniqueness (GNQ) metric derived from gradient descent, BS-Ghost GNQ for efficiency.
result GNQ successfully predicts sequence extractability and reveals risk heterogeneity.