Audited Conformal Prediction improves conditional coverage in pretrained models under distribution shift.
problem Uncertainty quantification for pretrained models under unknown distribution shift
method Leverages a small labeled dataset to train an audit model for marginal coverage, integrates outputs into conformal prediction framework
result Significantly higher conditional coverage than existing approaches
RUE audits machine learning predictions for reliability.
problem Ensuring trust in machine learning models for high-stakes applications.
method Resampling uncertainty estimation (RUE) algorithm to audit model reliability after training.
result RUE more effectively detects inaccurate predictions than existing tools.
Develops a method to audit indirect feature influence in complex models.
problem Auditing indirect feature influence in complex, black-box models.
method Disentangled influence audits using disentangled representations.
result Can detect proxy features and show which ones affect model outcomes most.
Study examines auditing fairness in evolving models, identifying strategic updates that preserve audit properties.
problem Auditing fairness in machine learning models that adapt to changing environments.
method Characterizes strategic updates that preserve audit properties, proposes a generic PAC auditing framework.
result Establishes distribution-free auditing bounds for statistical parity using the SP dimension.
Paper tackles transparency and auditability of machine learning in credit scoring.
problem Missed potential in using modern machine learning for credit scoring due to lack of transparency.
method Develops a framework for making black box machine learning models transparent, auditable, and explainable.
result Comparable interpretability can be achieved with machine learning while maintaining predictive power.
Audit financial machine learning workflows to detect spurious predictability.
problem Spurious predictability in financial machine learning models.
method Falsification audit testing predictive workflows against synthetic environments.
result Many apparent financial predictions are artifacts, not genuine.
ISAAC audits deep models for drug-target interactions, revealing structural differences.
problem Deep models for DTI often use irrelevant features, making them hard to evaluate.
method ISAAC uses intervention-based structural auditing to evaluate model sensitivity.
result ISAAC reveals significant structural differences in DTI models' reasoning.
Develops a technique to audit text-generation models trained on personal data.
problem Enforce data-protection regulations like GDPR and detect unauthorized data usage.
method Black-box auditing method that queries a model to detect if a user's data was used for training.
result Successfully audits well-generalized models without overfitting to training data.
Develops tools to audit ML models for bias and unfairness.
problem Auditing ML models for individual bias and unfairness.
method Formalizes the task as an optimization problem and develops inferential tools for the optimal value.
result Demonstrates the utility of tools in revealing biases in COMPAS recidivism prediction instrument.
Community moderation drifts towards majority, study finds.
problem How to ensure crowd-sourced moderation systems trust and reward accurate evaluations.
method Consensus-based auditing with a two-stage algorithm that weights contributors by the stability of their past residuals.
result Minority contributors' evaluations drift towards the majority, and their participation share falls on controversial topics.
Signed compression progress on a sealed audit is goodhart-resistant.
problem Intrinsic motivation for agents to improve their world models by compressing experience.
method Rewarding agents for the signed decrease of a fixed sealed-audit loss.
result Cumulative reward telescopes exactly to endpoint audit improvement, preventing infinite reward push while true audit performance stagnates.
AI-driven framework improves enterprise financial audits and risk identification.
problem Manual auditing is inefficient and limited by data complexity and evolving fraud tactics.
method Machine learning algorithms (SVM, RF, KNN) applied to a dataset of audit project counts, violations, and fraud instances.
result Random Forest achieves best performance with F1-score of 0.9012, identifying fraud and compliance anomalies.
New method improves model risk prediction using cross-audit projection.
problem Over-optimism in K-fold CV for binary classification. method Cross-audit projection (CAP) procedure combining resampling and asymptotic bias correction.
result CAP estimator achieves second-order asymptotic unbiasedness.
Wrapper improves black-box model auditability and decision trustworthiness.
problem Lack of transparency and auditability in machine learning models used in complex applications.
method Integrates uncertainty measures into black-box models to enhance auditability and decision trustworthiness.
result Improves trust in machine learning models by providing actionable mechanisms to reject uncertain predictions.
Study reveals AI skin cancer classifiers underperform for darker skin phototypes, advocating for fairness auditing.
problem AI bias in dermatology, particularly for darker skin phototypes.
method Predictive Representativity (PR) framework, evaluating classifiers on HAM10000 and BOSQUE Test sets.
result Substantial performance disparities by skin phototype, highlighting AI bias.
RLFA estimates misstated monetary fraction with weighted sampling without replacement.
problem Estimating misstated monetary fraction with given accuracy and confidence.
method Developed new confidence sequences for weighted average of unknown values using randomized weighted sampling and side information.
result Adaptive methods improve accuracy of estimates based on side information's predictive power.
JAWS audits predictive uncertainty under covariate shift using jackknife+ weighted methods.
problem Auditing predictive uncertainty under data distribution shifts.
method JAW and JAWA methods for distribution-free uncertainty quantification.
result JAW relaxes the jackknife+'s assumption of data exchangeability for covariate shift.
EL framework certifies and flags bias in ML models without distributional assumptions.
problem Systematic performance disparities across sensitive subpopulations in ML models.
method Empirical likelihood-based approach for non-parametric fairness auditing.
result EL framework outperforms bootstrap methods in certification and subpopulation discovery.
Study shows auditing fairness of personalized interventions is impossible due to unknown ground truths.
problem Auditing fairness of personalized interventions in social services, education, and healthcare.
method Point-identification of quantities under monotone treatment response assumption, providing sensitivity analysis for violations.
result Proves impossibility of auditing fairness using standard metrics and provides methods for auditing using partially-identified ROC and xROC curves.
AWARE-FX uses AI to audit foreign-exchange risk disclosures in corporate reports.
problem Weakly structured foreign-exchange risk disclosures in corporate reports.
method Combines lexicon, logic, encoders, and aggregation methods to convert text into traceable measures.
result FinBERT outperforms in most comparisons, improving F1 scores by up to 0.077.
Signed Evidence Flow (SEF) combines fitted prediction with signed feature attributions to measure evidence conflict and stability.
problem Modern data analysis lacks mechanisms to show the clarity, conflict, or stability of evidence behind predictions.
method Signed Evidence Flow (SEF) combines fitted prediction with signed feature attributions.
result SEF measures conflict and stability, and shows that conflict can improve loss prediction beyond confidence.
Neural networks help auditors efficiently assess financial statements by learning underlying data patterns.
problem Efficiently auditing large volumes of financial statements and journal entries.
method Vector Quantised-Variational Autoencoder (VQ-VAE) neural networks.
result VQ-VAE neural networks can learn a quantized representation of accounting data, uncovering latent factors and providing a representative audit sample.
Gen-LRA attacks synthetic data leakage without model knowledge.
problem Auditing synthetic data privacy leakage.
method Generative Likelihood Ratio Attack (Gen-LRA).
result Gen-LRA outperforms other attacks across metrics.
New approach to meaningful and robust algorithmic recourse.
problem Ineffective and unmeaningful algorithmic recourse explanations.
method Meaningful Algorithmic Recourse (MAR) and Effective Algorithmic Recourse (EAR).
result Proposes new constraints for algorithmic recourse that improve both prediction and target.
The paper addresses fairness in online learning by extending auditing schemes and presenting efficient algorithms.
problem Ensuring fairness in online learning while maximizing predictive accuracy.
method Extending auditing schemes to handle multiple auditors and presenting oracle-efficient algorithms.
result Presented algorithms achieve upper bounds on regret and fairness violations, improving on existing bounds.
AI models failed to profitably predict cryptocurrency extrema on Binance Spot.
problem Tackling the profitability of candle-based machine learning models for short-term cryptocurrency trading.
method Scripted fixed-seed model runs and deterministic simulators with human supervision.
result Strongest evidence found negative, with models underperforming buy-and-hold strategies.
Audit fees change based on company and economic factors during auditor switching.
problem Understanding how audit fees change when auditors switch firms.
method Examined the impact of auditor switching on audit fees, considering company characteristics and economic data.
result The direction and magnitude of audit fee changes during switching depend on economic stability and company characteristics.
Big data transforms accounting and auditing, enhancing insights but posing challenges.
problem Challenges in data privacy and security with increased data sources.
method Utilizing AI and machine learning for efficient data analysis and anomaly detection.
result Enhanced analytics tools and continuous learning are key to overcoming challenges.
New method audits DP guarantees without noise or subsampling info.
problem Auditing DP guarantees of ML models without prior info.
method Histogram-based density estimation for lower bounds.
result Natural generalization of membership inference auditing.
Data-trained predictive models see widespread use, but for the most part they are used as black boxes which output a prediction or score. It is therefore hard to acquire a deeper understanding of model behavior, and in particular how different features influence the model prediction. This is important when interpreting…
New model predicts dynamic tax evasion with audits and imitation.
problem Static treatment of tax compliance and evasion in Bertotti and Modanese model.
method Piecewise Deterministic Markov Processes (PDMPs) for audits and imitation mechanisms.
result Model shows persistent fluctuations and stationary distribution, not extreme equilibrium.
Paper uses 2-step Gradient Boosting to predict VAT tax gap.
problem Estimating tax evasion and revenue loss from tax avoidance.
method 2-steps Gradient Boosting model to correct selection bias.
result Significantly improved prediction of VAT tax gap.
This study improves audit sampling by using sequential procedures with statistical guarantees.
problem Improving audit efficiency and reliability with statistical methods.
method Formulated as a sequential testing problem, defining null and alternative hypotheses, stopping and decision rules, and exact boundary conditions.
result Exact design yields ex ante control of decision error probabilities, and simulation-based implementation approximates this design.
RESHAPE explains financial statement anomalies by aggregating explanations from AENNs.
problem Detecting and explaining accounting anomalies in financial audits is challenging.
method Proposes RESHAPE to explain model output on an aggregated attribute-level.
result RESHAPE provides more comprehensible explanations compared to existing methods.
Automates summarizing federal grant audits with machine learning.
problem Manual analysis of large federal grant audits is time-consuming and error-prone.
method Sentence clustering, k-means, proximity to centroids, human input for refinement.
result Automated summaries are comparable to human-generated ones using ROUGE metric.
New divergences help audit DP in high dimensions.
problem Challenges in auditing DP in high-dimensional data.
method Propose kernel Rényi divergence and its regularized version for auditing.
result Regularized kernel Rényi divergence can be estimated from samples in high dimensions.
Improved canary crafting for one-run privacy auditing reduces leakage estimates.
problem Detecting canaries in one-run privacy auditing to estimate leakage effectively.
method Optimizes canaries for detectability and diversity, using a greedy initialization and bilevel optimization.
result Achieves stronger leakage estimates at lower computational cost.
OpenAlpha validates decentralized capital strategies using game theory and market aggregation.
problem Decentralized capital management's lack of trust-minimised, adaptive deployment.
method Game-theoretic validation, adversarial auditing, market-based belief aggregation.
result Confidence scores from validation phases inform capital allocation rules.
Fine-tuning LLMs on privacy-sensitive data introduces privacy risk, and synthetic data audits can quantify this risk.
problem Fine-tuning LLMs on privacy-sensitive data introduces privacy risk.
method Generate synthetic canaries via high-temperature sampling from LLMs.
result Synthetic canaries are high-influence outliers that ensure strong audits.
Fairness audits fail under missing protected labels, especially at zero access.
problem Understanding the reliability of fairness audits with incomplete protected-label data.
method Introduced a seed-calibrated stress test to separate missingness effects from seed-to-seed movement.
result Missing protected labels do not significantly alter fairness mitigation methods, but they can lead to harmful intersectional outcomes.
The paper tackles intersectional fairness in machine learning, proposing methods to assess and mitigate bias.
problem Achieving optimal predictive performance is not enough; fairness with respect to multiple sensitive attributes is crucial.
method The paper introduces a comprehensive framework for auditing and achieving intersectional fairness in classification problems, including metrics, estimation methods, and post-processing techniques.
result The proposed methods can robustly estimate and mitigate intersectional bias in classification models without compromising predictive performance.
Study efficient auditing of ML fairness models.
problem Scalability of auditing ML models for fairness.
method Query-based auditing algorithms for estimating demographic parity.
result Optimal deterministic and practical randomized algorithms for fairness estimation.
Study cost-effective fairness audits with partial feedback, improving over random exploration.
problem Auditing fairness of classifiers with limited true labels.
method Introduces cost model, proposes near-optimal algorithms for black-box and mixture models.
result Significantly lower audit costs compared to natural baselines.
Survey of determinism issues in financial AI systems.
problem Vulnerabilities in reproducibility of financial AI systems.
method Literature review and first-party experiments on public financial datasets.
result Proposed a layered evaluation framework linking modality-specific metrics to audit readiness.
Algorithm identifies best arm with biased proxy and selective ground truth audits.
problem Fixed-confidence best-arm identification with biased proxy and selective ground truth.
method Propensity-weighted estimator and adaptive auditing algorithm.
result Plug-in Neyman rule achieves near-oracle audit efficiency.
Proposes PA-DSL for correcting noisy human labels in automated data labeling.
problem Noisy human labels in automated data labeling.
method Uses adjudicated cases to correct noisy human labels and debias analyses.
result Maintains nominal coverage and reduces RMSE by 10-17% relative to using only adjudicated labels.
New methods target conditional demographic parity using optimal transport distances.
problem Auditing and enforcing conditional demographic parity (CDP) in models with complex conditioning variables.
method Developed novel measures of conditional demographic disparity (CDD) based on optimal transport distances and regularization-based approaches.
result Validated methods airbit{} and airlp{} effectively target CDP in real-world datasets with continuous model outputs.
A new metric GNQ audits LLMs for privacy risks during training.
problem Auditing LLMs for privacy risks during training is computationally hard.
method Gradient Uniqueness (GNQ) metric derived from gradient descent, BS-Ghost GNQ for efficiency.
result GNQ successfully predicts sequence extractability and reveals risk heterogeneity.