Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

1345 · May 202619922001200920172026
48 results for IRB oversight

Machine learning uses crowdworkers; determining their status as human subjects is tricky.

problem Determining the appropriate status of ML crowdworkers as human subjects.
method Investigation of natural language processing studies to expose challenges and propose solutions.
result Potential loophole in the U.S. Common Rule for ML research oversight.

The Basel II internal ratings-based (IRB) approach to capital adequacy for credit risk plays an important role in protecting the Australian banking sector against insolvency. We outline the mathematical foundations of regulatory capital for credit risk, and extend the model specification of the IRB approach to a more g…

2014-12-03abs ↗pdf ↗

The 2007--2008 financial crisis has paved the way for the use of macroprudential policies in supervising the financial system as a whole. This paper views macroprudential oversight in Europe as a process, a sequence of activities with the ultimate aim of safeguarding financial stability. To conceptualize a process in t…

2013-12-29abs ↗pdf ↗

Paper introduces PHI to identify structurally distinct payment patterns in UK municipal procurement.

problem Vulnerability of public procurement to error, fraud, and corruption in high-volume transactions.
method Introduces Payment Heterogeneity Index (PHI) using Gaussian Mixture Model (GMM) and non-parametric statistics.
result Identifies a significant cohort with structurally distinct payment patterns, improving procurement oversight.

LLMs struggle with financial reasoning but can outperform the market with human oversight.

problem Financial reasoning failures in LLM-generated stock market predictions.
method Evaluated four LLMs using three prompting strategies and compared to human oversight.
result LLMs require human oversight to fully realize their potential in financial markets.

The purpose of this paper is to identify a relevant statistical correlation between rate of default, RD, and loss given default, LGD, in a major Brazilian financial institution Retail Home Equity exposure rated using the IRB approach, so that we may find a causal relationship between the two risk parameters. Therefore,…

2014-08-03abs ↗pdf ↗

AI agents manage portfolios, improving on human oversight.

problem Improving strategic asset allocation for institutional investors.
method 50 specialized agents produce capital market assumptions, construct portfolios, critique, and vote on each other's output.
result Meta-agent compares forecasts with realized returns and improves agent performance.

We prove a structure theorem for closed topological manifolds of cohomogeneity one; this result corrects an oversight in the literature. We complete the equivariant classification of closed, simply connected cohomogeneity one topological manifolds in dimensions 55, 66, and 77 and obtain topological characterizations…

2015-03-31abs ↗pdf ↗

This paper discusses the role of risk communication in macroprudential oversight and of visualization in risk communication. Beyond the soar in data availability and precision, the transition from firm-centric to system-wide supervision imposes vast data needs. Moreover, except for internal communication as in any orga…

2014-04-17abs ↗pdf ↗

No policy can simultaneously be fully autonomous, optimally calibrated, and helpful, proving a trilemma.

problem Proving impossibility of a policy achieving maximum helpfulness, optimal calibration, and full autonomy.
method Geometric proof showing that adding any non-affine autonomy incentive to a strictly proper scoring rule destroys strict properness.
result The Behavioral Credibility Trilemma: no policy can achieve all three goals simultaneously.

LR-Robot automates SLRs with AI, expert oversight, and multidimensional analysis.

problem Efficient but contextually limited outputs from existing SLR frameworks.
method Human-in-the-loop process, structured knowledge sources, retrieval-augmented generation.
result Empirical demonstration of AI-driven literature synthesis in option pricing.

This paper uses deep learning to detect money laundering in cross-border transactions.

problem Detecting money laundering in cross-border transactions is challenging due to complexity and scale.
method Unsupervised learning models, including CNNs and hybrid CNNGRU architectures, were tested for anomaly detection.
result Hybrid Convolutional-Recurrent Neural Integration Model (CRNIM) showed superior performance.

LR-Robot accelerates SLRs by combining expert oversight and AI, revealing trends and patterns in financial research.

problem Manual SLRs are impractical due to the scale and complexity of modern financial research.
method Domain experts define taxonomies and constraints, LLMs execute classification, and human evaluation ensures reliability.
result AI can understand and synthesize literature, revealing trends and core research directions.

Model proposes how regulators should oversee complex algorithms in high-stakes applications.

problem Regulating complex algorithms used in high-stakes applications like lending, testing, and hiring.
method Proposes a model where regulators are limited in learning about complex algorithms with misaligned preferences, and explores different regulatory approaches.
result Complex algorithms can improve welfare, but regulation should focus on the source of incentive misalignment for optimal results.

Paper presents a method for estimating long-term PDs with incomplete data.

problem Estimating long-term PDs with limited and incomplete historical data.
method Single risk factor approach for simultaneous calibration of PDs across sub-portfolios.
result Method yields long-term PDs without requiring complete historical data.

AI systems that explain their decisions can be monitored for harmful intentions.

problem Monitoring AI systems' decision-making processes for harmful intentions is imperfect and can miss some misbehavior.
method Monitoring the chain of thought (CoT) of AI systems that communicate in human language.
result CoT monitoring is a promising but fragile approach to AI safety.

Interpretable classification models are built with the purpose of providing a comprehensible description of the decision logic to an external oversight agent. When considered in isolation, a decision tree, a set of classification rules, or a linear model, are widely recognized as human-interpretable. However, such mode…

2018-10-22abs ↗pdf ↗

This paper proposes a method to use LLMs as auxiliary evaluators in place of human judges.

problem The need for cost-effective and scalable evaluation of AI systems.
method Formulates a two-stage sampling design with LLM evaluations and human ratings, using a doubly robust estimator.
result Proposes a method to determine optimal sample sizes for human and LLM ratings.

Classical stochastic gradient methods for optimization rely on noisy gradient approximations that become progressively less accurate as iterates approach a solution. The large noise and small signal in the resulting gradients makes it difficult to use them for adaptive stepsize selection and automatic stopping. We prop…

2016-10-18abs ↗pdf ↗

The proliferation of large data sets and Bayesian inference techniques motivates demand for better data sparsification. Coresets provide a principled way of summarizing a large dataset via a smaller one that is guaranteed to match the performance of the full data set on specific problems. Classical coresets, however, n…

2018-05-18abs ↗pdf ↗

The paper studies how to allocate human validation in AI-assisted tasks to minimize errors.

problem Heterogeneous reliability of AI-generated signals across tasks, products, and customer segments.
method Tuned prediction-powered inference, upper confidence bounds policy, Neyman square-root rule.
result The proposed policy outperforms uniform and epsilon-greedy allocation, closing most of the gap to the oracle when reliability is heterogeneous.

Paper proposes hybrid approach for transparent credit scoring models.

problem Lack of transparency in machine learning models limits their use in regulated environments.
method Post-hoc interpretation of black-box models guides feature selection, followed by training glass-box models.
result Reduces feature usage from 106 to 10 while maintaining comparable performance.

For sophisticated reinforcement learning (RL) systems to interact usefully with real-world environments, we need to communicate complex goals to these systems. In this work, we explore goals defined in terms of (non-expert) human preferences between pairs of trajectory segments. We show that this approach can effective…

2017-06-12abs ↗pdf ↗

No fair and strategy-proof automated market maker exists for more than two assets.

problem Designing a fair and strategy-proof automated market maker for multiple assets.
method Analyzing the weighted-product family of aggregation rules and their properties.
result No aggregation rule is both fair and strategy-proof for more than two assets.

The paper investigates ethical issues in large image datasets, focusing on pornographic content.

problem Ethical issues in large-scale computer vision datasets, particularly concerning pornographic content.
method Cross-sectional model-based quantitative census covering various factors in the ImageNet-ILSVRC-2012 dataset.
result The dataset contains verifiably pornographic images, including non-consensual and voyeuristic content.

LSBI approximates likelihood with linear functions for cosmological parameter estimation.

problem Estimating cosmological parameters from complex data.
method Sequential Linear Simulation-based Inference (LSBI) using Gaussian approximations.
result LSBI achieves convergence after 4-5 rounds of simulations, comparable to neural methods.

Reinforcement learning is a promising approach to developing hard-to-engineer adaptive solutions for complex and diverse robotic tasks. However, learning with real-world robots is often unreliable and difficult, which resulted in their low adoption in reinforcement learning research. This difficulty is worsened by the …

2018-03-19abs ↗pdf ↗

We embed arbitrary groups into regular graphs with prescribed automorphisms.

problem Embedding arbitrary groups into regular graphs with specific automorphisms.
method Constructing regular graphs with strong embeddings and automorphism groups isomorphic to any given finite group.
result For every d3d\geq 3 and every finite group GG, there exists a dd-regular graph ΓΓ with a strong embedding ββ such that Aut(Γ)Aut(β(Γ))G\mathrm{Aut}(Γ) \cong \mathrm{Aut}(β(Γ)) \cong G.

New framework detects crypto wash trading using liquidity measures.

problem Detecting and monitoring wash trading in crypto assets.
method Developed a new framework to detect wash trading through real-time liquidity fluctuation measures.
result Joint elevation in liquidity jump and diffusion indicates wash trading in crypto assets.

TraCeR uses transformers to analyze survival data with longitudinal covariates.

problem Handling longitudinal covariates and assessing model calibration in survival analysis.
method Transformer-based survival analysis framework with factorized self-attention architecture.
result TraCeR achieves significant performance improvements over state-of-the-art methods.

AutoML systems are currently rising in popularity, as they can build powerful models without human oversight. They often combine techniques from many different sub-fields of machine learning in order to find a model or set of models that optimize a user-supplied criterion, such as predictive performance. The ultimate g…

2019-08-28abs ↗pdf ↗

Proposes a new stochastic method to calibrate climate risks in financial models.

problem Estimating climate-related financial risks in bank loan portfolios.
method Stochastic forward-looking methodology to calibrate climate macro-correlation evolution from scientific data.
result A new framework to evaluate climate risks without specific scenario assumptions.

Study shows GPT's earnings forecasts are human-like but not always accurate.

problem Information friction in AI-generated financial analysis.
method Examined GPT's earnings forecasts following corporate earnings releases and proposed a diagnostic framework.
result GPT's narrative attention is consistent and human-like but not always associated with higher forecast accuracy.

Hybrid Bayesian-conformal framework improves uncertainty quantification in healthcare predictions.

problem Jointly satisfying distribution-free coverage guarantees and risk-adaptive precision in clinical decision-making.
method Integrates Bayesian hierarchical random forests with group-aware conformal calibration, using posterior uncertainties to weight conformity scores.
result Achieves target coverage (94.3% vs 95% target) with adaptive precision, 21% narrower intervals for low-uncertainty cases.

Proposes a value-oriented forecast reconciliation method for renewables in electricity markets.

problem Forecast reconciliation overlooks the value of forecasts in decision-making, leading to unfair outcomes.
method Value-oriented forecast reconciliation using a Nash bargaining framework and a primal-dual algorithm for parameter estimation.
result Consistently increases profits for all agents involved in an aggregated wind energy trading problem.

Preterm birth is the most common cause of neonatal death. Current diagnostic methods that assess the risk of preterm birth involve the collection of maternal characteristics and transvaginal ultrasound imaging conducted in the first and second trimester of pregnancy. Analysis of the ultrasound data is based on visual i…

2019-08-24abs ↗pdf ↗

HyFi cryptocurrencies backed by institutions show lower price risk than fully decentralized ones.

problem High volatility in decentralized finance (DeFi) cryptocurrencies.
method Panel EGLS models with fixed, random, and dynamic specifications using daily data for 18 major cryptocurrencies.
result HyFi-like assets exhibit lower price risk, especially during market stress.

The paper assesses fairness in risk score models, focusing on epistemic value.

problem Fairness of risk score models in communicating uncertainty.
method Identified key fairness desiderata, developed metrics for quantitative assessment, and applied methodology in two case studies.
result Introduced a novel calibration error metric for meaningful comparisons between groups of different sizes.

STOOD-X detects out-of-distribution samples without distributional assumptions and provides explainable visualizations.

problem Challenges in OOD detection, including restrictive assumptions, scalability issues, and lack of interpretability.
method Two-stage methodology combining statistical nonparametric test and explainability enhancements.
result Achieves competitive performance in high-dimensional and complex settings, with explainability framework enabling human oversight.