Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

2795588371,116 · Jun 202019922001200920172026
48 results for data breaches

Study analyzes data breach reporting patterns and frequency across U.S. states, finding increasing trends after 2020.

problem Contradictory conclusions in data breach frequency trends due to inconsistent data collection and reporting standards.
method Joint analysis of state Attorneys General's publications on data breaches across eight states with established notification laws.
result Frequency of data breaches is increasing after 2020, with commonalities and heterogeneities across states.

Study shows data breaches cause significant financial losses for firms, especially in health sector.

problem Understanding the economic impact of cyber incidents on listed firms.
method Event study using abnormal returns over 2012-2022, adjusting for event-induced variance and residual cross-correlation.
result Data breaches cause significant financial losses for firms, especially in health sector.

Paper tackles cybersecurity attack detection with an ensemble approach.

problem Challenges in multi-class classification for cyber security breaches.
method Designing a multi-node multi-class classification ensemble approach.
result Proposed approach outperforms full-data approach in multi-node data-censoring cases.

BreachRadar detects points-of-compromise in bank transactions to prevent fraud.

problem Detecting and preventing bank transaction fraud caused by data breaches.
method A distributed alternating algorithm that assigns probabilities to different locations being compromised.
result BreachRadar achieves over 90% precision and recall in detecting compromised cards.

Model for optimal cybersecurity investment considering clustered cyberattacks.

problem Optimal investment in cybersecurity to reduce system vulnerability under clustered cyberattacks.
method Developed a continuous-time stochastic model using a Hawkes process, extended Gordon-Loeb model, solved as a Markovian stochastic optimal control problem.
result Investment policies that account for attack clustering lead to more effective and responsive strategies, improving upon static and Poisson-based approaches.

Proposes a framework to explain KS deterioration in credit risk models.

problem Inconsistent and ad hoc diagnosis of KS decline in credit risk models.
method Counterfactual diagnostic framework attributing KS decline to sampling variability, portfolio composition, covariate shift, and residual deterioration.
result The proposed approach provides more interpretable and governance-relevant explanations than threshold-based review alone.

"How much is my data worth?" is an increasingly common question posed by organizations and individuals alike. An answer to this question could allow, for instance, fairly distributing profits among multiple data contributors and determining prospective compensation when data breaches happen. In this paper, we study the…

2019-02-27abs ↗pdf ↗

C-PP-COAD detects anomalies with limited real data, reducing dependency on real calibration data.

problem Limited real calibration data for online anomaly detection.
method Context-aware prediction-powered conformal online anomaly detection (C-PP-COAD).
result Significantly reduces dependency on real calibration data without compromising FDR control.

Unified AI system for data quality control and governance in regulated environments.

problem Isolated data quality control steps in existing systems.
method AI-driven framework integrating rule-based, statistical, and AI methods.
result Empirical gains in anomaly detection, reduced manual remediation, improved auditability.

Many machine learning applications are based on data collected from people, such as their tastes and behaviour as well as biological traits and genetic data. Regardless of how important the application might be, one has to make sure individuals' identities or the privacy of the data are not compromised in the analysis.…

2016-10-27abs ↗pdf ↗

Kernel two-sample testing is a useful statistical tool in determining whether data samples arise from different distributions without imposing any parametric assumptions on those distributions. However, raw data samples can expose sensitive information about individuals who participate in scientific studies, which make…

2018-08-01abs ↗pdf ↗

The paper proposes using density ratio estimation to evaluate synthetic data quality.

problem Improving the quality and utility of synthetic data for analysis.
method Density ratio estimation to measure synthetic data quality.
result Density ratio estimation yields more accurate global utility estimates than existing methods.

Given a state-of-the-art deep neural network classifier, we show the existence of a universal (image-agnostic) and very small perturbation vector that causes natural images to be misclassified with high probability. We propose a systematic algorithm for computing universal perturbations, and show that state-of-the-art …

2016-10-26abs ↗pdf ↗

Cyber attacks are growing in frequency and severity. Over the past year alone we have witnessed massive data breaches that stole personal information of millions of people and wide-scale ransomware attacks that paralyzed critical infrastructure of several countries. Combating the rising cyber threat calls for a multi-p…

2018-06-08abs ↗pdf ↗

FedPower improves eigenspace estimation privacy in federated learning.

problem Privacy breaches and communication challenges in federated eigenspace estimation.
method FedPower uses a power method with local power iterations and global aggregation, weighted by OPT, and adds Gaussian noise for privacy.
result FedPower provides convergence bounds and demonstrates effectiveness in experiments.
iCurrency?q-fin.GN

We discuss the idea of a purely algorithmic universal world iCurrency set forth in [Kakushadze and Liew, 2014] (https://ssrn.com/abstract=2542541) and expanded in [Kakushadze and Liew, 2017] (https://ssrn.com/abstract=3059330) in light of recent developments, including Libra. Is Libra a contender to become iCurrency? A…

2019-10-28abs ↗pdf ↗

New hybrid model combines GARCH and reinforcement learning for improved VaR estimation.

problem Inaccurate VaR estimation in volatile financial markets.
method Combines GARCH volatility models with DDQN reinforcement learning for dynamic risk forecasting.
result Significant improvement in VaR accuracy and reduction in breaches.

Parameter-transfer is a well-known and versatile approach for meta-learning, with applications including few-shot learning, federated learning, and reinforcement learning. However, parameter-transfer algorithms often require sharing models that have been trained on the samples from specific tasks, thus leaving the task…

2019-09-12abs ↗pdf ↗

Examines AI regulation in finance, highlighting risks and gaps in current laws.

problem Rapid AI adoption in finance introduces risks and compliance challenges.
method Reviews current legislation, industry guidelines, and real-world use cases.
result Need for adaptive, technology-neutral policies to balance innovation and consumer protection.

Study identifies key parameters and input dimensions making LLMs and VLMs brittle.

problem Vulnerability of large language and vision-language models to perturbations.
method Proposed FI measure based on information geometry to quantify sensitivity.
result Small subset of high FI parameters significantly contribute to brittleness.

Enhances cyber risk assessment with entity-specific features.

problem Lack of high-quality public cyber incident data.
method Develops an InsurTech framework to enrich cyber incident data with entity-specific attributes and implements machine learning models.
result InsurTech features improve prediction robustness and provide customized risk profiles.

Measures financial resilience using BSDEs and their properties.

problem Measuring financial resilience in dynamic risk environments.
method Developed stochastic calculus for BSDEs with jumps, revealing resilience rate as expectation of generator.
result Resilience rate can be represented as expectation of BSDE generator, revealing properties of dynamic risk measures.

Paper proposes GPM for simultaneous community detection and group synchronization.

problem Simultaneous community detection and group synchronization in networks.
method Generalized Power Method (GPM) for non-convex optimization.
result GPM achieves exact recovery in O(nlog2n)O(n\log^2n) time, outperforming SDP.

Study examines cyber losses across sectors, finds high severity and frequency.

problem Understanding the nature of cyber losses and their variability across sectors.
method Analysis of a leading industry dataset of cyber events, focusing on frequency and severity.
result Cyber risks are heavy-tailed, with high probability of extreme losses.

Neural networks improve VaR estimation accuracy and robustness.

problem Estimating Value at Risk (VaR) in financial markets.
method Generative regime switching framework with Monte-Carlo simulations, neural networks initialized via best model, balanced incentive function, reduced training data.
result Neural networks outperform traditional methods in VaR estimation, especially with less data.

We develop an optimal currency hedging strategy for fund managers who own foreign assets to choose the hedge tenors that maximize their FX carry returns within a liquidity risk constraint. The strategy assumes that the offshore assets are fully hedged with FX forwards. The chosen liquidity risk metric is Cash Flow at R…

2019-03-15abs ↗pdf ↗

DBNs improve VaR forecasting compared to traditional models, but SVaR forecasts are conservative.

problem Forecasting VaR and SVaR using dynamic Bayesian networks.
method DBN framework applied to S&P 500 index returns, comparing to autoregressive models and historical simulation.
result DBNs achieve comparable VaR forecasting accuracy to historical simulation models, but SVaR forecasts remain conservative.

AI-driven framework improves enterprise financial audits and risk identification.

problem Manual auditing is inefficient and limited by data complexity and evolving fraud tactics.
method Machine learning algorithms (SVM, RF, KNN) applied to a dataset of audit project counts, violations, and fraud instances.
result Random Forest achieves best performance with F1-score of 0.9012, identifying fraud and compliance anomalies.

RL-CVaR model improves insurance reserving under economic stress.

problem Managing insurance reserve setting under claim development uncertainty and macroeconomic stress.
method Reinforcement Learning (PPO) with CVaR constraints, trained under regime-aware curriculum.
result RL-CVaR policy reduces solvency violations and tail-risk compared to classical methods.

Investors with anxiety about drawdowns may use stop-loss and trailing stops as optimal selling strategies.

problem Investors' anxiety about drawdowns affects optimal selling strategies.
method Mathematical analysis of optimal stopping with random discounting.
result Stop-loss and trailing stops can be optimal selling strategies under anxiety about drawdowns.