Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

10.7%21.5%32.2%42.9% · Jun 202019922001200920172026
48 results for default data

We compare observed corporate cumulative default probabilities to those calculated using a stochastic model based on an extension of the work of Black and Cox and find that corporations default as if via diffusive dynamics. The model, based on a contingent-claims analysis of corporate capital structure, is easily calib…

2000-12-29abs ↗pdf ↗

RMT-Net tackles biased credit scoring data by learning from both default/non-default and rejection/approval tasks.

problem Missing-not-at-random selection bias in financial credit scoring data.
method Reject-aware Multi-Task Network (RMT-Net) that leverages the correlation between default/non-default and rejection/approval tasks.
result RMT-Net improves credit scoring models by learning from both default/non-default and rejection/approval tasks.

Temporal aggregation reveals latent default correlation from monthly data.

problem Understanding effective default correlation from monthly default data.
method Temporal coarse-graining of latent default-probability paths.
result Temporal coarse-graining improves identifiability and reduces over-allocation of long-horizon fluctuations.

Temporal coarse-graining of latent default paths explains effective correlation in corporate defaults.

problem Understanding effective default correlation in corporate defaults.
method Temporal coarse-graining of latent default-probability paths, applied to corporate default-count data.
result Temporal coarse-graining provides a scale-consistent baseline that improves identifiability and reduces over-allocation of long-horizon fluctuations.

New method estimates corporate default probabilities using indirect data.

problem Lack of direct default rate data for corporate companies.
method Modeling default probability dynamics using Bank of Russia overdue debt data.
result Validated method produces trustworthy default probability series.

The paper analyzes Lending Club's loan applicants to predict default risk.

problem Predicting default risk in loan applicants of Lending Club.
method Exploratory data analysis and machine learning (Logistic Regression, Random Forest) were used.
result A credit derivative based on Credit Default Swap was designed to hedge default risk.

Temporal coarse-graining of multi-sector default count data generates effective correlation matrices and rank copulas.

problem Explaining the difference in default dependence between monthly and annual aggregation.
method Dynamic low-rank state-space model with AR(1) latent credit-state factors.
result Effective correlation matrices and rank copulas are generated from monthly default count data.

Study predicts firm defaults using machine learning on Italian credit data.

problem Predicting firm defaults to inform bank lending policies.
method Used large granular credit data from Italian Central Credit Register, combined with public balance sheet data, and applied ensemble techniques and random forest models.
result Ensemble techniques and random forest provide the best results for predicting firm defaults.

Paper uses interbank contagion to predict U.S. bank defaults, finding it highly explanatory.

problem Predicting U.S. bank defaults using interbank contagion.
method Regression and neural network models were used to analyze U.S. commercial bank data.
result Interbank contagion is highly explanatory in default prediction, often outperforming established metrics.

Paper proposes a framework for precise daily default risk prediction of Chinese credit bonds.

problem Inadequate and inaccurate bond information disclosure creates risk of default for investors.
method Framework includes summarizing factors impacting defaults, constructing a risk index system, and using ConvLSTM neural network for prediction.
result The model provides more responsive and accurate daily default risk predictions than authoritative ratings.

Meta-learning symbolic default hyperparameters from dataset properties.

problem Empirical hyperparameter optimization is slow and requires manual configuration.
method Evolutionary algorithm to learn symbolic hyperparameter formulas from dataset properties.
result Meta-learning finds viable symbolic defaults for ML algorithms.

EMDLOT predicts bond defaults better than traditional methods.

problem Lack of interpretability and irregular temporal dependencies in financial data.
method Integrates time-series and textual data, uses Time-Aware LSTM, soft clustering, and multi-level attention.
result EMDLOT outperforms traditional and deep learning benchmarks in recall, F1-score, and mAP.

In this paper we present a novel approach for firm default probability estimation. The methodology is based on multivariate contingent claim analysis and pair copula constructions. For each considered firm, balance sheet data are used to assess the asset value, and to compute its default probability. The asset pricing …

2014-05-06abs ↗pdf ↗

NetDP predicts loan defaults using network data, addressing cold-start issues.

problem Cold-start problem in default prediction for new users.
method Combines unsupervised and supervised network representations, using parameter-server for scalability.
result Effectiveness in cold-start problem, especially for new users.

This paper proposes a deep learning model combining CNN and Transformer for improved credit default prediction.

problem Traditional machine learning models struggle with complex financial data and risk patterns.
method Combines CNN for local feature extraction and Transformer for global dependency modeling.
result The CNN+Transformer model outperforms traditional models in accuracy, AUC, and KS value.

Study integrates climate and text data to improve credit default prediction.

problem Improving credit risk assessment for mSEs with limited financial histories.
method Multimodal framework using LSTM, GRU, and transformer models.
result Integration of multiple data modalities improves credit default prediction.

In the aftermath of the global financial crisis, much attention has been paid to investigating the appropriateness of the current practice of default risk modeling in banking, finance and insurance industries. A recent empirical study by Guo et al.(2008) shows that the time difference between the economic and recorded …

2013-06-27abs ↗pdf ↗

According to theoretical models of valuing risky corporate securities, risk of default is primary component in overall yield spread. However, sizable empirical literature considers it otherwise by giving more importance to non-default risk factors. Current study empirically attempts to provide relative solution to this…

2013-03-14abs ↗pdf ↗

Machine learning improves joint default assessment by capturing non-linear dependencies.

problem Capturing non-linear dependencies among covariates for accurate joint default assessment.
method Application of machine learning techniques to credit card dataset, comparing with logistic regression.
result Machine learning outperforms logistic regression in assessing portfolio riskiness.

The present paper provides a multi-period contagion model in the credit risk field. Our model is an extension of Davis and Lo's infectious default model. We consider an economy of n firms which may default directly or may be infected by other defaulting firms (a domino effect being also possible). The spontaneous defau…

2009-04-10abs ↗pdf ↗

FSL-BDP models time-to-default without centralizing data, improving privacy mechanisms in federated settings.

problem Traditional credit risk models ignore default timing and violate data-protection rules.
method Federated Survival Learning with Bayesian Differential Privacy (FSL-BDP).
result FSL-BDP improves privacy mechanisms in federated settings, outperforming classical DP in most clients.

New model estimates corporate defaults using pure jump processes, capturing extreme events.

problem Estimating corporate defaults using standard diffusion models that underestimate short-term probabilities.
method Introduced pure jump processes with negative jumps only, derived formulas, calibrated parameters, and implemented practical tools.
result Models redistribute credit risk towards shorter maturities, improving short-term default probability estimates.

The paper uses CPI growth rates to improve LGD predictions for CRE loans.

problem Challenges in forecasting LGD for CRE loans due to extended resolution times and restricted data.
method Combines internal and public data, including CPI growth rates, to forecast CRE LGD.
result Incorporating CPI at the time of default improves LGD prediction accuracy.

Shorter time windows and carefully selected features outperform longer periods and extra features in mortgage default prediction.

problem The paradox of increased training data and features leading to worse model performance in time series prediction.
method Empirical study using Fannie Mae's mortgage data, comparing different time window lengths and feature combinations.
result Shorter time windows and carefully selected features yield superior prediction results in mortgage default prediction.

The study uses the Merton model to estimate PD and finds a phase transition affecting convergence speed.

problem Estimating the probability of default (PD) using limited historical data.
method Adopted the Merton model and analyzed phase transitions in default correlation.
result PD estimation converges slowly when temporal correlation decays by power law less than one.

The paper uses daily bond price data to estimate corporate default spreads, improving credit risk assessment.

problem Outdated credit risk information from quarterly accounting items.
method Adapting classic yield curve estimation methods to corporate bonds, using Bayesian estimation.
result High-frequency credit risk proxy via corporate default spreads improves model stability and prediction uncertainty.

We develop a dynamic point process model of correlated default timing in a portfolio of firms, and analyze typical default profiles in the limit as the size of the pool grows. In our model, a firm defaults at a stochastic intensity that is influenced by an idiosyncratic risk process, a systematic risk process common to…

2011-04-10abs ↗pdf ↗

This paper builds a machine learning model to predict credit defaults for unsecured lending.

problem High credit defaults and delinquency rates in unsecured lending due to imbalanced data.
method Employing machine learning techniques, particularly SMOTE for imbalanced data, and evaluating models like LGBM Classifier.
result LGBM Classifier model outperforms other models in predicting credit defaults.

iConViz helps banks manage default contagion risk in networked loans.

problem Managing default contagion risk in networked loans during economic downturns.
method Developed iConViz, an interactive tool, and a novel metric (contagion effect) to quantify and analyze the risk.
result iConViz facilitates closed-loop analysis and helps avoid ad hoc methods.

We propose two structural models for stochastic losses given default which allow to model the credit losses of a portfolio of defaultable financial instruments. The credit losses are integrated into a structural model of default events accounting for correlations between the default events and the associated losses. We…

2012-05-24abs ↗pdf ↗

The paper shows how to calculate risk-neutral default probabilities from bid and ask CDS quotes.

problem Calculating risk-neutral default probabilities from market quotes.
method Using conic finance framework and Poisson process to formulate and solve the calibration problem.
result A unique solution for risk-neutral default probabilities and implied liquidity.