Model assesses credit risk using behavioral data from Experian and Bank of Italy.
problem Improving credit risk assessment in financial institutions.
method Statistical and machine learning techniques applied to behavioral data from Experian and Bank of Italy.
result Demonstrates transferability of the model from private to central data.
Study uses synthetic data to estimate credit risk for underbanked consumers in Istanbul.
problem Estimating credit risk for underbanked consumers lacking formal credit records.
method Created synthetic dataset, used retrieval augmented generation, trained CatBoost, LightGBM, and XGBoost models.
result Alternative financial data improves credit risk estimation, raising AUC by 13%.
AI improves MSME credit scoring using bank statement data.
problem Lack of access to financing for MSMEs due to traditional credit scoring methods.
method Developed a cash flow-based pipeline using bank statement data for machine learning credit scoring.
result Bank statement features significantly improve credit scoring models, achieving AUROC of 0.806.
Many households in developing countries lack formal financial histories, making it difficult for firms to extend credit, and for potential borrowers to receive it. However, many of these households have mobile phones, which generate rich data about behavior. This article shows that behavioral signatures in mobile phone…
Alternative app data improves credit scoring for underserved borrowers.
problem Improving credit scoring for low-wealth and young individuals.
method Use of alternative data from app-based marketplaces, validated with TreeSHAP method.
result Alternative data sources predict financial behavior better than traditional bureau data.
New method estimates corporate default probabilities using indirect data.
problem Lack of direct default rate data for corporate companies.
method Modeling default probability dynamics using Bank of Russia overdue debt data.
result Validated method produces trustworthy default probability series.
DAISYnt evaluates synthetic data quality and privacy in regulated domains.
problem Balancing data quality and privacy in regulated domains.
method Developed a suite of advanced tests (DAISYnt) to evaluate synthetic data quality and privacy.
result DAISYnt sets a de facto standard for synthetic data evaluation in regulated domains.
Paper compares AI models for credit scoring and explains them.
problem Lack of interpretability in advanced AI models hinders credit risk management.
method Comparison of logistic regression, AI algorithms, and techniques to interpret AI models.
result Advanced tree-based models provide the best prediction of client default.
LDA-XGB1 balances fairness and accuracy in lending models.
problem Fair lending practices and model interpretability in binary classification.
method Biobjective optimization using binning and information value, leveraging XGBoost.
result Achieves effective balance between accuracy, fairness, and interpretability.
Paper uses Super-App data to improve income estimation models.
problem Improving accuracy of income estimation models.
method TreeSHAP method for Stochastic Gradient Boosting Interpretation.
result Alternative data from Super-Apps capture more information than traditional financial data.
The US Census Bureau corrupts data to protect privacy, but we show how to clean and analyze it effectively.
problem Analyzing Census data with intentional corruption to maintain privacy.
method Formulated a semiparametric model, proposed data cleaning, estimation, and inference procedures.
result Demonstrated that data cleaning can maintain precision and provided theoretical and empirical support.
In this article we present an alternative model for the distribution of household incomes in the United States. We provide arguments from two differing perspectives which both yield the proposed income distribution curve, and then fit this curve to empirical data on household income distribution obtained from the Unite…
2020 Census uses more noise to protect privacy than needed, improving data accuracy.
problem Ensuring privacy in census data while maintaining accuracy for policy decisions.
method Applied f-differential privacy to track and reduce noise across geographical levels. result The 2020 Census provides stronger privacy protections than its nominal guarantees suggest.
This paper presents approaches to determine a network based pricing for 3D printing services in the context of a two-sided manufacturing-as-a-service marketplace. The intent is to provide cost analytics to enable service bureaus to better compete in the market by moving away from setting ad-hoc and subjective prices. A…
Method debiases alternative data for fair credit underwriting.
problem Bias in alternative data affecting credit underwriting fairness.
method Causal inference applied to machine learning models.
result Improves model accuracy across racial groups without discrimination.
Large corporate credit models may be adapted for small business risk assessment.
problem Limited data and lack of credit analysts for small businesses.
method Adapting large corporate credit risk models for small businesses.
result Adapted models can predict small business credit risk effectively.
CCR-CNN uses CNN to predict corporate credit ratings from financial data.
problem Lack of data and limited model performance in predicting corporate credit ratings.
method Transform corporations into images and use CNN to analyze complex feature interactions.
result CCR-CNN outperforms state-of-the-art methods in predicting corporate credit ratings.
Framework integrates financial and annual report data for better corporate credit ratings.
problem Lack of insights from non-financial data in credit rating models.
method Uses FinBERT to extract features from annual reports and combines them with financial data.
result Improves credit rating accuracy by 8-12%.
A new algorithm improves credit scoring accuracy for imbalanced data.
problem Poor classification of minority class in credit scoring data sets.
method Weighted-Hybrid-Sampling-Boost (WHSBoost) algorithm with balanced data sampling.
result WHSBoost outperforms other methods in credit scoring accuracy.
This paper develops a machine learning model to assess credit risk in UAE commercial banks.
problem Lack of precision in conventional credit rating tools for accurate credit risk prediction.
method Constructs a credit risk assessment model using Linear Discriminant Analysis.
result Demonstrates improved accuracy in predicting good and bad creditors compared to conventional methods.
BSAC improves credit scoring models by leveraging autoencoders and addressing imbalanced datasets.
problem Imbalanced and heterogeneous credit scoring datasets.
method Bagging Supervised Autoencoder Classifier (BSAC) that uses autoencoders and undersampling.
result BSAC improves classification of loan applicants, demonstrating robustness and effectiveness.
The paper analyzes Lending Club's loan applicants to predict default risk.
problem Predicting default risk in loan applicants of Lending Club.
method Exploratory data analysis and machine learning (Logistic Regression, Random Forest) were used.
result A credit derivative based on Credit Default Swap was designed to hedge default risk.
Study integrates climate and text data to improve credit default prediction.
problem Improving credit risk assessment for mSEs with limited financial histories.
method Multimodal framework using LSTM, GRU, and transformer models.
result Integration of multiple data modalities improves credit default prediction.
Study shows how macroprudential policies affect credit growth in Israel, especially in housing and business sectors.
problem Impact of macroprudential policies on credit growth in Israel.
method Bank-level panel data analysis for Israel, 2004-2019; interaction of monetary and macroprudential policies.
result Accommodative monetary policy interacts with macroprudential policies to increase total credit growth.
Paper uses LightGBM for mobile user credit assessment.
problem Improving credit evaluation methods for communication operators.
method Data preprocessing, feature engineering, multiple machine learning models integration.
result Established a suitable fusion model for operator user credit evaluation.
Efficiently calculates privacy guarantees for 2020 Census data.
problem Evaluate privacy guarantees for 2020 U.S. Census data releases.
method Sieve-accelerated quadrature method to evaluate tail probabilities of high-dimensional convolutions.
result Achieves 1,824-fold speedup over prior methods while maintaining error tolerances.
Synthetic data improves credit scoring models' performance without compromising borrower privacy.
problem Scarcity of real data for credit scoring models due to privacy concerns.
method Privacy-preserving training with synthetic data.
result Credit scoring models trained with synthetic data show a reduction of 3% in AUC and 6% in KS compared to real data models.
This paper builds a machine learning model to predict credit defaults for unsecured lending.
problem High credit defaults and delinquency rates in unsecured lending due to imbalanced data.
method Employing machine learning techniques, particularly SMOTE for imbalanced data, and evaluating models like LGBM Classifier.
result LGBM Classifier model outperforms other models in predicting credit defaults.
Paper proposes an intelligent credit limit management system using causal inference.
problem Traditional credit limit management strategies are heuristic and not data-driven.
method Conditional independence testing, response model, log transformation, GBDT encoding, non-linear transformation on features, well-designed metric.
result The proposed approach effectively manages credit limits and incorporates diminishing marginal effects.
Credit scores misclassify borrowers, especially minorities, leading to inequitable access.
problem Misclassification of borrowers by credit scores, particularly minorities.
method Benchmarked a widely used credit score against a machine learning model.
result Machine learning model improves predictive accuracy for low-quality data, leading to more equitable access.
NetDP predicts loan defaults using network data, addressing cold-start issues.
problem Cold-start problem in default prediction for new users.
method Combines unsupervised and supervised network representations, using parameter-server for scalability.
result Effectiveness in cold-start problem, especially for new users.
Paper proposes a method to evaluate SME credit risk using meta paths.
problem Evaluate credit risk of small and medium-sized enterprises with limited data.
method Exploits the representative power of information networks and meta paths to infer SME financial status.
result Meta path feature effectively identifies SMEs with credit risks.
AI enhances bank credit risk management through deep learning and data analysis.
problem Inaccurate credit decisions and potential risks in bank credit risk management.
method Innovative application of AI technology, including deep learning and big data analysis.
result AI provides more accurate and comprehensive credit decision support, reducing risks and losses.
Paper uses RL for better credit scoring and underwriting.
problem Traditional underwriting methods are ungeneralizable in complex scenarios.
method Adapts RL principles for credit scoring, incorporating action space renewal and multi-choice actions.
result RL-based algorithms outperform traditional methods in aligned data scenarios.
Deep learning improves credit risk assessment without new data.
problem Improving credit risk assessment in banking without new data.
method Sequential deep learning using temporal convolutional networks.
result Sequential deep learning outperformed tree-based models in credit risk assessment.
Among other macroeconomic indicators, the monthly release of U.S. unemployment rate figures in the Employment Situation report by the U.S. Bureau of Labour Statistics gets a lot of media attention and strongly affects the stock markets. I investigate whether a profitable investment strategy can be constructed by predic…
Paper applies NEAT for dynamic credit evaluation using streaming data.
problem Dynamic credit evaluation using streaming data.
method Neuroevolution of Augmenting Topologies (NEAT) with enhancements.
result NEAT effectively handles dynamic credit evaluation with streaming data.
Method determines credit transition matrix from cumulative default probabilities.
problem Quantifying changes in bond credit ratings.
method Setup an ill-posed, linear inverse problem with entropy minimization.
result Method successfully determines CTM from cumulative default probabilities.
This paper reviews LLMs for credit risk assessment, creating a taxonomy.
problem Assessing credit risk using financial text analysis.
method Systematic review of 60 papers, focusing on model architectures, data types, and explainability mechanisms.
result Developed a taxonomy of LLM-based credit risk models.
This paper uses PCA and FA for feature selection in credit rating.
problem Selecting important features for credit rating prediction.
method Principal Component Analysis and Factor Analysis.
result Factor Analysis reduces feature set significantly without losing much accuracy.
Bayesian and simulation methods predict credit default probabilities.
problem Assessing credit risk in large customer portfolios.
method Two-phase approach: Bayesian estimation followed by Monte Carlo simulations.
result Estimation of true default rates through simulations.
New method learns credit prices offline without interaction.
problem Dynamic pricing of consumer credit.
method Offline deep reinforcement learning with Q-Learning.
result Effective personalized pricing policy learned without online interaction.
Study finds public procurement awards, especially NGEU-funded ones, boost new lending.
problem Understanding the impact of public procurement on new lending.
method Panel data local projections model, controlling for various factors.
result Public procurement awards, particularly NGEU-funded ones, significantly increase new lending.
Flexible models predict US Census survey response rates.
problem Predicting survey response rates in the US Census Bureau.
method Nonparametric additive models with structured interactions using ℓ0-based penalization.
result Models lead to predictions comparable to black-box methods but are interpretable.
Credit scoring is without a doubt one of the oldest applications of analytics. In recent years, a multitude of sophisticated classification techniques have been developed to improve the statistical performance of credit scoring models. Instead of focusing on the techniques themselves, this paper leverages alternative d…
Deep learning predicts employment changes and industry health.
problem Forecasting short-term employment changes and assessing long-term industry health.
method LSTNet, a multi-scale deep learning model, processes multivariate time series data.
result LSTNet outperforms baseline models in most sectors, especially stable ones.
Credit risk prediction is an effective way of evaluating whether a potential borrower will repay a loan, particularly in peer-to-peer lending where class imbalance problems are prevalent. However, few credit risk prediction models for social lending consider imbalanced data and, further, the best resampling technique t…
Derives metrics for DeFi vaults, addressing credit risk.
problem Credit risk in DeFi lending vaults.
method Three-level decomposition of vault risk; six structural features identified.
result Estimation architecture for credit risk metrics.