Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

2815638441,125 · Jun 202019922001200920172026
48 results for Credit Bureau Data

Model assesses credit risk using behavioral data from Experian and Bank of Italy.

problem Improving credit risk assessment in financial institutions.
method Statistical and machine learning techniques applied to behavioral data from Experian and Bank of Italy.
result Demonstrates transferability of the model from private to central data.

Study uses synthetic data to estimate credit risk for underbanked consumers in Istanbul.

problem Estimating credit risk for underbanked consumers lacking formal credit records.
method Created synthetic dataset, used retrieval augmented generation, trained CatBoost, LightGBM, and XGBoost models.
result Alternative financial data improves credit risk estimation, raising AUC by 13%.

AI improves MSME credit scoring using bank statement data.

problem Lack of access to financing for MSMEs due to traditional credit scoring methods.
method Developed a cash flow-based pipeline using bank statement data for machine learning credit scoring.
result Bank statement features significantly improve credit scoring models, achieving AUROC of 0.806.

Many households in developing countries lack formal financial histories, making it difficult for firms to extend credit, and for potential borrowers to receive it. However, many of these households have mobile phones, which generate rich data about behavior. This article shows that behavioral signatures in mobile phone…

2017-12-09abs ↗pdf ↗

Alternative app data improves credit scoring for underserved borrowers.

problem Improving credit scoring for low-wealth and young individuals.
method Use of alternative data from app-based marketplaces, validated with TreeSHAP method.
result Alternative data sources predict financial behavior better than traditional bureau data.

New method estimates corporate default probabilities using indirect data.

problem Lack of direct default rate data for corporate companies.
method Modeling default probability dynamics using Bank of Russia overdue debt data.
result Validated method produces trustworthy default probability series.

DAISYnt evaluates synthetic data quality and privacy in regulated domains.

problem Balancing data quality and privacy in regulated domains.
method Developed a suite of advanced tests (DAISYnt) to evaluate synthetic data quality and privacy.
result DAISYnt sets a de facto standard for synthetic data evaluation in regulated domains.

Paper compares AI models for credit scoring and explains them.

problem Lack of interpretability in advanced AI models hinders credit risk management.
method Comparison of logistic regression, AI algorithms, and techniques to interpret AI models.
result Advanced tree-based models provide the best prediction of client default.

LDA-XGB1 balances fairness and accuracy in lending models.

problem Fair lending practices and model interpretability in binary classification.
method Biobjective optimization using binning and information value, leveraging XGBoost.
result Achieves effective balance between accuracy, fairness, and interpretability.

The US Census Bureau corrupts data to protect privacy, but we show how to clean and analyze it effectively.

problem Analyzing Census data with intentional corruption to maintain privacy.
method Formulated a semiparametric model, proposed data cleaning, estimation, and inference procedures.
result Demonstrated that data cleaning can maintain precision and provided theoretical and empirical support.

In this article we present an alternative model for the distribution of household incomes in the United States. We provide arguments from two differing perspectives which both yield the proposed income distribution curve, and then fit this curve to empirical data on household income distribution obtained from the Unite…

2016-02-19abs ↗pdf ↗

2020 Census uses more noise to protect privacy than needed, improving data accuracy.

problem Ensuring privacy in census data while maintaining accuracy for policy decisions.
method Applied ff-differential privacy to track and reduce noise across geographical levels.
result The 2020 Census provides stronger privacy protections than its nominal guarantees suggest.

Large corporate credit models may be adapted for small business risk assessment.

problem Limited data and lack of credit analysts for small businesses.
method Adapting large corporate credit risk models for small businesses.
result Adapted models can predict small business credit risk effectively.

CCR-CNN uses CNN to predict corporate credit ratings from financial data.

problem Lack of data and limited model performance in predicting corporate credit ratings.
method Transform corporations into images and use CNN to analyze complex feature interactions.
result CCR-CNN outperforms state-of-the-art methods in predicting corporate credit ratings.

Framework integrates financial and annual report data for better corporate credit ratings.

problem Lack of insights from non-financial data in credit rating models.
method Uses FinBERT to extract features from annual reports and combines them with financial data.
result Improves credit rating accuracy by 8-12%.

A new algorithm improves credit scoring accuracy for imbalanced data.

problem Poor classification of minority class in credit scoring data sets.
method Weighted-Hybrid-Sampling-Boost (WHSBoost) algorithm with balanced data sampling.
result WHSBoost outperforms other methods in credit scoring accuracy.

This paper develops a machine learning model to assess credit risk in UAE commercial banks.

problem Lack of precision in conventional credit rating tools for accurate credit risk prediction.
method Constructs a credit risk assessment model using Linear Discriminant Analysis.
result Demonstrates improved accuracy in predicting good and bad creditors compared to conventional methods.

BSAC improves credit scoring models by leveraging autoencoders and addressing imbalanced datasets.

problem Imbalanced and heterogeneous credit scoring datasets.
method Bagging Supervised Autoencoder Classifier (BSAC) that uses autoencoders and undersampling.
result BSAC improves classification of loan applicants, demonstrating robustness and effectiveness.

The paper analyzes Lending Club's loan applicants to predict default risk.

problem Predicting default risk in loan applicants of Lending Club.
method Exploratory data analysis and machine learning (Logistic Regression, Random Forest) were used.
result A credit derivative based on Credit Default Swap was designed to hedge default risk.

Study integrates climate and text data to improve credit default prediction.

problem Improving credit risk assessment for mSEs with limited financial histories.
method Multimodal framework using LSTM, GRU, and transformer models.
result Integration of multiple data modalities improves credit default prediction.

Study shows how macroprudential policies affect credit growth in Israel, especially in housing and business sectors.

problem Impact of macroprudential policies on credit growth in Israel.
method Bank-level panel data analysis for Israel, 2004-2019; interaction of monetary and macroprudential policies.
result Accommodative monetary policy interacts with macroprudential policies to increase total credit growth.

Efficiently calculates privacy guarantees for 2020 Census data.

problem Evaluate privacy guarantees for 2020 U.S. Census data releases.
method Sieve-accelerated quadrature method to evaluate tail probabilities of high-dimensional convolutions.
result Achieves 1,824-fold speedup over prior methods while maintaining error tolerances.

Synthetic data improves credit scoring models' performance without compromising borrower privacy.

problem Scarcity of real data for credit scoring models due to privacy concerns.
method Privacy-preserving training with synthetic data.
result Credit scoring models trained with synthetic data show a reduction of 3% in AUC and 6% in KS compared to real data models.

This paper builds a machine learning model to predict credit defaults for unsecured lending.

problem High credit defaults and delinquency rates in unsecured lending due to imbalanced data.
method Employing machine learning techniques, particularly SMOTE for imbalanced data, and evaluating models like LGBM Classifier.
result LGBM Classifier model outperforms other models in predicting credit defaults.

Paper proposes an intelligent credit limit management system using causal inference.

problem Traditional credit limit management strategies are heuristic and not data-driven.
method Conditional independence testing, response model, log transformation, GBDT encoding, non-linear transformation on features, well-designed metric.
result The proposed approach effectively manages credit limits and incorporates diminishing marginal effects.

NetDP predicts loan defaults using network data, addressing cold-start issues.

problem Cold-start problem in default prediction for new users.
method Combines unsupervised and supervised network representations, using parameter-server for scalability.
result Effectiveness in cold-start problem, especially for new users.

Paper proposes a method to evaluate SME credit risk using meta paths.

problem Evaluate credit risk of small and medium-sized enterprises with limited data.
method Exploits the representative power of information networks and meta paths to infer SME financial status.
result Meta path feature effectively identifies SMEs with credit risks.

AI enhances bank credit risk management through deep learning and data analysis.

problem Inaccurate credit decisions and potential risks in bank credit risk management.
method Innovative application of AI technology, including deep learning and big data analysis.
result AI provides more accurate and comprehensive credit decision support, reducing risks and losses.

Study finds public procurement awards, especially NGEU-funded ones, boost new lending.

problem Understanding the impact of public procurement on new lending.
method Panel data local projections model, controlling for various factors.
result Public procurement awards, particularly NGEU-funded ones, significantly increase new lending.

Flexible models predict US Census survey response rates.

problem Predicting survey response rates in the US Census Bureau.
method Nonparametric additive models with structured interactions using ℓ0-based penalization.
result Models lead to predictions comparable to black-box methods but are interpretable.

Deep learning predicts employment changes and industry health.

problem Forecasting short-term employment changes and assessing long-term industry health.
method LSTNet, a multi-scale deep learning model, processes multivariate time series data.
result LSTNet outperforms baseline models in most sectors, especially stable ones.

Credit risk prediction is an effective way of evaluating whether a potential borrower will repay a loan, particularly in peer-to-peer lending where class imbalance problems are prevalent. However, few credit risk prediction models for social lending consider imbalanced data and, further, the best resampling technique t…

2018-04-28abs ↗pdf ↗