Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

2835678501,133 · Jun 202019922001200920172026
48 results for claims data

Deep Claim predicts payer responses from claims data using deep learning.

problem Predicting payer responses from claims data to improve healthcare performance.
method Learning complex dependencies in claim inputs to create a compact representation, then using deep learning to predict responses.
result Deep Claim improves claim denial prediction by 22.21%.

The paper introduces BCART models for aggregate claim amount, improving frequency-severity and joint modeling.

problem Modeling aggregate claim amount with frequency-severity and joint dependencies.
method Developed three types of BCART models: frequency-severity, sequential, and joint models. Used various distributions for claim severity data.
result Weibull distribution outperforms gamma and lognormal for right-skewed, heavy-tailed claim severity data.

Study tackles imbalanced data in car insurance claims prediction.

problem Predicting rare events (claims) in car insurance with imbalanced data.
method Various machine learning techniques (logistic-regression, decision tree, random forest, xgBoost, feed-forward network) applied to imbalanced dataset.
result Comparison of machine learning algorithms' performance in claim occurrence prediction.

Traditional non-life reserving models largely neglect the vast amount of information collected over the lifetime of a claim. This information includes covariates describing the policy, claim cause as well as the detailed history collected during a claim's development over time. We present the hierarchical reserving mod…

2019-10-28abs ↗pdf ↗

A new method for modeling insurance claim frequencies using random proportions.

problem Inaccurate fitting of classical distributions to insurance claim frequency data.
method Modeling claim frequencies using random proportions of insurance contracts and applying goodness-of-fit tests.
result A new statistical approach for better modeling insurance claim frequencies.

Bayesian CART models improve insurance claims frequency prediction and interpretation.

problem Improving accuracy and interpretability in insurance pricing models.
method Introducing Bayesian CART models for claims frequency, implementing MCMC algorithm for posterior tree exploration, and using DIC for model selection.
result Bayesian CART models can better classify policy-holders into risk groups.

Machine learning detects NASH patients from medical claims data.

problem Detecting undiagnosed NASH patients for screening and management.
method Gradient-boosted decision trees trained on administrative medical claims data.
result Model precision for NASH detection is significantly higher than NASH incidence.

We approximate the distribution of total expenditure of a retail company over warranty claims incurred in a fixed period [0, T], say the following quarter. We consider two kinds of warranty policies, namely, the non-renewing free replacement warranty policy and the non-renewing pro-rata warranty policy. Our approximati…

2010-08-05abs ↗pdf ↗

Develops a new GLM framework for claims reserving with adaptive estimation.

problem Accurate assessment of claims reserves with dynamic and dependent claim activity.
method Multivariate evolutionary GLM framework with adaptive particle filtering algorithm.
result Adaptive estimation of evolving factors improves claims reserve accuracy.

Reinforcement learning improves insurance claims reserving by learning from all claim trajectories.

problem Traditional reserving models learn only from settled claims, missing valuable data from ongoing claims.
method Formulated as a Markov decision process, uses reinforcement learning to update OCL estimates sequentially.
result Soft Actor-Critic implementation achieves competitive claim-level accuracy and strong aggregate performance.

Study uses healthcare claims data to identify Covid-19 risk factors without prior selection.

problem Identify risk factors for severe Covid-19 cases.
method Fine-grained hierarchical information from medical classification systems used to analyze over 33,000 covariates.
result Method has better predictive ability than pre-specified morbidity groups.

The study analyzes how bonus-malus systems and delayed claims settlement affect insurance companies' financial stability.

problem Analyzing the impact of bonus-malus systems and delayed claims settlement on insurance companies' financial stability.
method Examined a discrete-time risk model with time-varying premiums, evaluating two types of claims and settlement delays.
result Delayed settlement of by-claims leads to lower ruin probabilities under specific assumptions.

This paper suggests claim history will be deprecated in future auto insurance rates.

problem The role of historical claim records in auto insurance rates.
method Proposes a new risk variable elimination method and real-time road risk model design.
result Claim history will be considered a 'noise' factor and deprecated in Pay-How-You-Drive models.

The tail of the distribution of a sum of a random number of independent and identically distributed nonnegative random variables depends on the tails of the number of terms and of the terms themselves. This situation is of interest in the collective risk model, where the total claim size in a portfolio is the sum of a …

2007-03-01abs ↗pdf ↗

CANN models improve insurance claim count predictions using telematics data.

problem Improving insurance claim count predictions with telematics data.
method Combining classical actuarial models with neural networks for telematics data.
result CANN models outperform traditional models in predicting insurance claims.

Two machine learning models detect anomalies in ER claims, saving up to 40% in improper payments.

problem Improper health insurance payments from fraud and upcoding.
method Two machine learning models: an upcoding model based on severity code distributions and a random forest model for claim sorting.
result Random forest model saved 12% to 40% in improper payments compared to a baseline approach.

FL improves insurance claims loss prediction without sharing data.

problem Limited data volume and variety due to privacy concerns.
method Federated Learning (FL) to update a global model using local data insights.
result Improved claims loss forecasting compared to individual models.

This study compares the largest claims from two insurance portfolios using stochastic orderings.

problem Comparing the largest claims from two heterogeneous insurance portfolios.
method Used various stochastic orderings and established sufficient conditions associated with model parameters.
result Established sufficient conditions for comparing the largest claims from two insurance portfolios.

Model detects insurance fraud using social network analysis.

problem Fraudulent insurance claims by exaggeration or intentional damage.
method Network construction linking claims and parties, BiRank algorithm for fraud score computation, feature extraction from network and claims, supervised model building.
result Network features improve fraud detection performance.

The paper proposes an original methodology for constructing quantitative statistical models based on multidimensional distribution functions constructed on the basis of the insurance companies' data on inshurance policies (including policies with deductible) and claims incurred. Real data of some Russian insurance comp…

2019-08-14abs ↗pdf ↗

Model predicts individual insurance claim reserves using activation patterns.

problem Accurately predicting individual claim reserves in insurance contracts.
method Multinomial logistic regression to model claim activation and development.
result The model generates accurate predictions of total and per coverage reserves.

The article proposes a method to make valid insurance claim predictions without relying on specific models.

problem Prediction of insurance claims using statistical models can be unreliable due to model misspecification, selection effects, and lack of finite-sample validity.
method The article employs conformal prediction, a machine learning strategy that is model-free and tuning-parameter-free, ensuring finite-sample validity.
result The proposed method guarantees valid predictions at a pre-assigned coverage probability level and performs well in insurance applications, including meeting Solvency II requirements.

Study improves flood loss risk models using historical data and rainfall data.

problem Predicting financial losses from flooding events.
method Used neural networks, decision trees, and kernel-based regressors on NFIP dataset, incorporating rainfall data.
result Extreme Gradient Boosting provided the best results, and bias correction improved model performance.

In this work, we focus on fine-tuning an OpenAI GPT-2 pre-trained model for generating patent claims. GPT-2 has demonstrated impressive efficacy of pre-trained language models on various tasks, particularly coherent text generation. Patent claim language itself has rarely been explored in the past and poses a unique ch…

2019-07-01abs ↗pdf ↗

Enhances insurance loss models using InsurTech data and machine learning.

problem Traditional insurance loss models lack predictive accuracy due to limited data sources.
method Combining proprietary claims data with InsurTech data and applying machine learning techniques.
result Improved predictive accuracy of the loss model through machine learning.

Study improves motor insurance claim prediction using geographic data.

problem Limited location identifiers in public actuarial datasets.
method Zone-level modeling framework with environmental and orthoimagery data.
result Geographic information improves MTPL claim prediction accuracy.

New methods for quantifying insurance claim cost uncertainty using LightGBM and GLMs.

problem Quantifying prediction uncertainty in insurance claim costs.
method Proposed non-conformity measures for GLMs and GBMs with Tweedie loss.
result Locally weighted Pearson residuals outperform other methods in maintaining nominal coverage with smallest average width.