A new method for modeling insurance claim frequencies using random proportions.
problem Inaccurate fitting of classical distributions to insurance claim frequency data.
method Modeling claim frequencies using random proportions of insurance contracts and applying goodness-of-fit tests.
result A new statistical approach for better modeling insurance claim frequencies.
The paper introduces BCART models for aggregate claim amount, improving frequency-severity and joint modeling.
problem Modeling aggregate claim amount with frequency-severity and joint dependencies.
method Developed three types of BCART models: frequency-severity, sequential, and joint models. Used various distributions for claim severity data.
result Weibull distribution outperforms gamma and lognormal for right-skewed, heavy-tailed claim severity data.
Bayesian CART models improve insurance claims frequency prediction and interpretation.
problem Improving accuracy and interpretability in insurance pricing models.
method Introducing Bayesian CART models for claims frequency, implementing MCMC algorithm for posterior tree exploration, and using DIC for model selection.
result Bayesian CART models can better classify policy-holders into risk groups.
A new method for predicting insurance claims with statistical guarantees.
problem Creating accurate prediction intervals for insurance claims.
method Model-agnostic framework using split conformal prediction for frequency-severity modeling.
result Shows effectiveness on simulated and real datasets using various models.
It is illustrated a methodology to compute the pure premium for the automobile insurance (claim frequency and severity) using generalized linear models. It is obtained the pure premium for the partial damage loss cover (PPD) using a set of automobile insurance policies with an exposition of a year. It is found that the…
EBM improves car insurance claim severity and frequency prediction while maintaining interpretability.
problem Balancing predictive accuracy and interpretability in insurance claim modeling.
method Combines GAM and cyclic gradient boosting, providing interpretable predictions.
result EBM outperforms benchmark models in claim severity and frequency prediction.
Unified comparison of gradient boosting algorithms for insurance claims.
problem Improving predictive accuracy and computational efficiency in insurance claim prediction.
method Unified notation and comprehensive numerical study comparing 12 gradient boosting algorithms on 5 datasets.
result No trade-off between model adequacy and predictive accuracy.
Study improves motor insurance claim prediction using geographic data.
problem Limited location identifiers in public actuarial datasets.
method Zone-level modeling framework with environmental and orthoimagery data.
result Geographic information improves MTPL claim prediction accuracy.
The paper proposes an original methodology for constructing quantitative statistical models based on multidimensional distribution functions constructed on the basis of the insurance companies' data on inshurance policies (including policies with deductible) and claims incurred. Real data of some Russian insurance comp…
New model bridges pricing and reserving for insurance claims.
problem Incomplete claim data due to reporting and settlement delays.
method Develops an occurrence and development model to estimate both claims and premiums.
result Effective resolution of pricing and reserving inconsistencies.
Improved risk assessment for UBI using telematics data and AdaBoost.
problem Class imbalance in predicting claims frequency for UBI.
method Cost-sensitive multi-class AdaBoost (SAMME.C2) algorithm.
result SAMME.C2 outperforms other models in handling class imbalances.
In clinical and neuroscientific studies, systematic differences between two populations of brain networks are investigated in order to characterize mental diseases or processes. Those networks are usually represented as graphs built from neuroimaging data and studied by means of graph analysis methods. The typical mach…
New harmonic functions show nodal sets can be topologically complex despite frequency and regularity constraints.
problem Understanding the topology of nodal sets of harmonic functions with bounded frequency and regularity.
method Constructing harmonic functions on the unit ball with specific properties.
result The Betti numbers of the nodal set can be arbitrarily large, contradicting previous topological bounds.
Enhances non-life insurance pricing models using transformer models.
problem Improving predictive power of non-life insurance pricing models.
method Enhances actuarial non-life models with transformer models for tabular data.
result Transformer models outperform benchmark models in claim frequency prediction.
We present a novel approach to describing the microstructure of high frequency trading using two key elements. First we introduce a new notion of informed trader which we starkly contrast to current informed trader models. We describe the exact nature of the `superior information' high frequency traders have access to,…
Deep neural networks can generalize by reducing high-frequency noise over time, not always following a monotonic learning bias.
problem Understanding the learning dynamics and generalization of over-parameterized DNNs.
method Experimental analysis of deep double descent, focusing on the spectral bias of DNNs.
result The high-frequency components of DNNs diminish over training, leading to a second descent in test error.
Paper proposes a risk index combining frequency and severity of abnormal driving patterns.
problem Assessing driver risk based on telematics data.
method Combines frequency of abnormal driving patterns with severity quantified through tail rarity.
result Developed a risk index that enables reliable discrimination and ranking of drivers.
Study compares machine learning models for insurance pricing, including neural networks and GLMs.
problem Improving insurance pricing models using machine learning techniques.
method Benchmark study using four insurance datasets, comparing GLMs, GBM, FFNN, and CANN.
result CANNs provide better performance than GLMs and GBM, especially for frequency and severity modeling.
ResNets can approximate input distances under certain conditions, but existing theory is flawed.
problem Theoretical justification for regularizing ResNets to preserve input distances is flawed.
method Frequency analysis perspective to explain effectiveness of regularization schemes.
result Regularization schemes enforce a lower Lipschitz bound on low-frequency projections of images.
Study optimal reinsurance pricing under model uncertainty for multiple insurers.
problem Optimal reinsurance pricing in the presence of multiple sources of model uncertainty.
method Solves a continuous-time Stackelberg game for general reinsurance contracts, considering entropy penalties and ambiguity in insurers' models.
result Reinsurer prices under a distortion of the barycentre of insurers' models, maximizing expected wealth with an entropy penalty.
Estimating a large alphabet probability distribution from a limited number of samples is a fundamental problem in machine learning and statistics. A variety of estimation schemes have been proposed over the years, mostly inspired by the early work of Laplace and the seminal contribution of Good and Turing. One of the b…
Method proposed for pricing insurance products covering both foreseeable and unforeseeable risks.
problem Pricing insurance products that include unforeseeable risks.
method Mixed Poisson process with Bayesian setup and linear exponential family distributions.
result Bayesian premiums are more reactive to claim trends than traditional ones.
Many studies have shown that there are good reasons to claim very low predictability of currency nevertheless, the deviations from true randomness exist which have potential predictive and prognostic power [J.James, Quantitative finance 3 (2003) C75-C77]. We analyze the local trends which are of the main focus of the t…
This paper generalizes the framework for arbitrage-free valuation of bilateral counterparty risk to the case where collateral is included, with possible re-hypotecation. We analyze how the payout of claims is modified when collateral margining is included in agreement with current ISDA documentation. We then specialize…
Reinforcement learning improves insurance claims reserving by learning from all claim trajectories.
problem Traditional reserving models learn only from settled claims, missing valuable data from ongoing claims.
method Formulated as a Markov decision process, uses reinforcement learning to update OCL estimates sequentially.
result Soft Actor-Critic implementation achieves competitive claim-level accuracy and strong aggregate performance.
The study analyzes how bonus-malus systems and delayed claims settlement affect insurance companies' financial stability.
problem Analyzing the impact of bonus-malus systems and delayed claims settlement on insurance companies' financial stability.
method Examined a discrete-time risk model with time-varying premiums, evaluating two types of claims and settlement delays.
result Delayed settlement of by-claims leads to lower ruin probabilities under specific assumptions.
Deep Claim predicts payer responses from claims data using deep learning.
problem Predicting payer responses from claims data to improve healthcare performance.
method Learning complex dependencies in claim inputs to create a compact representation, then using deep learning to predict responses.
result Deep Claim improves claim denial prediction by 22.21%.
New method for individual claims reserving using machine learning.
problem Traditional claims reserving methods are limited in individual claim prediction.
method Restructured data utilization for CL prediction, using multi-period factors.
result Neural networks applied for individual claims reserving.
The tail of the distribution of a sum of a random number of independent and identically distributed nonnegative random variables depends on the tails of the number of terms and of the terms themselves. This situation is of interest in the collective risk model, where the total claim size in a portfolio is the sum of a …
Two machine learning models detect anomalies in ER claims, saving up to 40% in improper payments.
problem Improper health insurance payments from fraud and upcoding.
method Two machine learning models: an upcoding model based on severity code distributions and a random forest model for claim sorting.
result Random forest model saved 12% to 40% in improper payments compared to a baseline approach.
Optimizes insurance processing capacity to minimize costs.
problem Processing delays and backlogs in insurance claims.
method Optimal capacity selection to minimize delay-adjusted and fixed costs.
result Minimizes claims costs by balancing processing capacity and delays.
A non-parametric method for evaluation of the aggregate loss distribution (ALD) by combining and numerically inverting the empirical characteristic functions (CFs) is presented and illustrated. This approach to evaluate ALD is based on purely non-parametric considerations, i.e., based on the empirical CFs of frequency …
This study compares the largest claims from two insurance portfolios using stochastic orderings.
problem Comparing the largest claims from two heterogeneous insurance portfolios.
method Used various stochastic orderings and established sufficient conditions associated with model parameters.
result Established sufficient conditions for comparing the largest claims from two insurance portfolios.
We present a careful analysis of possible issues on the application of the self-excited Hawkes process to high-frequency financial data. We carefully analyze a set of effects leading to significant biases in the estimation of the "criticality index" n that quantifies the degree of endogeneity of how much past events tr…
Model detects insurance fraud using social network analysis.
problem Fraudulent insurance claims by exaggeration or intentional damage.
method Network construction linking claims and parties, BiRank algorithm for fraud score computation, feature extraction from network and claims, supervised model building.
result Network features improve fraud detection performance.
The paper models SaaS products as insurance, offering new pricing tools.
problem Modeling capped-usage SaaS products with insurance principles.
method Frequency-severity decomposition, premium calculation, Monte Carlo simulations.
result SaaS pricing can be analyzed using insurance actuarial methods.
We consider trading in a financial market with proportional transaction costs. In the frictionless case, claims are maximal if and only if they are priced by a consistent price process--the equivalent of an equivalent martingale measure. This result fails in the presence of transaction costs. A properly maximal claim i…
Investor maximizes utility from an unknown claim using robust optimization.
problem Maximizing utility from an unknown contingent claim.
method Robust optimization with quantile formulation and variational inequalities.
result Optimal trading strategy and utility indifference price determined.
Model predicts individual insurance claim reserves using activation patterns.
problem Accurately predicting individual claim reserves in insurance contracts.
method Multinomial logistic regression to model claim activation and development.
result The model generates accurate predictions of total and per coverage reserves.
BERT learns claim descriptions to identify patent novelty.
problem Identifying novel patent claims among existing documents.
method Training BERT on concatenated claims and descriptions, scoring BERT's output.
result BERT identifies relevant X documents for patent novelty.
Insurance companies must manage millions of claims per year. While most of these claims are non-fraudulent, fraud detection is core for insurance companies. The ultimate goal is a predictive model to single out the fraudulent claims and pay out the non-fraudulent ones immediately. Modern machine learning methods are we…
Paper introduces EEMs for pricing contingent claim returns.
problem Computing expected future prices of contingent claims.
method Dynamic change of measure approach to construct EEMs.
result EEMs provide physical and pricing expectations of contingent claim prices.
Traditional non-life reserving models largely neglect the vast amount of information collected over the lifetime of a claim. This information includes covariates describing the policy, claim cause as well as the detailed history collected during a claim's development over time. We present the hierarchical reserving mod…
In this work, we focus on fine-tuning an OpenAI GPT-2 pre-trained model for generating patent claims. GPT-2 has demonstrated impressive efficacy of pre-trained language models on various tasks, particularly coherent text generation. Patent claim language itself has rarely been explored in the past and poses a unique ch…
FiNCAT tool automatically identifies financial numerals in documents.
problem Differentiating between in-claim and out-of-claim numerals in financial documents.
method Extracts context embeddings of numerals using BERT, then uses Logistic Regression to classify.
result Achieved a Macro F1 score of 0.8223 on validation set.
LLMs help automate extraction of actuarial variables from unstructured claims data.
problem Manual processing of unstructured claims data is time-consuming and inconsistent.
method Two-stage processing architecture using LLMs, modular Python pipeline.
result LLM-based extraction achieved high accuracy and practical actuarial value.
New method simplifies individual claims reserving.
problem Insufficient flexibility and robustness in existing methods.
method Building on classical chain-ladder method, introduces new perspective.
result Advances toward a new standard for micro-level reserving.
Using a suitable change of probability measure, we obtain a novel Poisson series representation for the arbitrage- free price process of vulnerable contingent claims in a regime-switching market driven by an underlying continuous- time Markov process. As a result of this representation, along with a short-time asymptot…