Deviance Voronoi residuals improve earthquake insurance risk assessment.
problem Assessing earthquake insurance risk using spatio-temporal point process models.
method Extended Voronoi residuals and created simulation-based approach.
result Proposed formula for country-wide minimum capital test.
Deviance-style normalization for sparse, jointly overdispersed count matrices
problem Jointly overdispersed count matrices
method Dirichlet-multinomial deviance residualization
result Preserves exact sparsity, evaluates in constant time, recovers multinomial residual
Extends matrix factorization for deviance-based losses with GLM theory.
problem Improving data loss models beyond squared error.
method Adapts GLM theory to matrix factorization for deviance losses.
result Strong consistency and robustness of the proposed decomposition.
This paper aims to review the methodology behind the generalized linear models which are used in analyzing the actuarial situations instead of the ordinary multiple linear regression. We introduce how to assess the adequacy of the model which includes comparing nested models using the deviance and the scaled deviance. …
The paper addresses insurance pricing by improving machine learning models and metrics.
problem Lack of balance and confusion in insurance model performance metrics.
method Introduces autocalibration and Tweedie deviance minimization for insurance pricing models.
result Autocalibration corrects bias and ensures balance on local scales.
Discovering the causal structure among a set of variables is a fundamental problem in many areas of science. In this paper, we propose Kernel Conditional Deviance for Causal Inference (KCDC) a fully nonparametric causal discovery method based on purely observational data. From a novel interpretation of the notion of as…
Paper introduces NICc for fast cluster-based validation of prediction models.
problem Validation of prediction models on clustered data.
method Derived NICc to approximate leave-one-cluster-out deviance for standard regression models.
result NICc provides more accurate model size and variable selection, especially with strong clustering.
This paper generalizes beta divergence beyond its classical form associated with power variance functions of Tweedie models. Generalized form is represented by a compact definite integral as a function of variance function of the exponential dispersion model. This compact integral form simplifies derivations of many pr…
Bayesian model clusters brain activity time series.
problem Heterogeneous multivariate time series in brain imaging.
method Group-based Bayesian mixture of smoothing splines with covariate effects.
result Distinct brain activity patterns identified.
New method uses kernel deviance measures to discover causal relationships in heterogeneous data.
problem Discovering causal relationships in complex, heterogeneous datasets.
method KIIM-HT, a novel score measure based on heterogeneous transformations of RKHS embeddings.
result KIIM-HT outperforms previous methods in causal discovery tasks.
EGO-MDA identifies optimal spectral-bands for process discrimination.
problem Optimal spectral-bands for process discrimination.
method EGO-MDA, an unsupervised method using EGO and Mixture Discriminant Analysis.
result EGO-MDA achieves at least 70% improvement in median deviance.
Federated learning calibrates insurance indices from renewable energy producers' data.
problem Calibrating parametric insurance indices under heterogeneous renewable energy production losses.
method Federated learning framework using Tweedie GLMs and distributed optimization.
result Federated learning recovers comparable index coefficients under moderate heterogeneity.
Study optimizes CANN for actuarial tasks using RSM.
problem Optimizing hyperparameters for neural networks in actuarial science.
method Factorial design and response surface methodology (RSM).
result Reduced hyperparameter optimization from 288 to 188, achieving near-optimal performance.
Framework monitors insurance pricing models for drift and recalibration.
problem Maintaining predictive performance of pricing models in evolving insurance portfolios.
method Formalizes deviance loss and Murphy's score, studies Gini score, develops monitoring framework.
result Framework guides decisions on refitting or recalibrating pricing models.
The discovery of causal relationships is a fundamental problem in science and medicine. In recent years, many elegant approaches to discovering causal relationships between two variables from observational data have been proposed. However, most of these deal only with purely directed causal relationships and cannot det…
Introduces Soft-SVM for binary classification bridging logistic and SVM.
problem Data separability issues in binary classification.
method Soft-SVM regression using convex relaxation of hinge loss with softness and class-separation parameters.
result Soft-SVM performs well in classification and prediction errors.
Develops a GMM method to estimate roughness in stochastic volatility models.
problem Estimating roughness in stochastic volatility models with fractional Brownian motion.
method GMM approach for log-normal models with integrated variance and noisy realized variance.
result Consistent and asymptotically normal parameter estimator with bias correction.
Dynamic model captures spatial, temporal, and spatiotemporal volatility effects.
problem Analyzing volatility in spatial and temporal networks.
method Dynamic spatiotemporal and network ARCH model with common factors, Bayesian estimation.
result Model captures strong spatial/network interactions and spillover effects.
Bayesian CART models improve insurance claims frequency prediction and interpretation.
problem Improving accuracy and interpretability in insurance pricing models.
method Introducing Bayesian CART models for claims frequency, implementing MCMC algorithm for posterior tree exploration, and using DIC for model selection.
result Bayesian CART models can better classify policy-holders into risk groups.
Reasoning based on causality, instead of association has been considered as a key ingredient towards real machine intelligence. However, it is a challenging task to infer causal relationship/structure among variables. In recent years, an Independent Mechanism (IM) principle was proposed, stating that the mechanism gene…
The Partial Information Decomposition (PID) [arXiv:1004.2515] provides a theoretical framework to characterize and quantify the structure of multivariate information sharing. A new method (Idep) has recently been proposed for computing a two-predictor PID over discrete spaces. [arXiv:1709.06653] A lattice of maximum en…
New meta-score EPP interprets model performance differences.
problem Lack of interpretable benchmarks for model performance.
method Elo-based Predictive Power (EPP) meta-score, logistic regression.
result EPP scores have probabilistic interpretation and can be compared between data sets.
Mixture model-based clustering has become an increasingly popular data analysis technique since its introduction over fifty years ago, and is now commonly utilized within a family setting. Families of mixture models arise when the component parameters, usually the component covariance (or scale) matrices, are decompose…
Bregman perspective on CART provides a unified framework for impurity measures.
problem Unifying impurity measures in CART
method Bregman divergence approach
result Unified framework for impurity measures
A novel spatio-temporal graph neural network with a learnable Tweedie head improves vessel traffic flow prediction in sparse maritime data.
problem Accurate vessel traffic flow prediction in sparse maritime data.
method A model-agnostic learnable Tweedie head attached to ST-GNN backbones.
result The proposed head consistently improves RMSE across multiple ST-GNN backbones, especially on non-zero events.
Two new algorithms reduce feature space while preserving non-linear relationships.
problem High-dimensional data and overfitting issues.
method Bias-variance analysis for non-linear transformations and generalized linear models.
result Competitive performance on regression and classification tasks.
Study compares machine learning models for insurance pricing, including neural networks and GLMs.
problem Improving insurance pricing models using machine learning techniques.
method Benchmark study using four insurance datasets, comparing GLMs, GBM, FFNN, and CANN.
result CANNs provide better performance than GLMs and GBM, especially for frequency and severity modeling.
Improved fMRI analysis models enhance classification performance and select relevant brain regions.
problem Inaccurate selection of relevant brain components in MVPA models.
method Hybrid Sparsity-Ranked LASSO (JSRL) method integrating component-level and voxel-level activity.
result JSRL models achieve up to 51.7% improvement in cross-validated deviance R2 and 7.3% improvement in cross-validated AUC.