Proposes a flexible neural model for multi-state survival analysis.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A new framework models multi-state events and biomarkers.
In this paper we propose a multi-state model for the evaluation of the conversion option contract. The multi-state model is based on age-indexed semi-Markov chains that are able to reproduce many important aspects that influence the valuation of the option such as the duration problem, the time non-homogeneity and the …
The paper analyzes multivariate payments in multi-state life insurance using Markovian state processes.
Unified framework for fair pricing in long-term insurance products.
New models for insurance claims accounting for delays.
The paper calculates bonus values in complex insurance schemes.
Proposes a new model to analyze mortgage delinquency transitions.
A deep learning framework for survival analysis combining piecewise exponential models.
The idea of forward rates stems from interest rate theory. It has natural connotations to transition rates in multi-state models. The generalization from the forward mortality rate in a survival model to multi-state models is non-trivial and several definitions have been proposed. We establish a theoretical framework f…
In this work, we consider the class of multi-state autoregressive processes that can be used to model non-stationary time-series of interest. In order to capture different autoregressive (AR) states underlying an observed time series, it is crucial to select the appropriate number of states. We propose a new model sele…
Improves clustering fairness by learning fair clusters adaptively.
A Longitudinal Attribute-Conditioned Neural Network (LANTERN) framework for modeling health-state transition probabilities in irregular longitudinal data.
Study on stock trading model with uncertain market status, proving free boundaries and optimal strategies.
Fair clustering under the disparate impact doctrine requires that population of each protected group should be approximately equal in every cluster. Previous work investigated a difficult-to-scale pre-processing step for -center and -median style algorithms for the special case of this problem when the number of …
We study the problem of determining risk-minimizing investment strategies for insurance payment processes in the presence of taxes and expenses. We consider the situation where taxes and expenses are paid continuously and symmetrically and introduce the concept of tax- and expense-modified risk-minimization. Risk-minim…
SiBBlInGS discovers interpretable building blocks across states in multi-way data.
Selection of appropriate collective variables for enhancing sampling of molecular simulations remains an unsolved problem in computational biophysics. In particular, picking initial collective variables (CVs) is particularly challenging in higher dimensions. Which atomic coordinates or transforms there of from a list o…
Complex contagion model explains financial fire sales through continuous asset prices.
The paper models stochastic interest rates for life insurance using phase-type distributions.
The paper reveals that baselines significantly impact RL algorithms' convergence.
The study uses Bayesian Hidden Markov Models to predict cryptocurrency returns.
New algorithm predicts lung cancer progression and mortality.
Evidence suggests Rapid-Eye-Movement (REM) Sleep Behaviour Disorder (RBD) is an early predictor of Parkinson's disease. This study proposes a fully-automated framework for RBD detection consisting of automated sleep staging followed by RBD identification. Analysis was assessed using a limited polysomnography montage fr…
Excited-state dynamics simulations are a powerful tool to investigate photo-induced reactions of molecules and materials and provide complementary information to experiments. Since the applicability of these simulation techniques is limited by the costs of the underlying electronic structure calculations, we develop an…
Big data sets must be carefully partitioned into statistically similar data subsets that can be used as representative samples for big data analysis tasks. In this paper, we propose the random sample partition (RSP) data model to represent a big data set as a set of non-overlapping data subsets, called RSP data blocks,…
Prevents sensitive data generation in diffusion models using labeled and unlabeled data.
Study reveals Data Shapley's inconsistent performance in data selection tasks.
PRRO generates synthetic tabular data that improves SL performance and class distribution.
Defines data science as a natural ecosystem with challenges and missions.
Differences in data size per class, also known as imbalanced data distribution, have become a common problem affecting data quality. Big Data scenarios pose a new challenge to traditional imbalanced classification algorithms, since they are not prepared to work with such amount of data. Split data strategies and lack o…
Synthetic data enhances analytics but requires careful volume management.
New test ensures quality of shared data in machine learning.
Differential privacy allows quantifying privacy loss resulting from accessing sensitive personal data. Repeated accesses to underlying data incur increasing loss. Releasing data as privacy-preserving synthetic data would avoid this limitation, but would leave open the problem of designing what kind of synthetic data. W…
Paper creates fair synthetic data ensuring equal predictions across sensitive attributes.
DPA preserves data distribution in reduced dimensions.
Efficient synthetic data generation improves model performance on tabular data.
For most problems in science and engineering we can obtain data sets that describe the observed system from various perspectives and record the behavior of its individual components. Heterogeneous data sets can be collectively mined by data fusion. Fusion can focus on a specific target relation and exploit directly ass…
DAERNN models censored data using neural networks with data augmentation.
Data preprocessing techniques are devoted to correct or alleviate errors in data. Discretization and feature selection are two of the most extended data preprocessing techniques. Although we can find many proposals for static Big Data preprocessing, there is little research devoted to the continuous Big Data problem. A…
Data stream classification methods demonstrate promising performance on a single data stream by exploring the cohesion in the data stream. However, multiple data streams that involve several correlated data streams are common in many practical scenarios, which can be viewed as multi-task data streams. Instead of handli…
Data collection is a major bottleneck in machine learning and an active research topic in multiple communities. There are largely two reasons data collection has recently become a critical issue. First, as machine learning is becoming more widely-used, we are seeing new applications that do not necessarily have enough …
This paper quantifies uncertainty in Data Shapley using statistical inference.
DCoM uses deep neural networks to detect semantic data types from raw column values.
Causal inference from observational data is the goal of many data analyses in the health and social sciences. However, academic statistics has often frowned upon data analyses with a causal objective. The introduction of the term "data science" provides a historic opportunity to redefine data analysis in such a way tha…
Task-agnostic data valuation without validation requirements.
Data mining is about obtaining new knowledge from existing datasets. However, the data in the existing datasets can be scattered, noisy, and even incomplete. Although lots of effort is spent on developing or fine-tuning data mining models to make them more robust to the noise of the input data, their qualities still st…
New algorithm improves data imputation for complex multimodal data sets.