Study identifies 1,012 persistent wallet cohorts on Solana pump.fun, showing coordinated buying behavior.
problem Understanding coordinated buying behavior on Solana pump.fun.
method Two-stage detection pipeline: first-buyer-window extraction followed by persistent-cohort surfacing via graph co-occurrence.
result 1,012 persistent wallet cohorts identified, showing systematic co-buying across multiple launches.
Paper introduces OCRR Score for quantifying DeFi wallet credit risk.
problem Inability to assess credit risk in decentralized finance.
method Probabilistic measure based on historical and predictive on-chain activity.
result Dynamic adjustment of LTV and LT based on wallet risk profile.
Digital money could reduce germ spread during coronavirus.
problem Spreading of germs via paper money during coronavirus.
method Policy recommendations for mobile wallets, digital currencies, and data protection.
result Adopting digital money can help reduce germ spread.
The paper assesses VASPs' solvency using multiple data sources.
problem Insolvency risk in VASPs without systematic auditing.
method Cross-referencing cryptoasset wallets, balance sheets, and supervisory data.
result Inconsistent data between DLT transactions and balance sheets for some VASPs.
Cohort effects are important factors in determining the evolution of human mortality for certain countries. Extensions of dynamic mortality models with cohort features have been proposed in the literature to account for these factors under the generalised linear modelling framework. In this paper we approach the proble…
Paper evaluates deadline-ILS on insider trading contracts, finding it distinguishes signals from noise.
problem Deadlines in insider trading contracts and information leakage detection.
method Empirical evaluation using FFIC dataset, hazard-rate estimation, cross-market wallet analysis.
result Deadline-ILS distinguishes signal from proxy artefact, with a significant shift in magnitude.
Paper introduces SCI to distinguish market signals from coordination.
problem Unclear signals in prediction markets.
method Formalizes SCI, introduces weighted and time-varying extensions.
result Discriminates between market signals and coordination.
Deep neural networks improve sleep stage classification across diverse datasets.
problem Manual sleep scoring is subjective and lacks reliability; automatic systems generalize poorly.
method Developed a deep neural network using 15,684 polysomnography studies from five cohorts.
result Classification accuracy improved with more training data and multiple data sources.
AnChain.AI detects NFT wash trading with 0.14% of transactions flagged.
problem NFT market manipulation through wash trading.
method Algorithm flags transactions within 30 days of repurchase.
result 0.14% of NFT transactions are involved in wash trading.
The Lee Carter modelling framework is widely used because of its simplicity and robustness despite its inability to model specific cohort effects. A large number of extensions have been proposed that model cohort effects but there is no consensus. It is difficult to simultaneously account for cohort effects and age-adj…
We introduce a variable importance measure to quantify the impact of individual input variables to a black box function. Our measure is based on the Shapley value from cooperative game theory. Many measures of variable importance operate by changing some predictor values with others held fixed, potentially creating unl…
Research aims to ensure fair classification across explicit and implicit sensitive features.
problem Ensuring fairness in machine learning models when sensitive features are not explicitly provided.
method Defined explicit and implicit cohorts, used clustering of embeddings, modified loss function.
result Improved classification parity across explicit and implicit sensitive features.
COHORTNEY groups web users based on activity patterns.
problem Lack of academic discussion on cohort analysis for user behavior.
method Unsupervised non-parametric machine learning approach.
result COHORTNEY outperforms traditional methods in cohort analysis.
Agent-to-agent finance aims to manage payments and trust for AI agents.
problem Managing financial interactions between autonomous AI agents.
method Develops agent-to-agent finance concept and explores blockchain solutions.
result Agent-to-agent finance can address coordination frictions in financial markets.
ODVICE augments EHR cohorts using ontology to improve analysis robustness.
problem Limited records in cohorts for rare diseases hamper robust analysis.
method Ontology-driven Monte-Carlo graph spanning algorithm for data augmentation.
result ODVICE augmented cohorts show ~30% improvement in AUC over non-augmented datasets.
Proposes a new model for mortality forecasting considering age groups and cohort effects.
problem Longevity risk due to ageing population.
method Mixed-effects time-series approach with age groups dependency and random cohort effects.
result Remarkable improvements in forecast accuracy compared to the CBD model.
CAT framework improves AI medical screening fairness and reliability.
problem Imbalanced data, varying performance across cohorts, and patient-level inconsistencies in traditional metrics.
method CAT framework introduces patient-level assessment, entropy-based distribution weighting, and cohort-weighted sensitivity and specificity.
result Enhanced predictive reliability, fairness, and interpretability of AI-driven medical screening models.
Cohort analysis speeds up Bitcoin blockchain data queries.
problem Efficiently querying Bitcoin blockchain data for economic insights.
method Cohort analysis applied to Bitcoin transaction data.
result Creation of datasets and visualizations for key Bitcoin transaction indicators.
Mean-field approximations simplify insurance liability calculations.
problem High-dimensional system of equations makes insurance liability calculation infeasible.
method Use mean-field model to replace high-dimensional system with a low-dimensional non-linear system.
result Insurance liability converges to mean-field approximation as cohort size increases.
Brain imaging analysis on clinically acquired computed tomography (CT) is essential for the diagnosis, risk prediction of progression, and treatment of the structural phenotypes of traumatic brain injury (TBI). However, in real clinical imaging scenarios, entire body CT images (e.g., neck, abdomen, chest, pelvis) are t…
A new method for variable importance measures without impossible data.
problem Using impossible data for variable importance measures in black box models.
method Cohort Shapley, a method grounded in economic game theory using only observed data.
result Cohort Shapley provides a more trustworthy explanation of black box models' decisions.
Treatment recommendations within Clinical Practice Guidelines (CPGs) are largely based on findings from clinical trials and case studies, referred to here as research studies, that are often based on highly selective clinical populations, referred to here as study cohorts. When medical practitioners apply CPG recommend…
Develops a new method to model overlapping asymmetric datasets effectively.
problem Handling overlapping asymmetric datasets in data science.
method Twice penalized P-Spline approximation method.
result Improves model fit by over 65% in a real-life dataset.
Enhancing spectral embedding for low-dimensional embeddings in rare disease cohorts
problem Representing clinical concepts and patients in electronic health records
method Spectral-based unsupervised learning with flexible knowledge transfer
result Outperforms competing approaches in challenging scenarios
Advances in molecular "omics'" technologies have motivated new methodology for the integration of multiple sources of high-content biomedical data. However, most statistical methods for integrating multiple data matrices only consider data shared vertically (one cohort on multiple platforms) or horizontally (different …
Transformers simplify modeling of small longitudinal cohort data by reducing parameters and incorporating attention mechanisms.
problem Challenges in modeling longitudinal cohort data due to complex temporal dependencies and large dataset requirements.
method Simplified transformer architecture with attention mechanism, autoregressive model, and kernel-based temporal decay.
result The approach recovers contextual dependencies even with small datasets, identifying temporal patterns in stress and mental health.
In systems biomedicine, an experimenter encounters different potential sources of variation in data such as individual samples, multiple experimental conditions, and multi-variable network-level responses. In multiparametric cytometry, which is often used for analyzing patient samples, such issues are critical. While c…
Randomized Controlled Trials (RCTs) are the gold standard for comparing the effectiveness of a new treatment to the current one (the control). Most RCTs allocate the patients to the treatment group and the control group by uniform randomization. We show that this procedure can be highly sub-optimal (in terms of learnin…
Compact formulas for evaluating insurance policies' risks.
problem Quantifying demographic risk in insurance portfolios.
method Cohort-based approach with market-consistent valuation.
result Formal closed formula for idiosyncratic risk (accidental mortality).
Develops a GP framework for age and year-specific mortality surfaces.
problem Learning the covariance structure of age and year-specific mortality surfaces.
method Genetic programming algorithm to search for the most expressive GP kernel.
result Reveals the presence/absence of cohort effects in different populations.
The paper presents a systematic review of state-of-the-art approaches to identify patient cohorts using electronic health records. It gives a comprehensive overview of the most commonly de-tected phenotypes and its underlying data sets. Special attention is given to preprocessing of in-put data and the different modeli…
Deep learning (DL) methods have in recent years yielded impressive results in medical imaging, with the potential to function as clinical aid to radiologists. However, DL models in medical imaging are often trained on public research cohorts with images acquired with a single scanner or with strict protocol harmonizati…
Paper develops NN models for diabetes screening using NHANES data.
problem Developing accurate predictive models for diabetes in diverse populations.
method Proposes a neural network framework with survey weights, uncertainty quantification.
result Robust risk score models for diabetes in US population.
Typical cohorts in brain imaging studies are not large enough for systematic testing of all the information contained in the images. To build testable working hypotheses, investigators thus rely on analysis of previous work, sometimes formalized in a so-called meta-analysis. In brain imaging, this approach underlies th…
The paper compares two methods for handling missing data in causal discovery.
problem Handling missing data in causal discovery algorithms.
method Test-wise deletion and multiple imputation.
result Multiple imputation is more challenging for causal discovery than for estimation.
There is growing interest in the design of pension annuities that insure against idiosyncratic longevity risk while pooling and sharing systematic risk. This is partially motivated by the desire to reduce capital and reserve requirements while retaining the value of mortality credits; see for example Piggott, Valdez an…
Causal analysis reveals regional discrepancies in TOPCAT trial results.
problem Inconclusive results in TOPCAT trial for heart failure treatment.
method Causal discovery methods with domain knowledge integration.
result Significant causal effects shown for some subgroups globally.
At this moment, databanks worldwide contain brain images of previously unimaginable numbers. Combined with developments in data science, these massive data provide the potential to better understand the genetic underpinnings of brain diseases. However, different datasets, which are stored at different institutions, can…
Paper uses RNN to predict SaaS user lifetime value.
problem Predicting user lifetime value in SaaS applications.
method Recurrent Neural Network with multi-cell architecture, accounting for cohort, age-in-system, and contemporaneous information.
result Significantly improved prediction accuracy compared to existing models.
We explore inverse and quanto inverse crypto options, their pricing, and applications.
problem Market incompleteness in crypto options trading.
method Comparison of direct and inverse options, and introduction of currency-protected 'quanto' options.
result Pricing and hedging characteristics of inverse and quanto inverse options in a Black-Scholes framework.
Motivation: How do we integratively analyze large-scale multi-platform genomic data that are high dimensional and sparse? Furthermore, how can we incorporate prior knowledge, such as the association between genes, in the analysis systematically? Method: To solve this problem, we propose a Scalable Network Constrained T…
A new model-free variable importance method (IGCS) is introduced for high-dimensional data.
problem Model-free variable importance for high-dimensional data, especially when prediction functions are proprietary or expensive.
method Integrated Gradient (IG) version of Cohort Shapley (CS) method with O(nd) cost. result IGCS closely matches Cohort Shapley (CS) in performance, especially for binary predictors.
Background: Despite recent significant progress in the development of automatic sleep staging methods, building a good model still remains a big challenge for sleep studies with a small cohort due to the data-variability and data-inefficiency issues. This work presents a deep transfer learning approach to overcome thes…
Procedure removes training data dependency from deep networks, improving generalization.
problem Removing dependency on training data in deep networks for better generalization.
method Deterministic and stochastic parts to ensure forgetting, leveraging activation and weight dynamics.
result New bound on information extraction from black-box networks, ensuring forgetting in activations.
A new method assesses algorithmic fairness using game theory.
problem Evaluating algorithmic fairness without proprietary data.
method Cohort Shapley value, a game-theoretic approach.
result Identifies individual impact of protected attributes.
Inspection-L detects illicit cryptocurrency transactions using GNNs and self-supervised learning.
problem Detect illicit cryptocurrency transactions for anti-money laundering.
method Graph Neural Network (GNN) framework based on self-supervised Deep Graph Infomax (DGI) and Graph Isomorphism Network (GIN) with supervised learning algorithms.
result Inspection-L outperforms state-of-the-art methods in key classification metrics.
New algorithm improves causal discovery in biomedical data.
problem Stability and accuracy issues in causal discovery algorithms.
method Exploits temporal structure and tiered background knowledge.
result Increases accuracy in finite samples for causal structure estimation.
Study reveals structure of Bitcoin's crypto flow network.
problem Understanding crypto flows among Bitcoin users.
method Blockchain data, user identification, network construction, bow-tie structure, Hodge decomposition, non-negative matrix factorization.
result Users are located in upstream, downstream, and core of the crypto flow network.