ZICO learns DAGs from zero-inflated count data efficiently.
problem Learning network structures from zero-inflated count data.
method ZICO uses node-wise likelihoods with canonical links and a differentiable surrogate constraint for acyclicity.
result ZICO achieves superior performance and faster runtimes on simulated data.
New ZIPLN model accounts for zero-inflation in multivariate count data.
problem Zero-inflation in multivariate count data.
method Introduced Zero-Inflated PLN (ZIPLN) model with variational inference.
result ZIPLN significantly improves log-likelihood and reduces dispersion.
Paper introduces ZIPTF and C-ZIPTF for better tensor factorization of zero-inflated count data.
problem Inefficient tensor factorization for zero-inflated count data, especially in scRNA-seq.
method Zero Inflated Poisson Tensor Factorization (ZIPTF) and Consensus Zero Inflated Poisson Tensor Factorization (C-ZIPTF).
result ZIPTF and C-ZIPTF improve tensor factorization accuracy and consistency for zero-inflated count data.
New method improves mHealth user engagement using Thompson sampling for count data.
problem Optimizing mHealth interventions for distal outcomes through proximal context.
method Combines count data models with Thompson sampling for contextual bandits.
result Improves user engagement in mHealth trials compared to existing methods.
The paper models and predicts co-occurrence counts using Gamma regression.
problem Predicting relevance between items or users from high-dimensional sparse co-occurrence count data.
method Shared parameter alternating zero-inflated Gamma regression models (SA-ZIG) with Fisher scoring and learning rate adjustment.
result SA-ZIG with learning rate adjustment performs satisfactorily in predicting relevance.
Semiparametric STAR model improves mental health data analysis.
problem Overdispersed, zero-inflated, bounded count data in self-reported mental health surveys.
method STAR transformation and rounding of latent Gaussian model, nonparametric transformation estimation, EM algorithm for maximum likelihood.
result Substantial improvements in goodness-of-fit compared to existing models.
A composite loss framework is proposed for low-rank modeling of data consisting of interesting and common values, such as excess zeros or missing values. The methodology is motivated by the generalized low-rank framework and the hurdle method which is commonly used to analyze zero-inflated counts. The model is demonstr…
Bayesian framework for semiparametric regression of discrete data.
problem Complex distributional features of discrete data.
method Semiparametric modeling with nonparametric marginal and latent linear regression.
result Posterior consistency and analytical/posterior predictive distributions.
Paper proposes copula-based models for analyzing multivariate zero-inflated continuous data.
problem Challenges in analyzing multivariate zero-inflated continuous data with mixed discreteness and continuity.
method Proposes two copula-based density estimation models and rectified Gaussian copula.
result Demonstrates superior performance compared to conventional methods.
Zero-inflated datasets, which have an excess of zero outputs, are commonly encountered in problems such as climate or rare event modelling. Conventional machine learning approaches tend to overestimate the non-zeros leading to poor performance. We propose a novel model family of zero-inflated Gaussian processes (ZiGP) …
The seemingly disjoint problems of count and mixture modeling are united under the negative binomial (NB) process. A gamma process is employed to model the rate measure of a Poisson process, whose normalization provides a random probability measure for mixture modeling and whose marginalization leads to an NB process f…
Proposes a new model to predict travel demand with zero-inflated and long-tail characteristics.
problem Sparse and long-tailed travel demand data with many zeros.
method Spatial-Temporal Tweedie Graph Neural Network (STTD) using Tweedie distribution.
result STTD provides accurate predictions and precise confidence intervals.
Deep model tackles zero-inflated multi-species abundance estimation.
problem Predicting species distribution across landscapes with inflated zero counts.
method Proposes a novel deep learning model combining multivariate probit and log-normal distributions.
result Model outperforms existing methods on bird and fish population datasets.
New bandit algorithms improve sparse reward learning.
problem Sparse rewards hinder learning efficiency in real-world bandit applications.
method Developed algorithms based on Upper Confidence Bound and Thompson Sampling for zero-inflated distributions.
result Empirical performance of new algorithms is superior to existing methods.
In a previous analysis the problem of "zero-inflated" time data (caused by high frequency trading in the electronic order book) was handled by left-truncating the inter-arrival times. We demonstrated, using rigorous statistical methods, that the Weibull distribution describes the corresponding stochastic dynamics for a…
A Deep Zero-Inflated Model for Detecting North Atlantic Right Whale Presence
problem Balancing marine conservation and blue economy management
method Deep Zero-Inflated Bernoulli model
result Improved model adequacy and predictive performance
We propose a simple yet powerful framework for modeling integer-valued data, such as counts, scores, and rounded data. The data-generating process is defined by Simultaneously Transforming and Rounding (STAR) a continuous-valued process, which produces a flexible family of integer-valued distributions capable of modeli…
The M5 competition tackles overdispersed retail sales forecasting with GAMLSS.
problem Overdispersed and zero-inflated retail sales data.
method Distributional forecasting using GAMLSS framework.
result GAMLSS provides better probabilistic forecasting for count data.
A new model improves analysis of neural activity from calcium imaging.
problem Statistical modeling of deconvolved calcium signals for neural activity interpretation.
method Proposed a zero-inflated gamma (ZIG) model to characterize calcium responses as a mixture of a gamma distribution and a point mass.
result The ZIG model outperforms simpler models in neural encoding and decoding problems.
Robust Bayesian inference improves model performance on discrete data.
problem Misspecification of discrete-valued models leads to poor inference and prediction.
method Total Variation Distance (TVD) for discrepancy, efficient estimator and inference method.
result Our approach significantly improves predictive performance on various data.
Tornadoes are the most violent of all atmospheric storms. In a typical year, the United States experiences hundreds of tornadoes with associated damages on the order of one billion dollars. Community preparation and resilience would benefit from accurate predictions of these economic losses, particularly as populations…
SimCD simultaneously clusters cells and identifies differential gene expression in scRNA-seq data.
problem Separate clustering and differential expression analysis for scRNA-seq data leads to suboptimal results.
method Develops SimCD, a unified hierarchical gamma-negative binomial model for simultaneous cell clustering and differential expression analysis.
result SimCD outperforms existing methods in discovering cell clusters and capturing dynamic expression changes.
A new SBM for non-negative zero-inflated edge weights in networks.
problem Modeling international trading networks with non-negative zero-inflated edge weights.
method Restricted Tweedie distribution and nodal information accounting.
result Efficient two-step algorithm for estimating covariate effects.
New method predicts positive samples with missing labels.
problem Missing labels due to response-dependent factors.
method P(U)U-O-Mixture algorithm for joint estimation.
result Non-convex algorithm leads to optimal statistical error.
New hypergraph method improves scRNA-seq clustering.
problem Loss of higher-order information and overestimation in coexpression networks.
method Conceptualizing scRNA-seq data as hypergraphs and proposing novel clustering methods.
result Proposed methods outperform existing methods on simulated and real datasets.
In finance, durations between successive transactions are usually modeled by the autoregressive conditional duration model based on a continuous distribution omitting zero values. Zero or close-to-zero durations can be caused by either split transactions or independent transactions. We propose a discrete model allowing…
RainfallBench benchmarks GNSS-based precipitation nowcasting models, addressing complex meteorological challenges.
problem Evaluation of precipitation nowcasting models in meteorology is insufficient due to focus on periodic variables.
method RainfallBench dataset and specialized evaluation protocols for multi-scale, multi-resolution, and extreme rainfall events.
result Bi-Focus Precipitation Forecaster (BFPF) enhances rainfall time series forecasting by incorporating domain-specific priors.
Flow Matching for count data improves sample quality and efficiency.
problem Mapping between count distributions across batches or time points in high-dimensional count data.
method count-FM, a flow-matching framework based on a continuous-time birth-death process with local unit jumps.
result count-FM achieves better sample quality than representative baselines while using fewer parameters.
Enhanced Tweedie model for insurance claims using CatBoost.
problem Accurately modeling aggregate claims with zero-inflated data.
method Refined Tweedie model with boosting methods in CatBoost.
result Marked improvement in model performance for insurance analytics.
HIP method extended to multi-class, Poisson, and Zero-Inflated Poisson outcomes with an R Shiny app.
problem Subgroup heterogeneity in complex diseases like COPD.
method Integrating multiple data views while accounting for subgroup heterogeneity.
result Identified common and subgroup-specific markers of exacerbation frequency in males and females.
Proposes a robust EM algorithm for analyzing incomplete panel count data.
problem Missing reports in panel count data.
method Functional EM algorithm for non-parametric counting process mean function estimation.
result Robust to misspecification of Poisson process assumption and missing completely at random.
New model predicts travel demand uncertainty with high accuracy.
problem Uncertainty and sparsity in sparse travel demand prediction.
method Spatial-Temporal Zero-Inflated Negative Binomial Graph Neural Network (STZINB-GNN).
result STZINB-GNN outperforms benchmarks in predicting travel demand uncertainty.
The paper proposes count echo state networks for forecasting graduate student enrollments.
problem Forecasting graduate student enrollments from historical data.
method Developed hierarchical count echo state networks and compared them to Poisson autoregressions and negative binomial models.
result Hierarchical negative binomial based echo state network is the superior model.
p-SNE embeds Poisson count data into low dimensions preserving structure.
problem Embedding high-dimensional sparse Poisson data into a low-dimensional space.
method p-SNE (Poisson Stochastic Neighbor Embedding) using KL divergence and Hellinger distance.
result p-SNE recovers meaningful structure in real-world count datasets.
Novel Bayesian method for high-dimensional count data prediction.
problem Count data in high-dimensional settings requires feature selection.
method Pseudo-Bayesian framework with scaled Student prior and exponential weights.
result Strong performance compared to Lasso in various settings.
Count data take on non-negative integer values and are challenging to properly analyze using standard linear-Gaussian methods such as linear regression and principal components analysis. Generalized linear models enable direct modeling of counts in a regression context using distributions such as the Poisson and negati…
Better neural arithmetic logic units improve cell counting model generalization.
problem Neural networks struggle with high cell counts outside training data range.
method Introduced Neural Arithmetic Logic Units (NALU) for arithmetic operations in existing architectures.
result Improved cell counting accuracy for higher numeric ranges with better generalization.
Inpatient care is a large share of total health care spending, making analysis of inpatient utilization patterns an important part of understanding what drives health care spending growth. Common features of inpatient utilization measures include zero inflation, over-dispersion, and skewness, all of which complicate st…
The paper improves count data regression models for overdispersed data.
problem Improving regression models for overdispersed count data.
method Double ℓ1-regularized negative binomial regressions. result Oracle inequalities and consistency for Lasso estimators of partial regression coefficients.
Generative model identifies temporal count data components with regime-dependent contributions.
problem Modeling temporal count data with regime-dependent dynamics.
method Generative framework combining regime-adaptive dynamics with Poisson log-normal emissions.
result Established identifiability of the model and revealed co-variation patterns and regime shifts.
Multivariate count data are defined as the number of items of different categories issued from sampling within a population, which individuals are grouped into categories. The analysis of multivariate count data is a recurrent and crucial issue in numerous modelling problems, particularly in the fields of biology and e…
This work refines Cover's theory for binary classification on low-dimensional data.
problem The challenge of analyzing how low-dimensional data structures affect classification models.
method Refines Cover's function-counting theory to account for low-dimensional data structure.
result Derives dichotomy counts and analyzes the impact of data structure on classification models.
Proposes a method to handle sparse multiway count data with false zeros using zero-truncated Poisson regression.
problem Handling sparse multiway count data corrupted by false zeros.
method Zero-truncated Poisson regression with tensor completion.
result Accurate estimation of multiway count data from approximately IR2log22(I) non-zero counts. ENTED efficiently decomposes binary and count tensors using nonparametric Gaussian processes.
problem Handling high-dimensional and sparse binary and count data with traditional tensor decompositions.
method ENTED uses nonparametric Gaussian processes and sparse orthogonal variational inference to handle binary and count tensors.
result ENTED outperforms traditional methods in binary and count tensor completion tasks.
Bayesian model tackles spatial count data issues with flexible non-parametric techniques.
problem Challenges in traditional parametric models for spatial count data with unbalanced distributions and complex dependencies.
method Bayesian semi-parametric spatial dispersed count model combining non-parametric techniques and adapted count models.
result Demonstrates superior performance in managing dispersion and capturing intricate spatial patterns.
Graphical estimation of count time series dependencies.
problem Estimating dependencies between multivariate count time series.
method Parameter-driven generalized linear model with l1-type regularization and MCEM algorithm.
result Characterization of disease spread interdependence and sources/sinks in Greater Mumbai.
Counting the number of clusters, when these clusters overlap significantly is a challenging problem in machine learning. We argue that a purely mathematical quantum theory, formulated using the path integral technique, when applied to non-physics modeling leads to non-physics quantum theories that are statistical in na…
New phases identified in neural scaling laws with compute limits.
problem Understanding neural scaling laws under compute constraints.
method Solved neural scaling model with stochastic gradient descent, derived loss curves, analyzed model-parameter-count phases.
result Identified 4 phases (+3 subphases) in data-complexity/target-complexity phase-plane, derived exponents.