Multivariate count data are defined as the number of items of different categories issued from sampling within a population, which individuals are grouped into categories. The analysis of multivariate count data is a recurrent and crucial issue in numerous modelling problems, particularly in the fields of biology and e…
Graphical estimation of count time series dependencies.
problem Estimating dependencies between multivariate count time series.
method Parameter-driven generalized linear model with l1-type regularization and MCEM algorithm.
result Characterization of disease spread interdependence and sources/sinks in Greater Mumbai.
The Poisson distribution has been widely studied and used for modeling univariate count-valued data. Multivariate generalizations of the Poisson distribution that permit dependencies, however, have been far less popular. Yet, real-world high-dimensional count-valued data found in word counts, genomics, and crime statis…
Develops a method to model multivariate count processes with Cox processes and shot noise intensities.
problem Modeling and estimating dependent count processes using granular data.
method Multivariate Cox process with shot noise intensities, connected via Lévy copulas.
result Allows for over-dispersion, auto-correlation, and realistic features in count processes.
New ZIPLN model accounts for zero-inflation in multivariate count data.
problem Zero-inflation in multivariate count data.
method Introduced Zero-Inflated PLN (ZIPLN) model with variational inference.
result ZIPLN significantly improves log-likelihood and reduces dispersion.
We develop deep Poisson-gamma dynamical systems (DPGDS) to model sequentially observed multivariate count data, improving previously proposed models by not only mining deep hierarchical latent structure from the data, but also capturing both first-order and long-range temporal dependencies. Using sophisticated but simp…
Warped DLMs improve forecasting for count time series.
problem Limited options for modeling count time series data.
method Introduces a semiparametric methodology using warping of Gaussian DLMs.
result Demonstrates improved forecasting capabilities for count time series.
We introduce a multivariable Casson-Lin type invariant for links in S3. This invariant is defined as a signed count of irreducible SU(2) representations of the link group with fixed meridional traces. For 2-component links with linking number one, the invariant is shown to be a sum of multivariable …
Flexible models cluster RNA sequencing data.
problem Clustering discrete data from RNA sequencing studies.
method Finite mixtures of multivariate Poisson-log normal factor analyzers with constraints.
result Models give favorable clustering performance on real and simulated data.
We introduce a new dynamical system for sequentially observed multivariate count data. This model is based on the gamma--Poisson construction---a natural choice for count data---and relies on a novel Bayesian nonparametric prior that ties and shrinks the model parameters, thus avoiding overfitting. We present an effici…
Paper introduces MSPD for multivariate risk processes with dependencies.
problem Computing risk valuations with dynamic dependencies between frequency and severity.
method Combines Poisson imbedding, pseudo-chaotic expansion, and Malliavin calculus.
result Explicit general correlation formula for MSPDs.
Count data, for example the number of observed cases of a disease in a city, often arise in the fields of healthcare analytics and epidemiology. In this paper, we consider performing regression on multivariate data in which our outcome is a count. Specifically, we derive log-likelihood functions for finite mixtures of …
In this paper, we develop a new approach to learning high-dimensional Poisson directed acyclic graphical (DAG) models from only observational data without strong assumptions such as faithfulness and strong sparsity. A key component of our method is to decouple the ordering estimation or parent search where the problems…
Study uses regression and ML for COVID-19 mortality forecasting.
problem Forecasting COVID-19 mortality during the first wave in Spain.
method Cyclical curve log-regression, multivariate time series spatial residual correlation analysis, Bayesian approach, machine learning.
result Empirical analysis shows ML regression models perform better than traditional methods.
Fenrir efficiently estimates Bayesian MLN-DLMs for scalable inference.
problem Computational challenges in Bayesian MLN-DLMs for longitudinal count compositional data.
method Novel algorithm for MAP estimation and accurate posterior marginal approximation.
result Fenrir can be three orders of magnitude more efficient than Stan.
New test detects differences in heterogeneous datasets.
problem Detecting differences between two samples with unknown heterogeneity.
method Developed a nonparametric testing procedure that handles latent heterogeneity through a composite null.
result The test accurately detects differences in the presence of unknown heterogeneity.
Bayesian model for discrete data with conditional transformations.
problem Handling discrete ordinal and count data with excess zeros.
method Bayesian framework with conditional transformation functions and modular MCMC algorithm.
result Flexible modeling of linear and nonlinear covariate effects for ordinal and count data.
Proposes PSCCA for estimating correlations and canonical correlations in sparse count data.
problem Estimating correlations and canonical correlations in sparse count data from next-generation sequencing.
method Probabilistic approach for sparse count data sets (PSCCA).
result PSCCA outperforms other methods in estimating true correlations and canonical correlations at the natural parameter level.
Bayesian model fuses diverse microbiome data types.
problem Challenges in fusing different types of microbiome data.
method Flexible multinomial-Gaussian generative model with variational EM algorithm.
result Inferred latent variables provide common dimensionality reduction and predictive posterior distribution.
New exact tests detect changepoints in binary and count data, especially when normal approximations fail.
problem Detecting changepoints in multichannel binary and count data.
method Exact tests combining two-sample conditional tests with multiplicity correction.
result Exact tests are much more powerful than asymptotic tests in various settings.
PHIBP models complex microbiome data with shared parameters.
problem Complex, sparse count data in microbiome analysis.
method Bayesian nonparametric framework with shared species parameters.
result Flexible multivariate count model with tractable inference.
CANN models improve insurance claim count predictions using telematics data.
problem Improving insurance claim count predictions with telematics data.
method Combining classical actuarial models with neural networks for telematics data.
result CANN models outperform traditional models in predicting insurance claims.
MMM model clusters mixed-type longitudinal data efficiently.
problem Challenges in clustering multivariate longitudinal mixed-type data.
method MMM model reorganizes data into a three-way structure, using a mixture of matrix-variate normal distributions.
result MMM model handles various data types (continuous, ordinal, binary, nominal, count) and temporal dependence.
We study hedging and pricing of unattainable contingent claims in a non-Markovian regime-switching financial model. Our financial market consists of a bank account and a risky asset whose dynamics are driven by a Brownian motion and a multivariate counting process with stochastic intensities. The interest rate, drift, …
Study examines how different time series cross-validation methods affect anomaly detection in multivariate time series.
problem Evaluating anomaly detection in multivariate time series requires preserving temporal dependencies, especially for subsequence anomalies.
method Systematically investigates walk-forward and sliding window methods across various validation configurations and classifier types.
result Sliding window method consistently yields higher precision-recall scores and reduced fold-to-fold performance variance, particularly for deep learning models.
GGP models multivariate time series with latent sub-sequences for diverse behaviors.
problem Modeling multivariate time series with diverse behaviors and patterns.
method Graph Gamma Process (GGP) linear dynamical systems with latent sub-sequences.
result GGP models exhibit good predictive performance and reveal interpretable latent patterns.
We propose a multi-wing harmonium model for mining multimedia data that extends and improves on earlier models based on two-layer random fields, which capture bidirectional dependencies between hidden topic aspects and observed inputs. This model can be viewed as an undirected counterpart of the two-layer directed mode…
Assume that M(T) is a rational homology sphere plumbed 3-manifold associated with a connected negative definite graph T. We consider the combinatorial multivariable Poincaré series associated with T and its counting functions, which encode rich topological information. Using the `per…
STICC clusters geographic objects considering both spatial contiguity and attributes.
problem Discovering repeated geographic patterns with spatial contiguity.
method Spatial Toeplitz Inverse Covariance-Based Clustering (STICC) method.
result STICC significantly outperforms baseline methods in adjusted rand index and macro-F1 score.
Extends Benard-Conway invariant to all two-component links.
problem Counting irreducible SU(2) representations for two-component links.
method Counting irreducible SU(2) representations with fixed meridional traces.
result Invariant equals symmetrized multivariable link signature for (2, 2n)-torus links.
Estimates unknown population sizes using the hypergeometric distribution.
problem Estimating discrete distributions with unknown population sizes and category sizes.
method Proposes a novel solution using the hypergeometric likelihood, accounting for a data generating process with a latent variable.
result Empirically demonstrates superior performance in estimating population sizes and learning latent spaces compared to other methods.
Counting tripods on a flat torus using lattice point counting.
problem Counting finite BPS webs in flat torus geometry.
method Lattice point counting techniques in C2. result Asymptotic counting result for tripods on the torus.
TimeCNN improves forecasting by refining cross-variable interactions over time.
problem Multivariate time series forecasting struggles with dynamic and multifaceted cross-variable correlations.
method TimeCNN uses timepoint-independent convolution kernels to capture evolving relationships among variables.
result TimeCNN outperforms state-of-the-art models in real-world datasets with significant computational and speed advantages.
We provide an efficient algorithm for the classical problem, going back to Galton, Pearson, and Fisher, of estimating, with arbitrary accuracy the parameters of a multivariate normal distribution from truncated samples. Truncated samples from a d-variate normal N(μ,Σ) means a samples is only re…
Flow Matching for count data improves sample quality and efficiency.
problem Mapping between count distributions across batches or time points in high-dimensional count data.
method count-FM, a flow-matching framework based on a continuous-time birth-death process with local unit jumps.
result count-FM achieves better sample quality than representative baselines while using fewer parameters.
Bayesian method predicts runtime metrics for fog manufacturing.
problem Accurate prediction of runtime performance metrics in fog manufacturing.
method Bayesian sparse regression for multivariate mixed responses.
result Enhanced prediction and statistical inferences of runtime metrics.
Paper compares MICE-based methods to deep generative models for synthetic data in ratemaking.
problem Limited access to high-quality data for actuarial ratemaking.
method Benchmarked MICE-based models against VAEs and CTA-GANs.
result MICE-based models preserve marginal distributions and multivariate relationships better than deep models.
New theorem counts curves on orbifolds.
problem Counting curves on surfaces.
method Applied Mirzakhani's theorem to orbifolds.
result Curve counting theorem extends to orbifolds.
A new method, Count-MORL, improves offline reinforcement learning by using state-action frequency.
problem Improving offline reinforcement learning performance.
method Integrates count-based conservatism into model-based offline reinforcement learning.
result The learned policy is near-optimal and outperforms existing methods.
Proposes a method to reconcile count time series forecasts.
problem No formal framework for probabilistic reconciliation of count time series.
method Generalizes Bayes' rule for reconciling real-valued and count variables.
result Improves forecast accuracy for count variables compared to Gaussian reconciliation.
Study geodesic paths on flat surfaces, comparing length and singularity counts.
problem Comparing geometric length and singularity counts on geodesic paths.
method Apply counting limit laws to infinite graphs and then to flat surfaces.
result Statistical comparison of geometric length and singularity counts on geodesic paths.
Graph neural networks struggle with counting certain substructures in graphs.
problem Detecting and counting specific substructures in graphs.
method Study of graph neural networks' ability to count attributed graph substructures.
result Graph neural networks like MPNNs, 2-WL, and 2-IGNs have limitations in counting certain substructures.
Counts arcs in surfaces, proving convergence of geodesic currents.
problem Counting arcs of the same type in compact surfaces and related geometries.
method Derives convergence of geodesic currents to prove arc counts.
result Proves convergence of geodesic currents, leading to arc counting results.
Deviance-style normalization for sparse, jointly overdispersed count matrices
problem Jointly overdispersed count matrices
method Dirichlet-multinomial deviance residualization
result Preserves exact sparsity, evaluates in constant time, recovers multinomial residual
The paper proposes count echo state networks for forecasting graduate student enrollments.
problem Forecasting graduate student enrollments from historical data.
method Developed hierarchical count echo state networks and compared them to Poisson autoregressions and negative binomial models.
result Hierarchical negative binomial based echo state network is the superior model.
Counted essential surfaces in a knot's exterior, finding a unique pattern.
problem Counting essential surfaces in a knot's exterior.
method Counted essential surfaces by genus, using Euler totient function. Showed normal surfaces are connected by counting their components. Used Agol, Hass, and Thurston's tools to convert component counting into orbit counting.
result Found a unique pattern in the number of essential surfaces by genus.
New model for clustering dependent community Hawkes processes in temporal networks.
problem Modeling strong dependence and community structure in temporal networks.
method Dependent Community Hawkes (DCH) models combining stochastic block models and Hawkes processes.
result Spectral clustering error bound derived for DCH models.
Counting objects in digital images is a process that should be replaced by machines. This tedious task is time consuming and prone to errors due to fatigue of human annotators. The goal is to have a system that takes as input an image and returns a count of the objects inside and justification for the prediction in the…