Proposes a method to reconcile count time series forecasts.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New tree-structured Markov fields with Poisson marginals for counting variables.
Enhances knot counting invariant using skew braces.
Enhances psyquandle counting invariants using cocycles.
Bayesian model fuses diverse microbiome data types.
Count data, for example the number of observed cases of a disease in a city, often arise in the fields of healthcare analytics and epidemiology. In this paper, we consider performing regression on multivariate data in which our outcome is a count. Specifically, we derive log-likelihood functions for finite mixtures of …
Time series of counts arise in a variety of forecasting applications, for which traditional models are generally inappropriate. This paper introduces a hierarchical Bayesian formulation applicable to count time series that can easily account for explanatory variables and share statistical strength across groups of rela…
New ZIPLN model accounts for zero-inflation in multivariate count data.
A gamma process dynamic Poisson factor analysis model is proposed to factorize a dynamic count matrix, whose columns are sequentially observed count vectors. The model builds a novel Markov chain that sends the latent gamma random variables at time as the shape parameters of those at time , which are linked …
We define a family of generalizations of the two-variable quandle polynomial. These polynomial invariants generalize in a natural way to eight-variable polynomial invariants of finite biquandles. We use these polynomials to define a family of link invariants which further generalize the quandle counting invariant.
Semiparametric STAR model improves mental health data analysis.
Multivariate count data are defined as the number of items of different categories issued from sampling within a population, which individuals are grouped into categories. The analysis of multivariate count data is a recurrent and crucial issue in numerous modelling problems, particularly in the fields of biology and e…
ZICO learns DAGs from zero-inflated count data efficiently.
Structured high-cardinality data arises in many domains, and poses a major challenge for both modeling and inference. Graphical models are a popular approach to modeling structured data but they are unsuitable for high-cardinality variables. The count-min (CM) sketch is a popular approach to estimating probabilities in…
We prove that the space of complex irreducible polynomials of degree in variables satisfies two forms of homological stability: first, its cohomology stabilizes as increases, and second, its compactly supported cohomology stabilizes as increases. Our topological results are inspired by counting results …
Clustering with variable selection is a challenging yet critical task for modern small-n-large-p data. Existing methods based on sparse Gaussian mixture models or sparse K-means provide solutions to continuous data. With the prevalence of RNA-seq technology and lack of count data modeling for clustering, the current pr…
We develop deep Poisson-gamma dynamical systems (DPGDS) to model sequentially observed multivariate count data, improving previously proposed models by not only mining deep hierarchical latent structure from the data, but also capturing both first-order and long-range temporal dependencies. Using sophisticated but simp…
Learning Markov blanket (MB) structures has proven useful in performing feature selection, learning Bayesian networks (BNs), and discovering causal relationships. We present a formula for efficiently determining the number of MB structures given a target variable and a set of other variables. As expected, the number of…
NegBio-VAE models neural spike counts with negative binomial distribution.
Many popular first-order optimization methods (e.g., Momentum, AdaGrad, Adam) accelerate the convergence rate of deep learning models. However, these algorithms require auxiliary parameters, which cost additional memory proportional to the number of parameters in the model. The problem is becoming more severe as deep l…
Bayesian model tackles spatial count data issues with flexible non-parametric techniques.
New approach improves black-box planning efficiency by discovering focused macros.
We define a two-variable polynomial invariant of finite quandles. In many cases this invariant completely determines the algebraic structure of the quandle up to isomorphism. We use this polynomial to define a family of link invariants which generalize the quandle counting invariant.
We define counting and cocycle enhancement invariants of virtual knots using parity biquandles. These invariants are determined by pairs consisting of a biquandle 2-cocycle φ^0 and a map φ^1 with certain compatibility conditions leading to one-variable or two-variable polynomial invariants of virtual knots. We provide …
A fast algorithm for counting Markov equivalent DAGs and designing experiments.
The paper tackles extrapolation of gene knockouts effects on RNA counts.
This paper presents the Poisson-randomized gamma dynamical system (PRGDS), a model for sequentially observed count tensors that encodes a strong inductive bias toward sparsity and burstiness. The PRGDS is based on a new motif in Bayesian latent variable modeling, an alternating chain of discrete Poisson and continuous …
A directed acyclic graph (DAG) is the most common graphical model for representing causal relationships among a set of variables. When restricted to using only observational data, the structure of the ground truth DAG is identifiable only up to Markov equivalence, based on conditional independence relations among the v…
Proposes PSCCA for estimating correlations and canonical correlations in sparse count data.
Efficient Bayesian variable selection for binomial and negative binomial data.
The paper models and predicts co-occurrence counts using Gamma regression.
Accurate statistical models of neural spike responses can characterize the information carried by neural populations. But the limited samples of spike counts during recording usually result in model overfitting. Besides, current models assume spike counts to be Poisson-distributed, which ignores the fact that many neur…
Bayesian method tackles variable selection in high-dimensional data.
Supervised topic models with a logistic likelihood have two issues that potentially limit their practical use: 1) response variables are usually over-weighted by document word counts; and 2) existing variational inference methods make strict mean-field assumptions. We address these issues by: 1) introducing a regulariz…
We enhance the quandle coloring quiver invariant of oriented knots and links with quandle modules. This results in a two-variable polynomial invariant with specializes to the previous quandle module polynomial invariant as well as to the quandle counting invariant. We provide example computations to show that the enhan…
In this paper, we determine the bifurcation set of a real polynomial function of two variables for non-degenerate case in the sense of Newton polygons by using a toric compactification. We also count the number of singular phenomena at infinity, called "cleaving" and "vanishing" in the same setting. Finally, we give an…
Proposes a new DNN framework for count data with high-cardinality features.
In this note, we study the dynamics and associated zeta functions of conformally compact manifolds with variable negative sectional curvatures. We begin with a discussion of a larger class of manifolds known as convex co-compact manifolds with variable negative curvature. Applying results from dynamics on these spaces,…
Counting tripods on a flat torus using lattice point counting.
Develops a tool to identify abnormal blood smear results based on CBC tests.
Latent Dirichlet analysis, or topic modeling, is a flexible latent variable framework for modeling high-dimensional sparse count data. Various learning algorithms have been developed in recent years, including collapsed Gibbs sampling, variational inference, and maximum a posteriori estimation, and this variety motivat…
We study statistical aspects of state-dependent Hawkes processes, which are an extension of Hawkes processes where a self- and cross-exciting counting process and a state process are fully coupled, interacting with each other. The excitation kernel of the counting process depends on the state process that, reciprocally…
Flow Matching for count data improves sample quality and efficiency.
Estimates unknown population sizes using the hypergeometric distribution.
In X-ray binary star systems consisting of a compact object that accretes material from an orbiting secondary star, there is no straightforward means to decide if the compact object is a black hole or a neutron star. To assist this classification, we develop a Bayesian statistical model that makes use of the fact that …
Using the interpretation of certain generalised Donaldson-Thomas invariants, including stable pairs curve counts, as the monodromy of a flat connection on a formal principal bundle, we show that the conjectural Gopakumar-Vafa contributions of all genera to the Gromov-Witten partition function appear in the asymptotics …
TimeCNN improves forecasting by refining cross-variable interactions over time.
New theorem counts curves on orbifolds.