The paper proposes a variational autoencoder for discrete data analysis.
problem Sparse, high-dimensional, and overdispersed discrete data analysis.
method Variational autoencoder based on negative-binomial distribution.
result The proposed models achieve significantly better performance on text analysis, collaborative filtering, and multi-label learning compared to state-of-the-art baselines.
The paper improves count data regression models for overdispersed data.
problem Improving regression models for overdispersed count data.
method Double ℓ1-regularized negative binomial regressions. result Oracle inequalities and consistency for Lasso estimators of partial regression coefficients.
Deviance-style normalization for sparse, jointly overdispersed count matrices
problem Jointly overdispersed count matrices
method Dirichlet-multinomial deviance residualization
result Preserves exact sparsity, evaluates in constant time, recovers multinomial residual
The M5 competition tackles overdispersed retail sales forecasting with GAMLSS.
problem Overdispersed and zero-inflated retail sales data.
method Distributional forecasting using GAMLSS framework.
result GAMLSS provides better probabilistic forecasting for count data.
Proposes a new model to better handle overdispersed count time series.
problem Heterogeneous overdispersed count time series.
method Negative-Binomial Randomized Gamma Markov Process.
result Significantly improves predictive performance and fast convergence of inference algorithm.
We introduce overdispersed black-box variational inference, a method to reduce the variance of the Monte Carlo estimator of the gradient in black-box variational inference. Instead of taking samples from the variational distribution, we use importance sampling to take samples from an overdispersed distribution in the s…
Bayesian model learns cancer subtypes from diverse NGS data.
problem Overdispersed NGS count data and limited samples for specific cancer types.
method Bayesian Multi-Domain Learning (BMDL) model using hierarchical negative binomial factorization.
result BMDL achieves reproducible cancer subtyping without negative transfer effects.
The study analyzes a model for aggregate losses with dependent and overdispersed inter-losses times.
problem Analyzing aggregate loss models with dependent and overdispersed inter-losses times.
method The study uses a two-state Markovian arrival process (MAP2) and a Markov renewal process to model the inter-losses times. Severities are modeled using a heavy-tailed, double-Pareto Lognormal distribution. The model is estimated via direct maximization of the likelihood function.
result The model with dependence and overdispersion in inter-losses times leads to higher capital charges compared to a Poisson process.
Study optimizes smart contract adoption under high demand variability using Negative Binomial models.
problem Effective supply chain management under high demand variability.
method Combines dynamic Negative Binomial demand modeling with endogenous smart contract adoption optimization.
result The NB model outperforms other benchmarks in forecasting and optimizing smart contract adoption and order quantity.
Robust Bayesian inference improves model performance on discrete data.
problem Misspecification of discrete-valued models leads to poor inference and prediction.
method Total Variation Distance (TVD) for discrepancy, efficient estimator and inference method.
result Our approach significantly improves predictive performance on various data.
NegBio-VAE models neural spike counts with negative binomial distribution.
problem Limited biological plausibility of continuous latent variables in VAEs for neural spike modeling.
method Proposes a negative binomial latent-variable model with a dispersion parameter for overdispersed spike count modeling.
result NegBio-VAE outperforms competing models in reconstruction and generation tasks.
Bayesian framework for semiparametric regression of discrete data.
problem Complex distributional features of discrete data.
method Semiparametric modeling with nonparametric marginal and latent linear regression.
result Posterior consistency and analytical/posterior predictive distributions.
PHIBP models complex microbiome data with shared parameters.
problem Complex, sparse count data in microbiome analysis.
method Bayesian nonparametric framework with shared species parameters.
result Flexible multivariate count model with tractable inference.
Temporal aggregation reveals latent default correlation from monthly data.
problem Understanding effective default correlation from monthly default data.
method Temporal coarse-graining of latent default-probability paths.
result Temporal coarse-graining improves identifiability and reduces over-allocation of long-horizon fluctuations.
We present a new algorithm for boosting generalized additive models for location, scale and shape (GAMLSS) that allows to incorporate stability selection, an increasingly popular way to obtain stable sets of covariates while controlling the per-family error rate (PFER). The model is fitted repeatedly to subsampled data…
Temporal coarse-graining of latent default paths explains effective correlation in corporate defaults.
problem Understanding effective default correlation in corporate defaults.
method Temporal coarse-graining of latent default-probability paths, applied to corporate default-count data.
result Temporal coarse-graining provides a scale-consistent baseline that improves identifiability and reduces over-allocation of long-horizon fluctuations.
Unified NMF models for various noise distributions, improving feature extraction.
problem Inadequate assumptions for NMF under complex data distributions.
method Unified framework using MM-algorithms for traditional and convex NMF under Tweedie and Negative Binomial models.
result Unified multiplicative update rules for all models, including novel updates for convex NMF.
Bayesian bandits misspecification affects UX optimization, revealing new models.
problem Misspecification of value models in Bayesian bandits impacts UX optimization.
method Formulated UXO as a restless, sleeping bandit with unobserved confounders and optional stopping. Provided model extensions to address misspecifications.
result Common misspecifications lead to sub-optimal rewards, demonstrating overdispersion's effects on bandit performance.
New model predicts weekly earthquakes with better tail risk assessment.
problem Violation of Poisson assumption in seismic data.
method Neural network for per-cell overdispersion estimation.
result 8.6% reduction in mean pinball deviation, 12.5% lower CRPS in tail events.
Learning DAG or Bayesian network models is an important problem in multi-variate causal inference. However, a number of challenges arises in learning large-scale DAG models including model identifiability and computational complexity since the space of directed graphs is huge. In this paper, we address these issues in …
The paper proposes count echo state networks for forecasting graduate student enrollments.
problem Forecasting graduate student enrollments from historical data.
method Developed hierarchical count echo state networks and compared them to Poisson autoregressions and negative binomial models.
result Hierarchical negative binomial based echo state network is the superior model.
We present a Bayesian non-negative tensor factorization model for count-valued tensor data, and develop scalable inference algorithms (both batch and online) for dealing with massive tensors. Our generative model can handle overdispersed counts as well as infer the rank of the decomposition. Moreover, leveraging a repa…
Count data, for example the number of observed cases of a disease in a city, often arise in the fields of healthcare analytics and epidemiology. In this paper, we consider performing regression on multivariate data in which our outcome is a count. Specifically, we derive log-likelihood functions for finite mixtures of …
Neural Diffusion Intensity Models simplify Cox processes inference.
problem Intractable nonparametric estimation and posterior inference of latent stochastic intensity in Cox processes.
method Variational framework using neural SDEs, with theoretical guarantee of ELBO maximization coinciding with maximum likelihood estimation.
result Accurate recovery of latent intensity dynamics and posterior paths with significant speedup.
A new STAR framework models integer-valued data with flexible distributions.
problem Modeling integer-valued data with flexibility and accuracy.
method Simultaneously Transforming and Rounding (STAR) a continuous-valued process.
result STAR framework designs a new BART model for integer-valued data with impressive predictive accuracy.
Selective inference for group lasso estimators across various distributions and covariates.
problem Developing selective inference methods for group lasso estimators.
method Randomized group-regularized optimization problem with post-selection likelihood.
result Selective point estimator and Wald-type confidence regions for regression parameters.
This paper proposes learning to jump for generative modeling of sparse, skewed, heavy-tailed data.
problem Limited ability of diffusion models in modeling sparse, skewed, heavy-tailed data.
method Forward count thinning process and reverse count thickening process to train a deep neural network.
result Learning to jump performs better than learning to denoise for non-negative, sparse data.
Semiparametric STAR model improves mental health data analysis.
problem Overdispersed, zero-inflated, bounded count data in self-reported mental health surveys.
method STAR transformation and rounding of latent Gaussian model, nonparametric transformation estimation, EM algorithm for maximum likelihood.
result Substantial improvements in goodness-of-fit compared to existing models.
New method improves PCA for high-dimensional data with n < p.
problem PCA struggles in high-dimensional settings with n < p.
method Pairwise differences covariance estimation with four regularized versions.
result Proposed methods outperform existing estimators in high-dimensional data settings.
Modeling infection hotspots to quantify effects of contact tracing and testing.
problem Capturing the role of infection hotspots in disease transmission.
method Temporal point process modeling framework to represent visits and disease transmission.
result Estimation of transmission rates at sites and households using Bayesian optimization.
Hybrid model improves geopolitical conflict forecasting.
problem Forecasting geopolitical events from sparse, bursty data.
method Sparse Temporal Fusion Transformer (TFT) + Variational Nearest Neighbor Gaussian Process (VNNGP).
result Consistently outperforms standalone TFT in long-range horizons.
Bayesian models predict Collatz stopping times with high accuracy.
problem Predicting the total stopping time of Collatz sequences.
method Developed two complementary models: a hierarchical Negative Binomial regression and a mechanistic generative approximation.
result Bayesian models outperform generative approximations in predicting Collatz stopping times.
Data preprocessing improves data quality for robust data mining.
problem Noisy and incomplete data hinders data mining models.
method Overview of data cleaning, transformation, and preprocessing methods.
result Preprocessing significantly affects data mining model performance.
Big data sets must be carefully partitioned into statistically similar data subsets that can be used as representative samples for big data analysis tasks. In this paper, we propose the random sample partition (RSP) data model to represent a big data set as a set of non-overlapping data subsets, called RSP data blocks,…
A new method for handling imbalanced big data using ensembles and smart data.
problem Imbalanced data distribution in big data scenarios.
method Smart Data driven Decision Trees Ensemble (SD_DeTE) methodology.
result SD_DeTE outperforms Random Forest in handling imbalanced binary classification problems in big data.
Prevents sensitive data generation in diffusion models using labeled and unlabeled data.
problem Generating sensitive data in diffusion models using unlabeled data.
method Positive-Unlabeled Diffusion Models, approximating ELBO with labeled and unlabeled data.
result Prevents the generation of sensitive data without compromising image quality.
Study reveals Data Shapley's inconsistent performance in data selection tasks.
problem Inconsistency of Data Shapley's performance in data selection across different settings.
method Hypothesis testing framework and identification of utility functions.
result Data Shapley's performance is no better than random selection without specific constraints.
Survey on data collection challenges in machine learning.
problem Data scarcity and need for labeled data in machine learning.
method Comprehensive study of data acquisition, labeling, and improvement techniques.
result Identification of research challenges in data collection.
PRRO generates synthetic tabular data that improves SL performance and class distribution.
problem Low SL utility of synthetic data due to class imbalance and overlooked data relationships.
method Data pruning and column reordering to optimize SL utility.
result Synthetic data generated with PRRO enhances predictive performance and class distribution.
Defines data science as a natural ecosystem with challenges and missions.
problem Challenges and missions in data science due to 5D complexities and data life cycle phases.
method Systemic and data-centric view of data science as a fusion of data universe and its challenges, formalizing a general-purpose architecture.
result Essential data science as a natural ecosystem integrating specific disciplines and high-impact applications.
Synthetic data enhances analytics but requires careful volume management.
problem Accuracy of statistical methods on synthetic data vs. raw data.
method Synthetic Data Generation for Analytics framework using tabular diffusion models.
result Error rate decreases with more synthetic data but may stabilize or increase.
DPASF stream preprocesses Big Data streams efficiently.
problem Efficient preprocessing of streaming Big Data.
method Implemented six preprocessing algorithms in Apache Flink.
result Preprocessing improves data accuracy in streaming Big Data.
This paper introduces C-DSL to improve data mining outcomes by considering context.
problem Data collection ambiguities, data imbalance, hidden biases, lack of domain info, and data incompleteness.
method Developed Context-Driven Data Science Lifecycle (C-DSL) to address data quality issues.
result Tangible improvements to data mining outcomes were achieved through C-DSL.
Proposes using probabilistic models for privacy-preserving synthetic data.
problem Designing high-quality synthetic data for privacy preservation.
method Formulate the problem through probabilistic modelling, choosing a model for the data.
result Statistical discoveries can be reliably reproduced from synthetic data.
Unlabeled data helps stop active learning better than labeled data.
problem Reducing the need for manual annotation in text classification.
method Compared stopping methods based on labeled, unlabeled, and training data.
result Stopping methods using unlabeled data are more effective.
New test ensures quality of shared data in machine learning.
problem Ensuring quality of external data in machine learning tasks.
method Distribution-free two-sample testing procedures grounded in conformal outlier detection.
result Identifies valuable external data agents for model personalization.
Paper creates fair synthetic data ensuring equal predictions across sensitive attributes.
problem Ensuring fair predictions across sensitive attributes in synthetic data.
method Equalizing target probability distributions across sensitive attributes in synthetic data generation.
result Synthetic data provides strong fair predictions, equal across all thresholds.
A new method classifies multiple correlated data streams simultaneously.
problem Classifying multiple correlated data streams in practical scenarios.
method Double-Coupling Support Vector Machines (DC-SVM) considers both internal and external correlations.
result The proposed method outperforms traditional methods on artificial and real-world data streams.