Detecting and recovering labels in binomial logistic mixtures is challenging due to an information gap.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Extends CRR model with q-binomial random walks for asset pricing.
Efficient Bayesian variable selection for binomial and negative binomial data.
Two approaches improve parameter learning in various mixture models.
A new tree model, GRST, improves option pricing without log-normality assumptions.
Bayesian method tackles variable selection in high-dimensional data.
The paper models stock returns using -Gaussians and negative binomials.
Study optimizes smart contract adoption under high demand variability using Negative Binomial models.
The logistic regression model is known to converge to a Poisson point process model if the binary response tends to infinitely imbalanced. In this paper, it is shown that this phenomenon is universal in a wide class of link functions on binomial regression. The proof relies on the extreme value theory. For the logit, p…
Survival MDN uses invertible functions to speed up survival analysis models.
The thesis examines stochastic calculus in option pricing with logistic models and numerical methods.
Clustering with variable selection is a challenging yet critical task for modern small-n-large-p data. Existing methods based on sparse Gaussian mixture models or sparse K-means provide solutions to continuous data. With the prevalence of RNA-seq technology and lack of count data modeling for clustering, the current pr…
We propose a new data-augmentation strategy for fully Bayesian inference in models with binomial likelihoods. The approach appeals to a new class of Polya-Gamma distributions, which are constructed in detail. A variety of examples are presented to show the versatility of the method, including logistic regression, negat…
Paper optimizes clustering for multi-layer networks and discrete mixtures.
We present a family of expectation-maximization (EM) algorithms for binary and negative-binomial logistic regression, drawing a sharp connection with the variational-Bayes algorithm of Jaakkola and Jordan (2000). Indeed, our results allow a version of this variational-Bayes approach to be re-interpreted as a true EM al…
By developing data augmentation methods unique to the negative binomial (NB) distribution, we unite seemingly disjoint count and mixture models under the NB process framework. We develop fundamental properties of the models and derive efficient Gibbs sampling inference. We show that the gamma-NB process can be reduced …
We compare two statistical models of three binary random variables. One is a mixture model and the other is a product of mixtures model called a restricted Boltzmann machine. Although the two models we study look different from their parametrizations, we show that they represent the same set of distributions on the int…
The traditional Minkowski distances are induced by the corresponding Minkowski norms in real-valued vector spaces. In this work, we propose novel statistical symmetric distances based on the Minkowski's inequality for probability densities belonging to Lebesgue spaces. These statistical Minkowski distances admit closed…
The seemingly disjoint problems of count and mixture modeling are united under the negative binomial (NB) process. A gamma process is employed to model the rate measure of a Poisson process, whose normalization provides a random probability measure for mixture modeling and whose marginalization leads to an NB process f…
The paper introduces the concept of a cluster structure to define a joint distribution of the sample size and its exchangeable random partitions. The cluster structure allows the probability distribution of the random partitions of a subset of the sample to be dependent on the sample size, a feature not presented in a …
ETM models improve efficiency in semi-supervised logistic regression.
Introduces Soft-SVM for binary classification bridging logistic and SVM.
Consider semi-supervised learning for classification, where both labeled and unlabeled data are available for training. The goal is to exploit both datasets to achieve higher prediction accuracy than just using labeled data alone. We develop a semi-supervised logistic learning method based on exponential tilt mixture m…
Study calculates tail risk for various mixture distributions.
Researchers show mixtures of ranking models are generally identifiable.
Count data, for example the number of observed cases of a disease in a city, often arise in the fields of healthcare analytics and epidemiology. In this paper, we consider performing regression on multivariate data in which our outcome is a count. Specifically, we derive log-likelihood functions for finite mixtures of …
Paper establishes universal lower bounds and optimal rates for clustering sub-exponential mixture models.
We improve MoE models for classification with rigorous guarantees and practical methods.
PbP strategy improves logistic model prediction with missing values.
VBphenoR uses variational Bayes for EHR-based patient phenotyping.
We use the theory of normal variance-mean mixtures to derive a data augmentation scheme for models that include gamma functions. Our methodology applies to many situations in statistics and machine learning, including Multinomial-Dirichlet distributions, Negative binomial regression, Poisson-Gamma hierarchical models, …
We present a new modeling technique for solving the problem of ecological inference, in which individual-level associations are inferred from labeled data available only at the aggregate level. We model aggregate count data as arising from the Poisson binomial, the distribution of the sum of independent but not identic…
We consider a problem of ecological inference, in which individual-level covariates are known, but labeled data is available only at the aggregate level. The intended application is modeling voter preferences in elections. In Rosenman and Viswanathan (2018), we proposed modeling individual voter probabilities via a log…
The paper establishes convergence rates for MoE models in classification problems.
Improved regret bound for MNL MDPs with variance-aware approach.
Proposes efficient Bayesian logistic regression for large sparse datasets.
Probabilistic Bisection Algorithm performs root finding based on knowledge acquired from noisy oracle responses. We consider the generalized PBA setting (G-PBA) where the statistical distribution of the oracle is unknown and location-dependent, so that model inference and Bayesian knowledge updating must be performed s…
Develops Bayesian inference methods for gamma models.
The paper improves Fisher-Pitman tests for Poisson mixtures, detecting autism-related genes.
Proposes logistic-beta process for modeling dependent probabilities with beta marginals.
The study tightens bounds on binomial probabilities and minimums using KL-divergence.
The paper tackles learning mixtures of two multinomial logits, showing identifiability and presenting an algorithm.
Machine learning improves joint default assessment by capturing non-linear dependencies.
The paper introduces flat-topped PDFs for better fitting machine learning models.
PixelCNNs are a recently proposed class of powerful generative models with tractable likelihood. Here we discuss our implementation of PixelCNNs which we make available at https://github.com/openai/pixel-cnn. Our implementation contains a number of modifications to the original model that both simplify its structure an…
Improved logistic MoE with sigmoid gate shows better sample efficiency.
We construct a binomial tree model fitting all moments to the approximated geometric Brownian motion. Our construction generalizes the classical Cox-Ross-Rubinstein, the Jarrow-Rudd, and the Tian binomial tree models. The new binomial model is used to resolve a discontinuity problem in option pricing.
In this paper we propose a novel framework for the construction of sparsity-inducing priors. In particular, we define such priors as a mixture of exponential power distributions with a generalized inverse Gaussian density (EP-GIG). EP-GIG is a variant of generalized hyperbolic distributions, and the special cases inclu…