New method for efficient maximum likelihood estimation of -generalized probit regression.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This article proposes Multinomial Probit Bayesian Additive Regression Trees (MPBART) as a multinomial probit extension of BART - Bayesian Additive Regression Trees (Chipman et al (2010)). MPBART is flexible to allow inclusion of predictors that describe the observed units as well as the available choice alternatives. T…
EP method speeds up Bayesian probit regression in high dimensions.
Probit Monotone BART estimates binary outcomes using monotonic functions.
Probit regression was first proposed by Bliss in 1934 to study mortality rates of insects. Since then, an extensive body of work has analyzed and used probit or related binary regression methods (such as logistic regression) in numerous applications and fields. This paper provides a fresh angle to such well-established…
The study examines mixing times of data-augmentation Gibbs samplers for high-dimensional probit regression.
Bayesian methods improve inference for cumulative probit models on large datasets.
Binary classification models get more efficient predictive probabilities.
Efficiently identifies important variables in binary outcomes using variational Bayes.
Paper develops Bayesian inference for discrete-choice mnp models with Gaussian priors.
Linear Mixed Models (LMMs) are important tools in statistical genetics. When used for feature selection, they allow to find a sparse set of genetic traits that best predict a continuous phenotype of interest, while simultaneously correcting for various confounding factors such as age, ethnicity and population structure…
We analyze mixing times of three DA algorithms for regression models.
New algorithm learns sparse GLMs for binary outcomes efficiently.
Unified Skew-Gaussian process framework for various regression and classification tasks.
The logistic regression model is known to converge to a Poisson point process model if the binary response tends to infinitely imbalanced. In this paper, it is shown that this phenomenon is universal in a wide class of link functions on binomial regression. The proof relies on the extreme value theory. For the logit, p…
The logistic normal distribution has recently been adapted via the transformation of multivariate Gaus- sian variables to model the topical distribution of documents in the presence of correlations among topics. In this paper, we propose a probit normal alternative approach to modelling correlated topical structures. O…
Two methods use BART to model missing data in leaf photosynthetic trait data.
Deep model tackles zero-inflated multi-species abundance estimation.
This paper provides a theoretical and computational justification of the long held claim that of the similarity of the probit and logit link functions often used in binary classification. Despite this widespread recognition of the strong similarities between these two link functions, very few (if any) researchers have …
New conjugate priors improve Bayesian inference for multinomial probit models.
Algorithm improves variational inference in Wasserstein distance.
Efficient EP algorithm improves smoothing distribution inference in financial models.
MPVAE learns latent embeddings and label correlations for multi-label classification.
The multivariate probit model (MVP) is a popular classic model for studying binary responses of multiple entities. Nevertheless, the computational challenge of learning the MVP model, given that its likelihood involves integrating over a multidimensional constrained space of latent variables, significantly limits its a…
Learning multiple tasks across heterogeneous domains is a challenging problem since the feature space may not be the same for different tasks. We assume the data in multiple tasks are generated from a latent common domain via sparse domain transforms and propose a latent probit model (LPM) to jointly learn the domain t…
Bayesian method models binary response and covariates for two groups, estimating causal relationships.
Proposes a new tensor factorization model for better link prediction in knowledge graphs.
An efficient algorithm selects the correct number of latent dimensions in multidimensional probit models.
New algorithms improve MPBART for HIV patient data.
Black-box alpha (BB-) is a new approximate inference method based on the minimization of -divergences. BB- scales to large datasets because it can be implemented using stochastic gradient descent. BB- can be applied to complex probabilistic models with little effort since it only requires as input the likel…
Designs efficient algorithms for online and sliding window models of subspace embeddings for all p.
Characterizes uncertainty in high-dimensional linear classification models.
We consider probabilistic multinomial probit classification using Gaussian process (GP) priors. The challenges with the multiclass GP classification are the integration over the non-Gaussian posterior distribution, and the increase of the number of unknown latent variables as the number of target classes grows. Expecta…
Novel Bayesian model improves EEG-based BCI character selection.
Improves graph-based active learning for non-Gaussian models.
SoftBart improves BART for high-noise modeling in science.
Proposes efficient Bayesian logistic regression for large sparse datasets.
Efficient cross-validation for multi-penalty ridge regression.
SBPMT combines bagging and boosting for improved classification.
Statistical models with constrained probability distributions are abundant in machine learning. Some examples include regression models with norm constraints (e.g., Lasso), probit, many copula models, and latent Dirichlet allocation (LDA). Bayesian inference involving probability distributions confined to constrained d…
Generating user interpretable multi-class predictions in data rich environments with many classes and explanatory covariates is a daunting task. We introduce Diagonal Orthant Latent Dirichlet Allocation (DOLDA), a supervised topic model for multi-class classification that can handle both many classes as well as many co…
Graph-based semi-supervised learning is the problem of propagating labels from a small number of labelled data points to a larger set of unlabelled data. This paper is concerned with the consistency of optimization-based techniques for such problems, in the limit where the labels have small noise and the underlying unl…
We study convex empirical risk minimization for high-dimensional inference in binary models. Our first result sharply predicts the statistical performance of such estimators in the linear asymptotic regime under isotropic Gaussian features. Importantly, the predictions hold for a wide class of convex loss functions, wh…
We show how the problem of estimating conditional Kendall's tau can be rewritten as a classification task. Conditional Kendall's tau is a conditional dependence parameter that is a characteristic of a given pair of random variables. The goal is to predict whether the pair is concordant (value of ) or discordant (val…
We introduce a Bayesian nonparametric regression model for data with multiway (tensor) structure, motivated by an application to periodontal disease (PD) data. Our outcome is the number of diseased sites measured over four different tooth types for each subject, with subject-specific covariates available as predictors.…
The main challenges that arise when adopting Gaussian Process priors in probabilistic modeling are how to carry out exact Bayesian inference and how to account for uncertainty on model parameters when making model-based predictions on out-of-sample data. Using probit regression as an illustrative working example, this …
Variational inference (VI) is widely used as an efficient alternative to Markov chain Monte Carlo. It posits a family of approximating distributions and finds the closest member to the exact posterior . Closeness is usually measured via a divergence from to . While successful, this approach al…
Logistic regression is used thousands of times a day to fit data, predict future outcomes, and assess the statistical significance of explanatory variables. When used for the purpose of statistical inference, logistic models produce p-values for the regression coefficients by using an approximation to the distribution …