MPVAE learns latent embeddings and label correlations for multi-label classification.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The multivariate probit model (MVP) is a popular classic model for studying binary responses of multiple entities. Nevertheless, the computational challenge of learning the MVP model, given that its likelihood involves integrating over a multidimensional constrained space of latent variables, significantly limits its a…
EP method speeds up Bayesian probit regression in high dimensions.
The logistic normal distribution has recently been adapted via the transformation of multivariate Gaus- sian variables to model the topical distribution of documents in the presence of correlations among topics. In this paper, we propose a probit normal alternative approach to modelling correlated topical structures. O…
New conjugate priors improve Bayesian inference for multinomial probit models.
Two methods use BART to model missing data in leaf photosynthetic trait data.
Deep model tackles zero-inflated multi-species abundance estimation.
New method for efficient maximum likelihood estimation of -generalized probit regression.
New algorithms improve MPBART for HIV patient data.
Binary classification models get more efficient predictive probabilities.
Probit Monotone BART estimates binary outcomes using monotonic functions.
Efficient EP algorithm improves smoothing distribution inference in financial models.
Bayesian methods improve inference for cumulative probit models on large datasets.
Unified Skew-Gaussian process framework for various regression and classification tasks.
Proposes a new tensor factorization model for better link prediction in knowledge graphs.
Bayesian method models binary response and covariates for two groups, estimating causal relationships.
An efficient algorithm selects the correct number of latent dimensions in multidimensional probit models.
Learning multiple tasks across heterogeneous domains is a challenging problem since the feature space may not be the same for different tasks. We assume the data in multiple tasks are generated from a latent common domain via sparse domain transforms and propose a latent probit model (LPM) to jointly learn the domain t…
Paper develops Bayesian inference for discrete-choice mnp models with Gaussian priors.
This article proposes Multinomial Probit Bayesian Additive Regression Trees (MPBART) as a multinomial probit extension of BART - Bayesian Additive Regression Trees (Chipman et al (2010)). MPBART is flexible to allow inclusion of predictors that describe the observed units as well as the available choice alternatives. T…
Linear Mixed Models (LMMs) are important tools in statistical genetics. When used for feature selection, they allow to find a sparse set of genetic traits that best predict a continuous phenotype of interest, while simultaneously correcting for various confounding factors such as age, ethnicity and population structure…
We consider analysis of relational data (a matrix), in which the rows correspond to subjects (e.g., people) and the columns correspond to attributes. The elements of the matrix may be a mix of real and categorical. Each subject and attribute is characterized by a latent binary feature vector, and an inferred matrix map…
The study examines mixing times of data-augmentation Gibbs samplers for high-dimensional probit regression.
Improves graph-based active learning for non-Gaussian models.
Probit regression was first proposed by Bliss in 1934 to study mortality rates of insects. Since then, an extensive body of work has analyzed and used probit or related binary regression methods (such as logistic regression) in numerous applications and fields. This paper provides a fresh angle to such well-established…
Novel Bayesian model improves EEG-based BCI character selection.
SBPMT combines bagging and boosting for improved classification.
Efficiently identifies important variables in binary outcomes using variational Bayes.
We study convex empirical risk minimization for high-dimensional inference in binary models. Our first result sharply predicts the statistical performance of such estimators in the linear asymptotic regime under isotropic Gaussian features. Importantly, the predictions hold for a wide class of convex loss functions, wh…
We consider probabilistic multinomial probit classification using Gaussian process (GP) priors. The challenges with the multiclass GP classification are the integration over the non-Gaussian posterior distribution, and the increase of the number of unknown latent variables as the number of target classes grows. Expecta…
Skew Gaussian Processes improve classification performance by allowing asymmetry.
This paper provides a theoretical and computational justification of the long held claim that of the similarity of the probit and logit link functions often used in binary classification. Despite this widespread recognition of the strong similarities between these two link functions, very few (if any) researchers have …
New algorithm learns sparse GLMs for binary outcomes efficiently.
Graph-based semi-supervised learning is the problem of propagating labels from a small number of labelled data points to a larger set of unlabelled data. This paper is concerned with the consistency of optimization-based techniques for such problems, in the limit where the labels have small noise and the underlying unl…
We analyze mixing times of three DA algorithms for regression models.
Paper introduces symmetric divergence link models for probability distributions.
Characterizes uncertainty in high-dimensional linear classification models.
Classification of high dimensional data finds wide-ranging applications. In many of these applications equipping the resulting classification with a measure of uncertainty may be as important as the classification itself. In this paper we introduce, develop algorithms for, and investigate the properties of, a variety o…
The logistic regression model is known to converge to a Poisson point process model if the binary response tends to infinitely imbalanced. In this paper, it is shown that this phenomenon is universal in a wide class of link functions on binomial regression. The proof relies on the extreme value theory. For the logit, p…
Algorithm improves variational inference in Wasserstein distance.
We study the dynamical behavior of high-frequency data from the Korean Stock Price Index (KOSPI) using the movement of returns in Korean financial markets. The dynamical behavior for a binarized series of our models is not completely random. The conditional probability is numerically estimated from a return series of K…
We introduce a novel kernel that models input-dependent couplings across multiple latent processes. The pairwise joint kernel measures covariance along inputs and across different latent signals in a mutually-dependent fashion. A latent correlation Gaussian process (LCGP) model combines these non-stationary latent comp…
With the scale of data growing every day, reducing the dimensionality (a.k.a. sketching) of high-dimensional data has emerged as a task of paramount importance. Relevant issues to address in this context include the sheer volume of data that may consist of categorical samples, the typically streaming format of acquisit…
Generating user interpretable multi-class predictions in data rich environments with many classes and explanatory covariates is a daunting task. We introduce Diagonal Orthant Latent Dirichlet Allocation (DOLDA), a supervised topic model for multi-class classification that can handle both many classes as well as many co…
Designs efficient algorithms for online and sliding window models of subspace embeddings for all p.
Zero-inflated datasets, which have an excess of zero outputs, are commonly encountered in problems such as climate or rare event modelling. Conventional machine learning approaches tend to overestimate the non-zeros leading to poor performance. We propose a novel model family of zero-inflated Gaussian processes (ZiGP) …
The paper estimates CoVaR with various models for financial risk analysis.
In this paper, we consider the multivariate Bernoulli distribution as a model to estimate the structure of graphs with binary nodes. This distribution is discussed in the framework of the exponential family, and its statistical properties regarding independence of the nodes are demonstrated. Importantly the model can e…