BERET improves binary expansion test for multivariate independence.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Modified Metropolis algorithm ensures convergence for multivariate binary distributions with fixed-order updates.
Deep belief networks can approximate any multivariate density with binary hidden units.
BEGIN network models binary data without parametric assumptions.
Multivariate binary data is becoming abundant in current biological research. Logistic principal component analysis (PCA) is one of the commonly used tools to explore the relationships inside a multivariate binary data set by exploiting the underlying low rank structure. We re-expressed the logistic PCA model based on …
Proposes MELODIC family for simultaneous binary logistic regression.
A new method for binary ICA using non-stationary sources.
A method for representing and comparing categorical trajectories using multivariate functional principal components.
Proposes an L1-regularized functional SVM for binary classification with functional covariates.
Archetypal analysis helps understand binary data sets.
Cluster-based ZSL for multivariate data predicts unseen classes.
We consider analysis of relational data (a matrix), in which the rows correspond to subjects (e.g., people) and the columns correspond to attributes. The elements of the matrix may be a mix of real and categorical. Each subject and attribute is characterized by a latent binary feature vector, and an inferred matrix map…
A new clustering algorithm for functional data using binary trees.
Graphical models are commonly used tools for modeling multivariate random variables. While there exist many convenient multivariate distributions such as Gaussian distribution for continuous data, mixed data with the presence of discrete variables or a combination of both continuous and discrete variables poses new cha…
Introduces FairCOCCO for fair learning with multitype, multivariate sensitive attributes.
In this paper, we consider the multivariate Bernoulli distribution as a model to estimate the structure of graphs with binary nodes. This distribution is discussed in the framework of the exponential family, and its statistical properties regarding independence of the nodes are demonstrated. Importantly the model can e…
Large deviation principles for multivariate stochastic volatility models.
A new method prunes neural networks efficiently without losing effectiveness.
A neural network with a single hidden layer can't represent certain multivariable functions.
Sequential or online dimensional reduction is of interests due to the explosion of streaming data based applications and the requirement of adaptive statistical modeling, in many emerging fields, such as the modeling of energy end-use profile. Principal Component Analysis (PCA), is the classical way of dimensional redu…
Bayesian VAR model discovers Granger causality with uncertainty-aware binary graphs.
EP method speeds up Bayesian probit regression in high dimensions.
Dimension reduction of multivariate data supervised by auxiliary information is considered. A series of basis for dimension reduction is obtained as minimizers of a novel criterion. The proposed method is akin to continuum regression, and the resulting basis is called continuum directions. With a presence of binary sup…
Identifies interpretable generative model for multivariate data.
The multivariate probit model (MVP) is a popular classic model for studying binary responses of multiple entities. Nevertheless, the computational challenge of learning the MVP model, given that its likelihood involves integrating over a multidimensional constrained space of latent variables, significantly limits its a…
The study examines how bias affects hypothesis formation in neural networks.
This paper describes a recursive estimation procedure for multivariate binary densities (probability distributions of vectors of Bernoulli random variables) using orthogonal expansions. For covariates, there are basis coefficients to estimate, which renders conventional approaches computationally prohibitive …
Brain decoding is a popular multivariate approach for hypothesis testing in neuroimaging. It is well known that the brain maps derived from weights of linear classifiers are hard to interpret because of high correlations between predictors, low signal to noise ratios, and the high dimensionality of neuroimaging data. T…
There has been a lot of work fitting Ising models to multivariate binary data in order to understand the conditional dependency relationships between the variables. However, additional covariates are frequently recorded together with the binary data, and may influence the dependence relationships. Motivated by such a d…
It is the main goal of this article to address the bipartite ranking issue from the perspective of functional data analysis (FDA). Given a training set of independent realizations of a (possibly sampled) second-order random function with a (locally) smooth autocorrelation structure and to which a binary label is random…
New exact tests detect changepoints in binary and count data, especially when normal approximations fail.
MMM model clusters mixed-type longitudinal data efficiently.
EagleEye detects localized density anomalies in multivariate data.
GGP models multivariate time series with latent sub-sequences for diverse behaviors.
Quantum algorithm estimates multivariate mean with near-optimal efficiency.
Paper proposes a simple estimator for DPP correlation kernels.
Typical dimensionality reduction (DR) methods are often data-oriented, focusing on directly reducing the number of random variables (features) while retaining the maximal variations in the high-dimensional data. In unsupervised situations, one of the main limitations of these methods lies in their dependency on the sca…
MULTIFIT tests independence between two random vectors using multiscale Fisher's test.
Williams and Beer (2010) proposed a nonnegative mutual information decomposition, based on the construction of redundancy lattices, which allows separating the information that a set of variables contains about a target variable into nonnegative components interpretable as the unique information of some variables not p…
In recent years, there has been a growing interest in identifying anomalous structure within multivariate data streams. We consider the problem of detecting collective anomalies, corresponding to intervals where one or more of the data streams behaves anomalously. We first develop a test for a single collective anomaly…
Two methods extend multivariate Kelly optimization to large problem sizes.
Enhanced metrics for multiclass classification improve on existing methods.
Hyper-parameters play a major role in the learning and inference process of latent Dirichlet allocation (LDA). In order to begin the LDA latent variables learning process, these hyper-parameters values need to be pre-determined. We propose an extension for LDA that we call 'Latent Dirichlet allocation Gibbs Newton' (LD…
Multivariate binary distributions can be decomposed into products of univariate conditional distributions. Recently popular approaches have modeled these conditionals through neural networks with sophisticated weight-sharing structures. It is shown that state-of-the-art performance on several standard benchmark dataset…
Study introduces a benchmark suite for evaluating neural MI estimators on real-world unstructured datasets.
This paper investigates the ability of generative networks to convert their input noise distributions into other distributions. Firstly, we demonstrate a construction that allows ReLU networks to increase the dimensionality of their noise distribution by implementing a "space-filling" function based on iterated tent ma…
Extracting actionable insight from Electronic Health Records (EHRs) poses several challenges for traditional machine learning approaches. Patients are often missing data relative to each other; the data comes in a variety of modalities, such as multivariate time series, free text, and categorical demographic informatio…
Proposes a method to estimate time-dependent probability density functions using binary classifiers.