Novel approach for estimating joint probability densities using tensor decompositions and dictionaries.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Estimates joint probability distribution from 1-way marginals using low-rank tensors and random projections.
Method estimates joint probability density from samples using low-rank decomposition and random projections.
Develops a new framework for estimating joint probability distributions.
This work proposes a new method to estimate joint probability from pairwise marginals, reducing sample complexity.
We present a novel approach for estimating conditional probability tables, based on a joint, rather than independent, estimate of the conditional distributions belonging to the same table. We derive exact analytical expressions for the estimators and we analyse their properties both analytically and via simulation. We …
FJS method improves multinomial classification accuracy.
An important application of Lebesgue integral quadrature arXiv:1807.06007 is developed. Given two random processes, and , two generalized eigenvalue problems can be formulated and solved. In addition to obtaining two Lebesgue quadratures (for and ) from two eigenproblems, the projections of - and…
SJS model predicts label shifts in multinomial datasets.
We analyse time series of CDS spreads for a set of major US and European institutions on a pe- riod overlapping the recent financial crisis. We extend the existing methodology of ε-drawdowns to the one of joint ε-drawups, in order to estimate the conditional probabilities of abrupt co-movements among spreads. We correc…
This paper presents a Bayesian method for estimating the rank of a low-rank tensor model of joint PMF.
Estimating the joint probability mass function (PMF) of a set of random variables lies at the heart of statistical learning and signal processing. Without structural assumptions, such as modeling the variables as a Markov chain, tree, or other graphical model, joint PMF estimation is often considered mission impossible…
There has been a lot of recent interest in designing neural network models to estimate a distribution from a set of examples. We introduce a simple modification for autoencoder neural networks that yields powerful generative models. Our method masks the autoencoder's parameters to respect autoregressive constraints: ea…
Proposes methods to estimate posterior probability and propensity score functions without assuming constant propensity score.
This work extends stochastic localization to joint probability measures for data analysis.
There has recently been considerable interest in completing a low-rank matrix or tensor given only a small fraction (or few linear combinations) of its entries. Related approaches have found considerable success in the area of recommender systems, under machine learning. From a statistical estimation point of view, the…
The ability to estimate joint, conditional and marginal probability distributions over some set of variables is of great utility for many common machine learning tasks. However, estimating these distributions can be challenging, particularly in the case of data containing a mix of discrete and continuous variables. Thi…
The most direct way to express arbitrary dependencies in datasets is to estimate the joint distribution and to apply afterwards the argmax-function to obtain the mode of the corresponding conditional distribution. This method is in practice difficult, because it requires a global optimization of a complicated function,…
Missing data and noisy observations pose significant challenges for reliably predicting events from irregularly sampled multivariate time series (longitudinal) data. Imputation methods, which are typically used for completing the data prior to event prediction, lack a principled mechanism to account for the uncertainty…
The paper proposes a method to estimate joint probability from unpaired data using entropic transport kernels.
Paper estimates AI hallucinations in conditional generation tasks.
The paper bounds and identifies joint probabilities in causal inference with monotonicity assumptions.
Generative models for graphs have been typically committed to strong prior assumptions concerning the form of the modeled distributions. Moreover, the vast majority of currently available models are either only suitable for characterizing some particular network properties (such as degree distribution or clustering coe…
GFlowNets sample diverse candidates in active learning.
Hidden regular variation is a sub-model of multivariate regular variation and facilitates accurate estimation of joint tail probabilities. We generalize the model of hidden regular variation to what we call hidden domain of attraction. We exhibit examples that illustrate the need for a more general model and discuss de…
OPAA estimates probability densities using functional analysis.
A probabilistic query may not be estimable from observed data corrupted by missing values if the data are not missing at random (MAR). It is therefore of theoretical interest and practical importance to determine in principle whether a probabilistic query is estimable from missing data or not when the data are not MAR.…
Estimates copula density for complex data distributions.
Maximum mean discrepancy (MMD) has been widely adopted in domain adaptation to measure the discrepancy between the source and target domain distributions. Many existing domain adaptation approaches are based on the joint MMD, which is computed as the (weighted) sum of the marginal distribution discrepancy and the condi…
Protein contacts contain important information for protein structure and functional study, but contact prediction from sequence remains very challenging. Both evolutionary coupling (EC) analysis and supervised machine learning methods are developed to predict contacts, making use of different types of information, resp…
Classifier chains are popular and effective method to tackle a multi-label classification problem. The aim of this paper is to study the asymptotic properties of the chain model in which the conditional probabilities are of the logistic form. In particular we find conditions on the number of labels and the distribution…
Estimating a constrained relation is a fundamental problem in machine learning. Special cases are classification (the problem of estimating a map from a set of to-be-classified elements to a set of labels), clustering (the problem of estimating an equivalence relation on a set) and ranking (the problem of estimating a …
First passage models, where corporate assets undergo correlated random walks and a company defaults if its assets fall below a threshold provide an attractive framework for modeling the default process. Typical one year default correlations are small, i.e., of order a few percent, but nonetheless including correlations…
I consider two problems in machine learning and statistics: the problem of estimating the joint probability density of a collection of random variables, known as density estimation, and the problem of inferring model parameters when their likelihood is intractable, known as likelihood-free inference. The contribution o…
The paper tackles joint learning of linear systems, improving accuracy with pooled data.
Random forests is a common non-parametric regression technique which performs well for mixed-type data and irrelevant covariates, while being robust to monotonic variable transformations. Existing random forest implementations target regression or classification. We introduce the RFCDE package for fitting random forest…
The goal of online display advertising is to entice users to "convert" (i.e., take a pre-defined action such as making a purchase) after clicking on the ad. An important measure of the value of an ad is the probability of conversion. The focus of this paper is the development of a computationally efficient, accurate, a…
Proposes a new model to better handle correlation risk in credit risk calculations.
NMF and PCC linked, improving data denoising and feature stability.
DynForest predicts event probabilities from longitudinal data, handling endogenous predictors.
The article explains the probabilistic method of default probability estimation by Pluto and Tasche.
GANF uses normalizing flows to detect anomalies in multiple time series.
AR-CSM models use derivatives of univariate log-conditionals to estimate joint distributions efficiently.
Expands Hidden Markov Model to include Markov chain observations.
Proposes IPT for modeling complex joint distributions.
Develops a new framework for joint portfolio risk forecasting.
Quantum probability theory reveals hidden structure in joint probability distributions.
Econometric framework integrates heavy-tailed distributions with behavioral probability weighting for better asset pricing.