Estimates joint causal effects using single-variable interventions on nonlinear models.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Estimates effects of multiple interventions with hidden confounders using single-variable interventions.
Single proxy variable helps estimate causal effects from confounders.
Kernel methods estimate causal effects with a single proxy for deterministic confounders.
In data science and machine learning, hierarchical parametric models, such as mixture models, are often used. They contain two kinds of variables: observable variables, which represent the parts of the data that can be directly measured, and latent variables, which represent the underlying processes that generate the d…
New invariants for RNA foldings and stuck links defined.
A Python package solves source duplication in single channel LVMs using spectral regularisation.
The paper identifies causal effects in latent variable models using higher-order cumulants.
Variable importance is central to scientific studies, including the social sciences and causal inference, healthcare, and other domains. However, current notions of variable importance are often tied to a specific predictive model. This is problematic: what if there were multiple well-performing predictive models, and …
New method uses latent variables to estimate treatment effects from single-arm trials.
New method identifies causal variables from multi-node interventions, expanding on previous single-node approaches.
The latent Dirichlet allocation (LDA) model is a widely-used latent variable model in machine learning for text analysis. Inference for this model typically involves a single-site collapsed Gibbs sampling step for latent variables associated with observations. The efficiency of the sampling is critical to the success o…
New method tackles OOD robustness with a single additional variable.
This study presents a new lossy image compression method that utilizes the multi-scale features of natural images. Our model consists of two networks: multi-scale lossy autoencoder and parallel multi-scale lossless coder. The multi-scale lossy autoencoder extracts the multi-scale image features to quantized variables a…
We consider the task of estimating a Gaussian graphical model in the high-dimensional setting. The graphical lasso, which involves maximizing the Gaussian log likelihood subject to an l1 penalty, is a well-studied approach for this task. We begin by introducing a surprising connection between the graphical lasso and hi…
New model generates realistic single-cell gene expression data.
Harmoniums model multiple time-to-event variables in survival analysis.
We classify the dispersive Poisson brackets with one dependent variable and two independent variables, with leading order of hydrodynamic type, up to Miura transformations. We show that, in contrast to the case of a single independent variable for which a well known triviality result exists, the Miura equivalence class…
This note is devoted to Keller-Lieb-Thirring spectral estimates for Schrödinger operators on infinite cylinders: the absolute value of the ground state level is bounded by a function of a norm of the potential. Optimal potentials with small norms are shown to depend on a single variable. The proof is a perturbation arg…
In regression modelling approach, the main step is to fit the regression line as close as possible to the target variable. In this process most algorithms try to fit all of the data in a single line and hence fitting all parts of target variable in one go. It was observed that the error between predicted and target var…
New method uses surrogate outcomes and single-record data to improve suicide risk modeling.
A linear non-Gaussian structural equation model called LiNGAM is an identifiable model for exploratory causal analysis. Previous methods estimate a causal ordering of variables and their connection strengths based on a single dataset. However, in many application domains, data are obtained under different conditions, t…
SPPCSO addresses multicollinearity in high-dimensional data, improving model stability and predictive accuracy.
Single-colored ADO-3 invariant matches Links-Gould polynomial for 5-braid closures.
We characterize and study variable importance (VIMP) and pairwise variable associations in binary regression trees. A key component involves the node mean squared error for a quantity we refer to as a maximal subtree. The theory naturally extends from single trees to ensembles of trees and applies to methods like rando…
New methods for estimating causal effects in hidden variable DAGs.
In many cases, feature selection is often more complicated than identifying a single subset of input variables that would together explain the output. There may be interactions that depend on contextual information, i.e., variables that reveal to be relevant only in some specific circumstances. In this setting, the con…
Improved GPLVM model for single-cell RNA-seq data.
Kernel testing compares cell states in single-cell data.
Proposes a method to combine datasets with missing values using Gaussian process latent variables.
We compare observed corporate cumulative default probabilities to those calculated using a stochastic model based on an extension of the work of Black and Cox and find that corporations default as if via diffusive dynamics. The model, based on a contingent-claims analysis of corporate capital structure, is easily calib…
The paper studies implicit regularization in over-parameterized models for high-dimensional data.
New model clusters cells and individuals, revealing genetic influences on cell types.
Unified Bayesian Optimisation for mixed variables improves performance.
Improved hierarchical discrete VAEs for better stability and performance.
This paper presents an algorithm for the unsupervised learning of latent variable models from unlabeled sets of data. We base our technique on spectral decomposition, providing a technique that proves to be robust both in theory and in practice. We also describe how to use this algorithm to learn the parameters of two …
We present an extension of sparse Canonical Correlation Analysis (CCA) designed for finding multiple-to-multiple linear correlations within a single set of variables. Unlike CCA, which finds correlations between two sets of data where the rows are matched exactly but the columns represent separate sets of variables, th…
In this work, we propose a simple but effective method to interpret black-box machine learning models globally. That is, we use a compact binary tree, the interpretation tree, to explicitly represent the most important decision rules that are implicitly contained in the black-box machine learning models. This tree is l…
The Lugannani-Rice formula is a saddlepoint approximation method for estimating the tail probability distribution function, which was originally studied for the sum of independent identically distributed random variables. Because of its tractability, the formula is now widely used in practical financial engineering as …
AugurOne trains single image generators without GANs using image warps.
We analyze the dynamics of an algorithm for approximate inference with large Gaussian latent variable models in a student-teacher scenario. To model nontrivial dependencies between the latent variables, we assume random covariance matrices drawn from rotation invariant ensembles. For the case of perfect data-model matc…
Observed associations in a database may be due in whole or part to variations in unrecorded (latent) variables. Identifying such variables and their causal relationships with one another is a principal goal in many scientific and practical domains. Previous work shows that, given a partition of observed variables such …
The uncertainty or the variability of the data may be treated by considering, rather than a single value for each data, the interval of values in which it may fall. This paper studies the derivation of basic description statistics for interval-valued datasets. We propose a geometrical approach in the determination of s…
A new methodology for incorporating LGD correlation effects into the Basel II risk weight functions is introduced. This methodology is based on modelling of LGD and default event with a single loss variable. The resulting formulas for capital charges are numerically compared to the current proposals by the Basel Commit…
Develops a new sampling method for gauge theories.
We derive a formula for the weight system of the multivariable Alexander polynomial using determinants, show that it obeys known relations, and satisfies some of the same relations as the single variable polynomial.
New findings show single-treatment effects are unidentifiable in factorial experiments.
Neural networks learn faster with correlated latent variables.