Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

92183275366 · Jun 202019922001200920182026
48 results for common variability

We extend common entropy concept and propose algorithms to distinguish causation from correlation.

problem Discovering the simplest latent variable for conditional independence of observed variables.
method Renyi common entropy, iterative algorithm, constraint-based methods modification.
result Improved constraint-based methods for causal inference in small samples.

New method identifies common cause in causal insufficiency, revealing complex phase transitions.

problem Identifying common cause in causal insufficiency with observed joint probability.
method Generalized maximum likelihood method, closely related to maximum entropy principle.
result Identifies consistent common cause that aligns with the common cause principle.

D-GCCA improves multi-view data analysis by separating common and distinctive components.

problem Analyzing multi-view high-dimensional data with latent factors.
method Decomposes each view's data matrix into common and distinctive sources with orthogonality constraints.
result Consistent estimators with good performance and efficient computation.

Algorithm estimates common mean from Gaussian variables with unknown variances.

problem Estimating common mean from Gaussian variables with different unknown variances.
method Intuitive and efficient algorithm using Subset-of-Signals model as benchmark.
result Improved estimation error by polynomial factors compared to previous work.

Novel method identifies structural differences between networks using structural equation models.

problem Identifying structural differences between networks characterized by structural equation models.
method Reparameterization and algorithm design with calibration and construction stages to identify differential structures.
result Our method outperformed independently constructed networks on synthetic data and demonstrated applicability on a real data set.

CLOUD method detects causal relationships in various data types without latent variable assumptions.

problem Detecting causal relationships in the presence of unobserved common causes.
method CLOUD method using Normalized Maximum Likelihood (NML) Code for various data types (discrete, mixed, continuous).
result CLOUD method is more effective than existing methods in inferring causal relationships.

Paper presents a method to estimate mixed-variable distributions.

problem Estimating joint, conditional, and marginal distributions from mixed data.
method Graph representation of data, eigenvector equations for distribution estimation.
result Method successfully estimates distributions for various machine learning tasks.

A game theory study examines gradual concessions in variable contribution games under uncertainty.

problem Gradualism in contribution games due to free rider effect.
method Stochastic game analysis of variable contribution games, extending Nerlove-Arrow model.
result Equilibrium characterized by regular control strategies leading to gradual concession.

The paper develops a method to model high-dimensional data with many variables and weak signals.

problem Modeling high-dimensional dependent data with many explanatory variables and low signal-to-noise ratio.
method Penalized regression for high-dimensional data, factor modeling of residuals, high-dimensional white noise testing, projected Principal Component Analysis.
result Established asymptotic properties of the proposed method for high-dimensional data.

In recent years, there is a growing interest in learning Bayesian networks with continuous variables. Learning the structure of such networks is a computationally expensive procedure, which limits most applications to parameter learning. This problem is even more acute when learning networks with hidden variables. We p…

2012-07-11abs ↗pdf ↗

Binary sequence correlation estimation fails but trinary data succeeds.

problem Estimating correlation in binary sequences generated by thresholding a hidden continuous sequence.
method Formal analysis and numerical experiments on likelihood maximization and discretization effects.
result Consistent estimation of correlation is possible with trinary data but not with binary data.

This work restricts hidden cardinality in causal models to infer causal relations.

problem Causal relations between variables with a common unobserved cause cannot be directly inferred.
method Derive inequality constraints from d-separation in causal models with known cardinalities of unobserved variables.
result Inference of causal relations is possible with additional assumptions about cardinalities.

Study proposes a method for identifying important variables in multi-class classification problems.

problem Lack of studies on variable selection in nonparametric classification models, especially for multi-class problems.
method Sparse non-parametric density estimation approach for identifying high impacts variables.
result Proposed method identifies important variables for each class in multi-class classification problems.

We propose a mixture of latent trait models with common slope parameters (MCLT) for model-based clustering of high-dimensional binary data, a data type for which few established methods exist. Recent work on clustering of binary data, based on a dd-dimensional Gaussian latent variable, is extended by incorporating com…

2014-04-11abs ↗pdf ↗

Bayesian method models binary response and covariates for two groups, estimating causal relationships.

problem Estimating causal relationships between binary response and covariates in observational data.
method Gaussian DAG-probit model with MCMC sampling for posterior distribution estimation.
result Validated method on simulated and real datasets, showing value of grouping variable in causality.

Stochastic neural networks with infinite width become deterministic, reducing training variance.

problem Understanding how stochasticity in neural networks affects learning and regularization.
method Theoretical analysis of stochastic neural networks with infinite width.
result As the width of an optimized stochastic neural network increases, its predictive variance on the training set decreases to zero.

Observed associations in a database may be due in whole or part to variations in unrecorded (latent) variables. Identifying such variables and their causal relationships with one another is a principal goal in many scientific and practical domains. Previous work shows that, given a partition of observed variables such …

2012-10-19abs ↗pdf ↗

New MCMC methods use auxiliary variables to sample from intractable distributions.

problem Sampling from distributions with unknown normalizing constants.
method Unified Markov chain Monte Carlo framework with auxiliary variables.
result New algorithms outperform existing methods on synthetic and real datasets.

The paper derives theoretical foundations for two common machine learning variable importance measures.

problem Understanding variable importance in machine learning problems.
method The paper derives closed-form expressions for Permute-and-Predict (PaP) and Leave-One-Covariate-Out (LOCO) methods.
result Theoretical derivations explain the behavior of PaP and LOCO under collinearity, linking them to coefficients and predictor variability.

Paper extends stochastic dominance for compound binomial distributions.

problem Stochastic dominance for infinite-mean random variables.
method Investigates properties and inclusion relationships of distribution classes, extends results to compound binomial distributions.
result Establishes necessary and sufficient conditions for first-order stochastic dominance preservation.

In a variety of disciplines such as social sciences, psychology, medicine and economics, the recorded data are considered to be noisy measurements of latent variables connected by some causal structure. This corresponds to a family of graphical models known as the structural equation model with latent variables. While …

2014-08-09abs ↗pdf ↗

In a variety of disciplines such as social sciences, psychology, medicine and economics, the recorded data are considered to be noisy measurements of latent variables connected by some causal structure. This corresponds to a family of graphical models known as the structural equation model with latent variables. While …

2010-02-25abs ↗pdf ↗

Whitening, or sphering, is a common preprocessing step in statistical analysis to transform random variables to orthogonality. However, due to rotational freedom there are infinitely many possible whitening procedures. Consequently, there is a diverse range of sphering methods in use, for example based on principal com…

2015-12-02abs ↗pdf ↗

The paper compares traditional regression with modern neural network methods for financial hedging and risk compression.

problem Finding optimal hedge ratios and managing portfolio risk using traditional regression methods has limitations.
method The paper introduces regularization techniques and common factor analyses using neural networks to improve upon regression methods.
result Neural network methods provide better performance in hedge ratio estimation and risk compression compared to traditional regression.

New method reparameterizes discrete variables to reduce gradient variance.

problem Low variance gradient estimation for discrete variables in neural networks.
method Marginalizing out the variable of interest to bypass discontinuity, resulting in a new reparameterization trick.
result The new reparameterization reduces gradient variance significantly, theoretically not larger than likelihood-ratio method.

Foster and Hart proposed an operational measure of riskiness for discrete random variables. We show that their defining equation has no solution for many common continuous distributions including many uniform distributions, e.g. We show how to extend consistently the definition of riskiness to continuous random variabl…

2013-01-08abs ↗pdf ↗

Discond-VAE separates continuous and discrete factors in data.

problem Separating shared and class-specific variations in real-world data.
method Introduces private and public latent variables to represent continuous and discrete factors, respectively.
result Discond-VAE successfully disentangles class-dependent continuous factors from discrete factors.

Paper relaxes identifiability conditions for causal models with latent variables.

problem Challenges in identifying causal graphical models with latent variables.
method Proposes a double triangular graphical condition for nonparametric measurement models with binary latent variables.
result Guarantees identifiability of the entire causal graphical model under relaxed conditions.

ARSM estimator improves gradient backpropagation for categorical variables.

problem Improving gradient backpropagation through categorical variables.
method ARSM combines variable augmentation, REINFORCE, Rao-Blackwellization, and variable swapping.
result ARSM outperforms existing estimators and provides variance reduction methods.