When response variables are nominal and populations are cross-classified with respect to multiple polytomies, questions often arise about the degree of association of the responses with explanatory variables. When populations are known, we introduce a nominal association vector and matrix to evaluate the dependence of …
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Proposes PLA-GGM for estimating variable associations with confounders.
High-dimensional, large-sample astrophysical databases of galaxy clusters, such as the Chandra Deep Field South COMBO-17 database, provide measurements on many variables for thousands of galaxies and a range of redshifts. Current understanding of galaxy formation and evolution rests sensitively on relationships between…
A concentration graph associated with a random vector is an undirected graph where each vertex corresponds to one random variable in the vector. The absence of an edge between any pair of vertices (or variables) is equivalent to full conditional independence between these two variables given all the other variables. In…
The monodromy conjecture states that every pole of the topological (or related) zeta function induces an eigenvalue of monodromy. This conjecture has already been studied a lot; however, in full generality it is proven only for zeta functions associated to a polynomial in two variables. In this article we consider zeta…
With any non necessarily orientable unpunctured marked surface (S,M) we associate a commutative algebra, called quasi-cluster algebra, equipped with a distinguished set of generators, called quasi-cluster variables, in bijection with the set of arcs and one-sided simple closed curves in (S,M). Quasi-cluster variables a…
HNet detects significant associations in mixed data types efficiently.
New sampler reduces MCMC complexity for Bayesian variable selection.
The central aim in this paper is to address variable selection questions in nonlinear and nonparametric regression. Motivated by statistical genetics, where nonlinear interactions are of particular interest, we introduce a novel and interpretable way to summarize the relative importance of predictor variables. Methodol…
In this paper we develop a general conceptual approach to the problem of existence of action-angle variables for dynamical systems, which establishes and uses the fundamental conservation property of associated torus actions: anything which is preserved by the system is also preserved by the associated torus actions. T…
The study uses DCC for financial market analysis, revealing hidden correlations.
We study a norm for structured sparsity which leads to sparse linear predictors whose supports are unions of prede ned overlapping groups of variables. We call the obtained formulation latent group Lasso, since it is based on applying the usual group Lasso penalty on a set of latent variables. A detailed analysis of th…
Bayesian deep learning accounts for input uncertainty using Errors-in-Variables models.
IEN speeds up T-Rex+GVS for fast, efficient GWAS.
New measure assesses predictive dependence between continuous variables, capturing non-functional relationships.
We characterize and study variable importance (VIMP) and pairwise variable associations in binary regression trees. A key component involves the node mean squared error for a quantity we refer to as a maximal subtree. The theory naturally extends from single trees to ensembles of trees and applies to methods like rando…
We consider the problem in precision health of grouping people into subpopulations based on their degree of vulnerability to a risk factor. These subpopulations cannot be discovered with traditional clustering techniques because their quality is evaluated with a supervised metric: the ease of modeling a response variab…
Drinfeld associator is a key tool in computing the Kontsevich integral of knots. A Drinfeld associator is a series in two non-commuting variables, satisfying highly complicated algebraic equations - hexagon and pentagon. The logarithm of a Drinfeld associator lives in the Lie algbera L generated by the symbols a,b,c mo…
Paper proposes semi-supervised learning with triplet Markov chains.
When Daan Krammer and Stephen Bigelow independently proved that braid groups are linear, they used the Lawrence-Krammer-Bigelow representation for generic values of its variables q and t. The t variable is closely connected to the traditional Garside structure of the braid group and plays a major role in Krammer's alge…
Paper proposes AR model for graph sequences.
Paper proposes mechanism learning to reverse causal inference in ML.
We construct, for a homogeneous Lagrangian of arbitrary order in two independent variables, a differential 2-form with the property that it is closed precisely when the Lagrangian is null. This is similar to the property of the `fundamental Lepage equivalent' associated with first-order Lagrangians defined on jets of s…
Observed associations in a database may be due in whole or part to variations in unrecorded (latent) variables. Identifying such variables and their causal relationships with one another is a principal goal in many scientific and practical domains. Previous work shows that, given a partition of observed variables such …
We propose an approach to the aggregation of risks which is based on estimation of simple quantities (such as covariances) associated to a vector of dependent random variables, and which avoids the use of parametric families of copulae. Our main result demonstrates that the method leads to bounds on the worst case Valu…
New particle algorithms optimize latent variable models.
A new method selects important variables for clustering from dependency networks.
A new method identifies class-specific covariates in multi-class prediction tasks.
In this study, we propose an automatic learning method for variables selection based on Lasso in epidemiology context. One of the aim of this approach is to overcome the pretreatment of experts in medicine and epidemiology on collected data. These pretreatment consist in recoding some variables and to choose some inter…
DAG models with hidden variables present many difficulties that are not present when all nodes are observed. In particular, fully observed DAG models are identified and correspond to well-defined sets ofdistributions, whereas this is not true if nodes are unobserved. Inthis paper we characterize exactly the set of dist…
Bayesian method for robust causal inference using many-dimensional instrumental variables.
This paper enhances LSTM neural networks for multi-variable time series data, providing interpretable insights.
This work introduces a novel estimation method, called LOVE, of the entries and structure of a loading matrix A in a sparse latent factor model X = AZ + E, for an observable random vector X in Rp, with correlated unobservable factors Z \in RK, with K unknown, and independent noise E. Each row of A is scaled and sparse.…
Unified geometric framework for integrability of conservative and dissipative systems.
To any completely integrable second-order system of real or complex partial differential equations in n > 1 independent variables and in one dependent variable, Mohsen Hachtroudi associated in 1937 a normal projective (Cartan) connection, and he computed its curvature. By means of a natural transfer of jet polynomials …
Method identifies causal drivers from background features.
Proposes DVC for better variable selection in non-grid data.
DEDACT breaks down feature importance into direct and associative components.
We propose a probabilistic graphical model realizing a minimal encoding of real variables dependencies based on possibly incomplete observation and an empirical cumulative distribution function per variable. The target application is a large scale partially observed system, like e.g. a traffic network, where a small pr…
New method approximates partition function of graphical models using gauge functions and polynomials.
Two categorifications are given for the arrow polynomial, an extension of the Kauffman bracket polynomial for virtual knots. The arrow polynomial extends the bracket polynomial to infinitely many variables, each variable corresponding to an integer {\it arrow number} calculated from each loop in an oriented state summa…
Efficient adjustment sets found for cost-minimized causal estimations.
This paper categorifies Chebyshev polynomials using diagrammatic algebra.
Pulling back the weight system associated with the spinor representation of the Lie algebra so(7) by the universal Vassiliev-Kontsevich invariant yields a numerical link invariant with values in formal power series. Computing some skein relations satisfied by this invariant, I derive a recursive algorithm for its evalu…
In this paper, we face the problem of simulating discrete random variables with general and varying distributions in a scalable framework, where fully parallelizable operations should be preferred. The new paradigm is inspired by the context of discrete choice models. Compared to classical algorithms, we add paralleliz…
Partial Least Squares (PLS) methods have been heavily exploited to analyse the association between two blocs of data. These powerful approaches can be applied to data sets where the number of variables is greater than the number of observations and in presence of high collinearity between variables. Different sparse ve…
We establish a correspondence between Young diagrams and differential operators of infinitely many variables. These operators form a commutative associative algebra isomorphic to the algebra of the conjugated classes of finite permutations of the set of natural numbers. The Schur functions form a complete system of com…
The study improves the perceptron's storage capacity by optimizing variable selection.