A game theory study examines gradual concessions in variable contribution games under uncertainty.
problem Gradualism in contribution games due to free rider effect.
method Stochastic game analysis of variable contribution games, extending Nerlove-Arrow model.
result Equilibrium characterized by regular control strategies leading to gradual concession.
A new distance for mixed-variable, hierarchical datasets with meta variables.
problem Heterogeneous datasets limit generalizability and performance in machine learning and optimization.
method Developed a modeling framework for mixed-variable and hierarchical domains with meta variables, and a novel distance function.
result The novel distance function allows comparison of heterogeneous datasets, improving model performance.
X-SHAP assesses multiplicative variable contributions in machine learning models.
problem Understanding multiplicative interactions in machine learning models.
method Model-agnostic method that extends SHAP to assess multiplicative contributions.
result X-SHAP proves useful in capturing multiplicative feature importance.
In this work, we propose an end-to-end block-based auto-encoder system for image compression. We introduce novel contributions to neural-network based image compression, mainly in achieving binarization simulation, variable bit rates with multiple networks, entropy-friendly representations, inference-stage code optimiz…
We develop nested automatic differentiation (AD) algorithms for exact inference and learning in integer latent variable models. Recently, Winner, Sujono, and Sheldon showed how to reduce marginalization in a class of integer latent variable models to evaluating a probability generating function which contains many leve…
Study explores K-means clustering of variables and its relation to PCA.
problem Exploring the relationship between K-means clustering of variables and PCA.
method Apply PCA to original data and K-means to transposed data, quantify variable contributions to principal components.
result Identifies how variable clusters contribute to principal components identified by PCA.
In this work, we propose a simple but effective method to interpret black-box machine learning models globally. That is, we use a compact binary tree, the interpretation tree, to explicitly represent the most important decision rules that are implicitly contained in the black-box machine learning models. This tree is l…
New method disentangles feature importance scores in machine learning.
problem Misinterpretation of feature importance scores due to interactions and dependencies.
method Derive DIP (Disentangled Importance) decomposition of feature importance scores.
result DIP decomposition uniquely separates standalone contributions from interactions and dependencies.
Proposes CLIQUE for improved local variable importance in multi-class classification.
problem Lack of methods to characterize local structure in model loss space.
method CLIQUE (Conditional Local Importance by Quantile Expectations)
result CLIQUE emphasizes locally dependent information and captures interaction behavior.
The paper proposes a method to precisely decompose confounders and estimate treatment effects.
problem Estimating treatment effects from observational data with confounder identification and balancing.
method Learning decomposed representations to identify and balance confounders and non-confounders.
result The method achieves more precise treatment effect estimation than existing methods.
The article proposes modified Gower's coefficients for handling mixed type variables in nearest neighbor methods.
problem Handling mixed type variables in nearest neighbor methods, especially imputation and statistical matching.
method Suggests modifications to the Gower's distance for interval and ratio scaled variables to address unbalanced contributions and outlier sensitivity.
result Improved distance calculations reduce the unbalanced contribution of different variable types and attenuate outlier effects.
Improves Gower's similarity for mixed-type variables with automatic weighting.
problem Handling missing values and unbalanced variable contributions in Gower's similarity for mixed-type data.
method Automatic weighting scheme minimizing differences in correlation between contributing dissimilarities and weighted Gower's dissimilarity.
result Improved performance in classification and imputation of missing values.
Develops a Bayesian non-parametric approach for signal separation with varying components.
problem Signal separation with varying components across different input locations.
method Augments Gaussian Process Latent Variable Models with weighted sums of pure component signals and incorporates priors for linear weights.
result Framework allows for non-linear variations in signals and incorporates useful priors for linear weights.
In many cases, feature selection is often more complicated than identifying a single subset of input variables that would together explain the output. There may be interactions that depend on contextual information, i.e., variables that reveal to be relevant only in some specific circumstances. In this setting, the con…
Quantitatively assessing relationships between latent variables and observed variables is important for understanding and developing generative models and representation learning. In this paper, we propose latent-observed dissimilarity (LOD) to evaluate the dissimilarity between the probabilistic characteristics of lat…
Many widely studied graphical models with latent variables lead to nontrivial constraints on the distribution of the observed variables. Inspired by the Bell inequalities in quantum mechanics, we refer to any linear inequality whose violation rules out some latent variable model as a "hidden variable test" for that mod…
New model uncovers non-Euclidean neural representations.
problem Discovering latent neural states in complex, non-Euclidean spaces.
method Manifold GPLVM for identifying latent variables and neural contributions.
result mGPLVM correctly recovers non-Euclidean latent structures in neural data.
Develops a new sampling method for gauge theories.
problem Sampling from SU(N) gauge theories. method Gauge-equivariant flows for SU(N) variables. result Constructs a class of flows respecting matrix conjugation symmetry.
Sparse models help in selecting fewer variables for efficient predictions.
problem Overfitting and high computational costs in learning models.
method Automated variable selection for sparse predictive models.
result Sparse models improve model efficiency and interpretability.
New framework for inference with LAR, explaining variable contributions and providing stopping rules.
problem LAR's lack of well-understood termination point and basic behavioral properties.
method Developed a novel framework for inference with LAR, providing new mathematical properties and stopping rules.
result LAR estimates of non-zero population correlations have independent normal distributions for inference, and zero-valued correlations have a non-normal joint distribution.
A new measure of causal influence quantifies intrinsic contributions in DAGs.
problem Quantifying intrinsic causal contributions in Directed Acyclic Graphs (DAGs).
method Recursive decomposition of node contributions, structure-preserving interventions, Shapley symmetrization.
result A measure of intrinsic causal contribution that is invariant to node relabeling.
Estimates joint causal effects using single-variable interventions on nonlinear models.
problem Estimating joint causal effects from single-variable interventions.
method Identifiability result and practical estimator for decomposing causal effects.
result Joint effects can be inferred without joint interventional data for nonlinear additive models.
Group model selection is the problem of determining a small subset of groups of predictors (e.g., the expression data of genes) that are responsible for majority of the variation in a response variable (e.g., the malignancy of a tumor). This paper focuses on group model selection in high-dimensional linear models, in w…
Enhances sensitivity analysis for correlated inputs.
problem Estimating sensitivity indices in models with correlated inputs.
method Proposes an extension of Sobol' estimator using a linear correlation model.
result Improves accuracy in variance-based sensitivity analysis.
A method uses Shapley values and Mahalanobis distances to explain multivariate outliers.
problem Explaining multivariate outlyingness in data.
method Decomposing squared Mahalanobis distance using Shapley values.
result Shapley values provide variable contributions to outlying observations.
This paper enhances LSTM neural networks for multi-variable time series data, providing interpretable insights.
problem Accurate prediction of multi-variable time series data with interpretable insights.
method Variable-wise hidden states and a mixture attention mechanism to model the generative process of the target variable.
result Enhanced prediction performance by capturing the dynamics of different variables.
Dealing with datasets of very high dimension is a major challenge in machine learning. In this paper, we consider the problem of feature selection in applications where the memory is not large enough to contain all features. In this setting, we propose a novel tree-based feature selection approach that builds a sequenc…
This paper examines from an experimental perspective random forests, the increasingly used statistical method for classification and regression problems introduced by Leo Breiman in 2001. It first aims at confirming, known but sparse, advice for using random forests and at proposing some complementary remarks for both …
Developing an explainable outlier detection method for interval-valued data using Shapley value-based approach.
problem Outlier detection in interval-valued data.
method Proposed a novel approach based on Shapley value for interval-valued data.
result Fine-grained interpretation of outliers with variable contributions.
Estimates linear model from noisy covariates and instruments using spectral regularization.
problem Estimating a linear model from many noisy covariates and instruments.
method Two-stage least squares with spectral regularization of canonical correlations.
result Upper and lower bounds on estimation error, proving optimality of the method with noisy data.
The main contribution of our paper is to give a partial classification of the quasi-exactly solvable Lie algebras of first order differential operators in three variables, and to show how this can be applied to the construction of new quasi-exactly solvable Schrödinger operators in three dimensions.
Partial covariance factorizes in path diagrams, simplifying analysis.
problem Understanding partial covariance in complex diagrams.
method Factorization of partial covariance over nodes and edges.
result Simpson's paradox cannot occur in singly-connected diagrams.
We apply belief propagation to a Bayesian bipartite graph composed of discrete independent hidden variables and discrete visible variables. The network is the Discrete counterpart of Independent Component Analysis (DICA) and it is manipulated in a factor graph form for inference and learning. A full set of simulations …
The paper studies geometric constants under modified Ricci flows with variable parameters.
problem Understanding geometric constants under variable coupling parameters in Ricci flows.
method Introduced modified Ricci flows with variable coefficients, derived evolution formulas, and proved monotonicity conditions.
result Conditions for maintaining monotonicity of geometric constants under modified Ricci flows.
Data analysis and machine learning have become an integrative part of the modern scientific methodology, offering automated procedures for the prediction of a phenomenon based on past observations, unraveling underlying patterns in data and providing insights about the problem. Yet, caution should avoid using machine l…
Estimates causal contributions of multiple causes on outcome changes.
problem Quantifying the effect of multiple causes on an outcome change.
method Develops a multiply robust estimation strategy combining regression and re-weighting methods.
result The method recovers the target parameter under partial misspecification and is consistent and asymptotically normal.
We apply a wild bootstrap method to the Lancaster three-variable interaction measure in order to detect factorisation of the joint distribution on three variables forming a stationary random process, for which the existing permutation bootstrap method fails. As in the i.i.d. case, the Lancaster test is found to outperf…
Paper introduces new approximations for lognormal sums, matching comonotonicity and moments.
problem Approximating sums of lognormal random variables accurately.
method Introduces new approximations based on weighted distribution theory, emphasizing comonotonicity and moment matching.
result Approximations perform better than classical methods, especially in the right tail of the distribution.
Ising models describe the joint probability distribution of a vector of binary feature variables. Typically, not all the variables interact with each other and one is interested in learning the presumably sparse network structure of the interacting variables. However, in the presence of latent variables, the convention…
Causal inference concerns the identification of cause-effect relationships between variables. However, often only linear combinations of variables constitute meaningful causal variables. For example, recovering the signal of a cortical source from electroencephalography requires a well-tuned combination of signals reco…
The paper introduces methods to identify key variables discriminating between two datasets.
problem Identifying variables that distinguish between two datasets.
method Introduces a mathematical notion of discriminating variables and proposes two methods for their selection.
result Proposed methods improve upon existing techniques in two-sample variable selection.
We use the score function for causal discovery, tackling challenges with hidden variables.
problem Causal discovery from observational data with hidden variables.
method Fine-tuning identifiability results, establishing conditions for inferring causal relations from the score, proposing a flexible algorithm.
result Empirical validation of the proposed algorithm for causal discovery on linear, nonlinear, and latent variable models.
Study on risk contributions of portfolios using lambda quantile risk measures.
problem No known allocation rule for non-positively homogeneous risk measures.
method Defined lambda quantiles on portfolio compositions, derived derivatives, and introduced generalized Euler contributions.
result Explicit formulae for the derivatives of lambda quantiles, showing their homogeneity properties.
The paper aims to explore the impacts of bi-demographic structure on the current account and growth. Using a SVAR modeling, we track the dynamic impacts between these underlying variables. New insights have been developed about the dynamic interrelation between population growth, current account and economic growth. Th…
Financial institutions have to allocate so-called "economic capital" in order to guarantee solvency to their clients and counter parties. Mathematically speaking, any methodology of allocating capital is a "risk measure", i.e. a function mapping random variables to the real numbers. Nowadays "value-at-risk", which is d…
New tree-structured Markov fields with Poisson marginals for counting variables.
problem Counting variables with complex dependencies.
method Tree-structured Markov random fields with Poisson marginals.
result Straightforward sampling and joint probability calculations.
An AI approach selects variables in linear models.
problem Selecting significant variables in linear regression models.
method Artificial Neural Network trained to determine variable significance based on OLS estimates.
result The AI approach outperforms traditional methods in accuracy and variable selection.
CIPNN model tackles continuous latent variables, solving intractable posterior problems.
problem Solving intractable posterior calculation for continuous latent variables.
method Derives analytical solution for posterior of continuous latent variables, proposes CIPNN and CIPAE.
result CIPNN model demonstrates great classification capability, solving problems for continuous latent variables.