Improves risk and variability measures continuity and consistency.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New method detects causal relationships from noisy measurements.
Robust variable selection for high-dimensional data with missing and measurement errors.
Study tackles causal structure learning in linear models with unobserved variables and measurement error.
Derives derivatives of risk measures for various types of portfolio losses.
The aim of this paper is to introduce a risk measure that extends the Gini-type measures of risk and variability, the Extended Gini Shortfall, by taking risk aversion into consideration. Our risk measure is coherent and catches variability, an important concept for risk management. The analysis is made under the Choque…
We introduce and compare new variability measures based on risk quantiles.
In this paper we derive variability measures for the conditional probability distributions of a pair of random variables, and we study its application in the inference of causal-effect relationships. We also study the combination of the proposed measures with standard statistical measures in the the framework of the Ch…
Simple conditions for comonotonic additive risk measures from acceptance sets.
New measure assesses predictive dependence between continuous variables, capturing non-functional relationships.
Foster and Hart proposed an operational measure of riskiness for discrete random variables. We show that their defining equation has no solution for many common continuous distributions including many uniform distributions, e.g. We show how to extend consistently the definition of riskiness to continuous random variabl…
Electronic Medical Records (EMR) are a rich source of patient information, including measurements reflecting physiologic signs and administered therapies. Identifying which variables are useful in predicting clinical outcomes can be challenging. Advanced algorithms such as deep neural networks were designed to process …
A new variable importance measure for DRFs detects broader impacts on output distributions.
Proposes a new dependency function for measuring non-linear relationships.
Optimal sampling strategy improves prediction accuracy with surrogate variables under measurement constraints.
We introduce a variable importance measure to quantify the impact of individual input variables to a black box function. Our measure is based on the Shapley value from cooperative game theory. Many measures of variable importance operate by changing some predictor values with others held fixed, potentially creating unl…
A new method selects important variables for clustering from dependency networks.
Paper extends FOFC algorithm to work with mixed data types.
Study critical exponents on hyperbolic surfaces with long boundaries using Weil-Petersson measures.
We present a novel method for variable selection in regression models when covariates are measured with error. The iterative algorithm we propose, MEBoost, follows a path defined by estimating equations that correct for covariate measurement error. Via simulation, we evaluated our method and compare its performance to …
Estimating the strength of dependency between two variables is fundamental for exploratory analysis and many other applications in data mining. For example: non-linear dependencies between two continuous variables can be explored with the Maximal Information Coefficient (MIC); and categorical variables that are depende…
We propose a procedure for assigning a relevance measure to each explanatory variable in a complex predictive model. We assume that we have a training set to fit the model and a test set to check the out of sample performance. First, the individual relevance of each variable is computed by comparing the predictions in …
We proposed a new statistical dependency measure called Copula Dependency Coefficient(CDC) for two sets of variables based on copula. It is robust to outliers, easy to implement, powerful and appropriate to high-dimensional variables. These properties are important in many applications. Experimental results show that C…
Since the quasiconvex risk measures is a bigger class than the well known convex risk measures, the study of quasiconvex risk measures makes sense especially in the financial markets with volatility. In this paper, we will study the quasiconvex risk measures defined on a special space where the variable …
The paper extends Pearson correlation to multi-variables, useful for noise measurement and feature selection.
New local MDI variable importances derived from global scores match Shapley values.
Paper justifies ideal point forecasts as measurable, clarifying conditions for their existence.
When response variables are nominal and populations are cross-classified with respect to multiple polytomies, questions often arise about the degree of association of the responses with explanatory variables. When populations are known, we introduce a nominal association vector and matrix to evaluate the dependence of …
Measuring dependence between two random variables is very important, and critical in many applied areas such as variable selection, brain network analysis. However, we do not know what kind of functional relationship is between two covariates, which requires the dependence measure to be equitable. That is, it gives sim…
The equivalence between multiportfolio time consistency of a dynamic multivariate risk measure and a supermartingale property is proven. Furthermore, the dual variables under which this set-valued supermartingale is a martingale are characterized as the worst-case dual variables in the dual representation of the risk m…
New dispersion indices based on inaccuracy and divergence introduced for information measures.
Forré introduces a new conditional independence notion for mixed variables.
Global sensitivity analysis with variance-based measures suffers from several theoretical and practical limitations, since they focus only on the variance of the output and handle multivariate variables in a limited way. In this paper, we introduce a new class of sensitivity indices based on dependence measures which o…
Random Forest variable importance is improved by class balancing techniques.
The paper introduces Shapley curves for measuring variable importance in nonparametric settings.
Defines a new metric to measure importance of predictors in complex machine learning models.
The Vol-Det Conjecture relates the volume and the determinant of a hyperbolic alternating link in . We use exact computations of Mahler measures of two-variable polynomials to prove the Vol-Det Conjecture for many infinite families of alternating links. We conjecture a new lower bound for the Mahler measure of cer…
Hierarchical-CPI improves variable importance measurement for medical data.
Paper proposes a uniqueness Shapley measure to compare variable importance.
New risk measures for incomplete markets without lattice structures.
The paper derives theoretical foundations for two common machine learning variable importance measures.
Latent variable models are used to estimate variables of interest quantities which are observable only up to some measurement error. In many studies, such variables are known but not precisely quantifiable (such as "job satisfaction" in social sciences and marketing, "analytical ability" in educational testing, or "inf…
The paper addresses risk sharing and variability measures among agents with general risk preferences.
Motivation: Algorithms that discover variables which are causally related to a target may inform the design of experiments. With observational gene expression data, many methods discover causal variables by measuring each variable's degree of statistical dependence with the target using dependence measures (DMs). Howev…
A new graphical method compares stochastic variables visually.
The variability of the clusters generated by clustering techniques in the domain of latitude and longitude variables of fatal crash data are significantly unpredictable. This unpredictability, caused by the randomness of fatal crash incidents, reduces the accuracy of crash frequency (i.e., counts of fatal crashes per c…
This work presents entropic constraints from DAGs with hidden variables.
RI-based variable ranking and selection outperforms lasso in high-dimensional datasets.