Consider an experiment involving a potentially small number of subjects. Some random variables are observed on each subject: a high-dimensional one called the "observed" random variable, and a one-dimensional one called the "outcome" random variable. We are interested in the dependencies between the observed random var…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The paper sets limits on the accuracy of macroeconomic forecasts based on statistical moments and trade volumes.
Proposes a new dependency function for measuring non-linear relationships.
We propose an approach to the aggregation of risks which is based on estimation of simple quantities (such as covariances) associated to a vector of dependent random variables, and which avoids the use of parametric families of copulae. Our main result demonstrates that the method leads to bounds on the worst case Valu…
In many cases, feature selection is often more complicated than identifying a single subset of input variables that would together explain the output. There may be interactions that depend on contextual information, i.e., variables that reveal to be relevant only in some specific circumstances. In this setting, the con…
Investigates VaR behavior for sums of one-sided random variables, showing impossibilities and conditions for super-additivity.
Diversification improves profits for heavy-tailed investments.
CMRFs extend PGMs for topological data, capturing both conditional and marginal dependencies.
SHAFF estimates Shapley effects efficiently even with dependent variables.
New bounds on continuous random variables' right-tail probabilities.
Estimating the strength of dependency between two variables is fundamental for exploratory analysis and many other applications in data mining. For example: non-linear dependencies between two continuous variables can be explored with the Maximal Information Coefficient (MIC); and categorical variables that are depende…
Measuring dependence between two random variables is very important, and critical in many applied areas such as variable selection, brain network analysis. However, we do not know what kind of functional relationship is between two covariates, which requires the dependence measure to be equitable. That is, it gives sim…
DynForest R package predicts outcomes with time-dependent predictors.
The study examines how market trade randomness influences price and return volatility.
We introduce the Randomized Dependence Coefficient (RDC), a measure of non-linear dependence between random variables of arbitrary dimension based on the Hirschfeld-Gebelein-Rényi Maximum Correlation Coefficient. RDC is defined in terms of correlation of random non-linear copula projections; it is invariant with respec…
New class of heavy-tailed distributions shows weighted averages dominate individual variables.
Unexpectedly, weighted Pareto variables are stochastically dominant.
Random feature matrices' singular values concentrate near their full expectation in high dimensions.
New criteria distinguish cause from effect in data, overcoming statistical limitations.
We prove semi-empirical concentration inequalities for random variables which are given as possibly nonlinear functions of independent random variables. These inequalities describe concentration of random variable in terms of the data/distribution-dependent Efron-Stein (ES) estimate of its variance and they do not requ…
We provide a theoretical foundation for non-parametric estimation of functions of random variables using kernel mean embeddings. We show that for any continuous function , consistent estimators of the mean embedding of a random variable lead to consistent estimators of the mean embedding of . For Matérn ke…
New measure assesses predictive dependence between continuous variables, capturing non-functional relationships.
New tree-structured Markov fields with Poisson marginals for counting variables.
In this manuscript, we analytically and numerically study statistical properties of an heteroskedastic process based on the celebrated ARCH generator of random variables whose variance is defined by a memory of -exponencial, form (). Specifically, we inspect the self-correlation function o…
The paper examines bounds for stop-loss payoffs using transformed random variables.
We introduce an approximate search algorithm for fast maximum a posteriori probability estimation in probabilistic programs, which we call Bayesian ascent Monte Carlo (BaMC). Probabilistic programs represent probabilistic models with varying number of mutually dependent finite, countable, and continuous random variable…
Bayesian networks, and especially their structures, are powerful tools for representing conditional independencies and dependencies between random variables. In applications where related variables form a priori known groups, chosen to represent different "views" to or aspects of the same entities, one may be more inte…
Random forest hyperparameters affect variable selection in omics studies.
New method improves variable importance in random forests.
New AI-block models for clustering high-dimensional variables based on maxima of random processes.
Two ANOVA-based algorithms boost random Fourier feature models for function approximation.
We revisit the Kolmogorov-Smirnov and Cramér-von Mises goodness-of-fit (GoF) tests and propose a generalisation to identically distributed, but dependent univariate random variables. We show that the dependence leads to a reduction of the "effective" number of independent observations. The generalised GoF tests are not…
Machine learning provides algorithms that can learn from data and make inferences or predictions on data. Bayesian networks are a class of graphical models that allow to represent a collection of random variables and their condititional dependencies by directed acyclic graphs. In this paper, an inference algorithm for …
Information theory provides ideas for conceptualising information and measuring relationships between objects. It has found wide application in the sciences, but economics and finance have made surprisingly little use of it. We show that time series data can usefully be studied as information -- by noting the relations…
RaSE screens variables via random subspaces, identifying joint effects.
The reparameterization trick enables optimizing large scale stochastic computation graphs via gradient descent. The essence of the trick is to refactor each stochastic node into a differentiable function of its parameters and a random variable with fixed distribution. After refactoring, the gradients of the loss propag…
Simple conditions for comonotonic additive risk measures from acceptance sets.
Fermat-Torricelli points help assess investment risks by smoothing series data.
Optimized variable orderings improve autoregressive model performance.
DIET tests conditional independence using marginal dependence measures of residual information.
We consider the setting where a collection of time series, modeled as random processes, evolve in a causal manner, and one is interested in learning the graph governing the relationships of these processes. A special case of wide interest and applicability is the setting where the noise is Gaussian and relationships ar…
We present two alternative ways to apply PAC-Bayesian analysis to sequences of dependent random variables. The first is based on a new lemma that enables to bound expectations of convex functions of certain dependent random variables by expectations of the same functions of independent Bernoulli random variables. This …
The paper provides bounds on the CDF of a variable under nonstationary conditions.
We present a methodology for clustering N objects which are described by multivariate time series, i.e. several sequences of real-valued random variables. This clustering methodology leverages copulas which are distributions encoding the dependence structure between several random variables. To take fully into account …
In this paper, we study the stochastic combinatorial multi-armed bandit (CMAB) framework that allows a general nonlinear reward function, whose expected value may not depend only on the means of the input random variables but possibly on the entire distributions of these variables. Our framework enables a much larger c…
Tree ensemble methods such as random forests [Breiman, 2001] are very popular to handle high-dimensional tabular data sets, notably because of their good predictive accuracy. However, when machine learning is used for decision-making problems, settling for the best predictive procedures may not be reasonable since enli…
Gaussian copulas are widely used in the industry to correlate two random variables when there is no prior knowledge about the co-dependence between them. The perturbed Gaussian copula approach allows introducing the skew information of both random variables into the co-dependence structure. The analytical expression of…
Better signal detection in undersampled data using joint and cross covariances.