New findings on how certain functionals behave in random variable spaces.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Given a monotone convex function on the space of essentially bounded random variables with the Lebesgue property (order continuity), we consider its extension preserving the Lebesgue property to as big solid vector space of random variables as possible. We show that there exists a maximum such extension, with explicit …
Paper generalizes bipolar theorems for non-negative random variables.
We establish general versions of a variety of results for quasiconvex, lower-semicontinuous, and law-invariant functionals. Our results extend well-known results from the literature to a large class of spaces of random variables. We sometimes obtain sharper versions, even for the well-studied case of bounded random var…
CMRFs extend PGMs for topological data, capturing both conditional and marginal dependencies.
Develops a new method for nonlinear dimension reduction using random features.
AugBagg improves random forest accuracy with added noise variables.
Formula for integrating random variables on hyperbolic surfaces.
We propose a new method for input variable selection in nonlinear regression. The method is embedded into a kernel regression machine that can model general nonlinear functions, not being a priori limited to additive models. This is the first kernel-based variable selection method applicable to large datasets. It sides…
Bayesian non-parametric model selects latent dimensions automatically.
Study critical exponents on hyperbolic surfaces with long boundaries using Weil-Petersson measures.
Random walk constructs Morse functions on surfaces.
We develop a new statistical test for comparing variables with varying scales.
Bayesian optimization adapted for discrete spaces using random mappings.
Random forests are a statistical learning method widely used in many areas of scientific research because of its ability to learn complex relationships between input and output variables and also its capacity to handle high-dimensional data. However, current random forest approaches are not flexible enough to handle he…
Study on inequalities for multinomial variables.
We describe a method to perform functional operations on probability distributions of random variables. The method uses reproducing kernel Hilbert space representations of probability distributions, and it is applicable to all operations which can be applied to points drawn from the respective distributions. We refer t…
Foster and Hart proposed an operational measure of riskiness for discrete random variables. We show that their defining equation has no solution for many common continuous distributions including many uniform distributions, e.g. We show how to extend consistently the definition of riskiness to continuous random variabl…
In this paper, we study the problem of expected utility maximization of an agent who, in addition to an initial capital, receives random endowments at maturity. Contrary to previous studies, we treat as the variables of the optimization problem not only the initial capital but also the number of units of the random end…
Develops a new framework for conditional independence.
New risk measures for incomplete markets without lattice structures.
Sharp concentration results for sums of heavy-tailed random variables.
Bayesian non-linear latent variable modeling for complex data.
Paper extends stochastic dominance for compound binomial distributions.
New tree-structured Markov fields with Poisson marginals for counting variables.
The standard margin-based structured prediction commonly uses a maximum loss over all possible structured outputs. The large-margin formulation including latent variables not only results in a non-convex formulation but also increases the search space by a factor of the size of the latent space. Recent work has propose…
In this work, we consider an extension of graphical models to random graphs, trees, and other objects. To do this, many fundamental concepts for multivariate random variables (e.g., marginal variables, Gibbs distribution, Markov properties) must be extended to other mathematical objects; it turns out that this extensio…
We prove a motivic stabilization result for the cohomology of the local systems on configuration spaces of varieties over attached to character polynomials. Our approach interprets the stabilization as a probabilistic phenomenon based on the asymptotic independence of certain *motivic random variables*, an…
We consider the problem of extracting a low-dimensional, linear latent variable structure from high-dimensional random variables. Specifically, we show that under mild conditions and when this structure manifests itself as a linear space that spans the conditional means, it is possible to consistently recover the struc…
Information bottleneck (IB) is a technique for extracting information in one random variable that is relevant for predicting another random variable . IB works by encoding in a compressed "bottleneck" random variable from which can be accurately decoded. However, finding the optimal bottleneck variab…
The sectional curvature of a compact Riemannian manifold M can be seen as a random variable on the Grassmann bundle of 2-planes in TM endowed with the Fubini-Study volume density. In this article we calculate the moments of this random variable by integrating suitable local Riemannian invariants and discuss the distrib…
The paper analyzes graph Laplacians on manifolds with curvature bounds and applies to non-collapsed spaces.
New class of heavy-tailed distributions shows weighted averages dominate individual variables.
Two new methods for analyzing repeated measures data using embeddings into Reproducing Kernel Hilbert Spaces.
Consider an experiment involving a potentially small number of subjects. Some random variables are observed on each subject: a high-dimensional one called the "observed" random variable, and a one-dimensional one called the "outcome" random variable. We are interested in the dependencies between the observed random var…
We provide a general construction of time-consistent sublinear expectations on the space of continuous paths. It yields the existence of the conditional G-expectation of a Borel-measurable (rather than quasi-continuous) random variable, a generalization of the random G-expectation, and an optional sampling theorem that…
Diversification improves profits for heavy-tailed investments.
Improved DSSMs for easier interpretable latent variables.
The paper sets limits on the accuracy of macroeconomic forecasts based on statistical moments and trade volumes.
We introduce a novel mechanism to tighten the local polytope relaxation for MAP inference in Markov random fields with low state space variables. We consider a surjection of the variables to a set of hyper-variables and apply the local polytope relaxation over these hyper-variables. The state space of each individual h…
This study compares machine learning methods for high-cardinality categorical variables.
Random Forest variable importance is improved by class balancing techniques.
Variational Auto-Encoder (VAE) has been widely applied as a fundamental generative model in machine learning. For complex samples like imagery objects or scenes, however, VAE suffers from the dimensional dilemma between reconstruction precision that needs high-dimensional latent codes and probabilistic inference that f…
New random walk results on rank one symmetric spaces.
A note on extending Chernoff bound for unit interval random variables.
Random feature models approximate functions in Banach spaces efficiently.
This paper examines from an experimental perspective random forests, the increasingly used statistical method for classification and regression problems introduced by Leo Breiman in 2001. It first aims at confirming, known but sparse, advice for using random forests and at proposing some complementary remarks for both …
This paper shows how to estimate distances in latent space of random graphs using entropic OT.