Kernel embeddings of distributions and the Maximum Mean Discrepancy (MMD), the resulting distance between distributions, are useful tools for fully nonparametric two-sample testing and learning on distributions. However, it is rarely that all possible differences between samples are of interest -- discovered difference…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New metrics improve regression evaluation across different data distributions.
Distributed machine learning is an approach allowing different parties to learn a model over all data sets without disclosing their own data. In this paper, we propose a weighted distributed differential privacy (WD-DP) empirical risk minimization (ERM) method to train a model in distributed setting, considering differ…
ABROCA assesses algorithmic bias, revealing skewed distributions that inflate results.
Paper characterizes DLN distribution, its properties, and estimation methods.
This work examines the sensitivity of energy distance to mean differences compared to covariance differences.
We introduce principal differences analysis (PDA) for analyzing differences between high-dimensional distributions. The method operates by finding the projection that maximizes the Wasserstein divergence between the resulting univariate populations. Relying on the Cramer-Wold device, it requires no assumptions about th…
Paper proposes a new method to aggregate multiple sources with different label distributions.
Analyzed US firm data 1970-2019, identifying scale effects and distributional forms.
SAPAG attacks distributed learning by reconstructing true training data from gradients.
This work proposes a new method to match distributions across different spaces using cycle-consistent maps.
New algorithms improve distributional TD learning with linear approximations.
Quantile TD learning outperforms classical TD learning for value estimation.
We study distributions of realized variance (squared realized volatility) and squared implied volatility, as represented by VIX and VXO indices. We find that Generalized Beta distribution provide the best fits. These fits are much more accurate for realized variance than for squared VIX and VXO -- possibly another indi…
Distributed learning adapts to diverse devices, improving performance.
Transfer learning aims to learn robust classifiers for the target domain by leveraging knowledge from a source domain. Since the source and the target domains are usually from different distributions, existing methods mainly focus on adapting the cross-domain marginal or conditional distributions. However, in real appl…
An important goal common to domain adaptation and causal inference is to make accurate predictions when the distributions for the source (or training) domain(s) and target (or test) domain(s) differ. In many cases, these different distributions can be modeled as different contexts of a single underlying system, in whic…
Deep networks adapt to function regularity and data distribution.
Are two sets of observations drawn from the same distribution? This problem is a two-sample test. Kernel methods lead to many appealing properties. Indeed state-of-the-art approaches use the distance between kernel-based distribution representatives to derive their test statistics. Here, we show that distan…
In open set learning, a model must be able to generalize to novel classes when it encounters a sample that does not belong to any of the classes it has seen before. Open set learning poses a realistic learning scenario that is receiving growing attention. Existing studies on open set learning mainly focused on detectin…
The analysis of the USA 2001 income distribution shows that it can be described by at least two main components, which obey the generalized Tsallis statistics with different values of the q parameter. Theoretical calculations using the gas kinetics model with a distributed saving propensity factor and two ensembles rep…
Training on mixed distributions improves test performance even when components are unrelated.
The study improves model performance prediction on unseen distributions.
Functional brain networks are well described and estimated from data with Gaussian Graphical Models (GGMs), e.g. using sparse inverse covariance estimators. Comparing functional connectivity of subjects in two populations calls for comparing these estimated GGMs. Our goal is to identify differences in GGMs known to hav…
MaxRM uses random forests to minimize maximum risk across different environments.
In this paper, we present the results of Monte Carlo simulations for two popular techniques of long-range correlations detection - classical and modified rescaled range analyses. A focus is put on an effect of different distributional properties on an ability of the methods to efficiently distinguish between short and …
We propose a random walk model of asset returns where the parameters depend on market stress. Stress is measured by, e.g., the value of an implied volatility index. We show that model parameters including standard deviations and correlations can be estimated robustly and that all distributions are approximately normal.…
New algorithm for RL using mean embeddings of return distributions.
Exponential distribution is ubiquitous in the framework of multi-agent systems. An alternative approach with an economic motivation to derive the exponential distribution in the framework of iterations in the space of distributions is disclosed.
Kernel tests assess equivalence between distributions without assuming specific moments.
Domain generalization is the problem of machine learning when the training data and the test data come from different data domains. We present a simple theoretical model of learning to generalize across domains in which there is a meta-distribution over data distributions, and those data distributions may even have dif…
This work analyzes how multi-agent reinforcement learning can bridge the gap to reality in distributed multi-robot systems.
An analytic solution for asset allocation with Laplace distribution.
A new measure scales MMD to assess distribution closeness.
Recent interest in the external validity of prediction models (i.e., the problem of different train and test distributions, known as dataset shift) has produced many methods for finding predictive distributions that are invariant to dataset shifts and can be used for prediction in new, unseen environments. However, the…
In this paper we perform a statistical analysis of the high-frequency returns of the IBEX35 Madrid stock exchange index. We find that its probability distribution seems to be stable over different time scales, a stylized fact observed in many different financial time series. However, an in-depth analysis of the data us…
Analyzes financial return distributions over various time scales.
Survey of performative prediction, a machine learning setup causing distribution shifts.
Efficient bandit exploration for various distributions without distribution-specific tuning.
Different models of capital exchange among economic agents have been proposed recently trying to explain the emergence of Pareto's wealth power law distribution. One important factor to be considered is the existence of risk aversion. In this paper we study a model where agents posses different levels of risk aversion,…
We propose methods for distributed graph-based multi-task learning that are based on weighted averaging of messages from other machines. Uniform averaging or diminishing stepsize in these methods would yield consensus (single task) learning. We show how simply skewing the averaging weights or controlling the stepsize a…
Recently, Mike and Farmer have constructed a very powerful and realistic behavioral model to mimick the dynamic process of stock price formation based on the empirical regularities of order placement and cancelation in a purely order-driven market, which can successfully reproduce the whole distribution of returns, not…
Since their introduction a year ago, distributional approaches to reinforcement learning (distributional RL) have produced strong results relative to the standard approach which models expected values (expected RL). However, aside from convergence guarantees, there have been few theoretical results investigating the re…
We devise a distributional variant of gradient temporal-difference (TD) learning. Distributional reinforcement learning has been demonstrated to outperform the regular one in the recent study \citep{bellemare2017distributional}. In the policy evaluation setting, we design two new algorithms called distributional GTD2 a…
Properties of distributions of the number of trades in different intraday time intervals for five stocks traded in MICEX are studied. The dependence of the mean number of trades on the capital turnover is analyzed. Correlation analysis using factorial and moments demonstrates the multifractal nature of these dist…
We propose a \textbf{uni}fied \textbf{f}ramework for \textbf{i}mplicit \textbf{ge}nerative \textbf{m}odeling (UnifiGem) with theoretical guarantees by integrating approaches from optimal transport, numerical ODE, density-ratio (density-difference) estimation and deep neural networks. First, the problem of implicit gene…
Study on continuous sequence classification with distribution uncertainty.
Envy is a rather complex and irrational emotion. In general, it is very difficult to obtain a measure of this feeling, but in an economical context envy becomes an observable which can be measured. When various individuals compare their possessions, envy arises due to the inequality of their different allocations of co…