Reduces identity testing of reversible Markov chains to simpler symmetric chain tests.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Identity testing for reversible Markov chains without symmetry assumption.
New sampling and identity-testing methods for mixtures of distributions that don't satisfy approximate tensorization of entropy.
The study examines property testing and estimation under non-identically distributed samples, finding necessary and sufficient sample complexities.
We investigate the problems of identity and closeness testing over a discrete population from random samples. Our goal is to develop efficient testers while guaranteeing Differential Privacy to the individuals of the population. We describe an approach that yields sample-efficient differentially private testers for the…
In this work we present novel differentially private identity (goodness-of-fit) testers for natural and widely studied classes of multivariate product distributions: Gaussians in with known covariance and product distributions over . Our testers have improved sample complexity compared to …
We exhibit an efficient procedure for testing, based on a single long state sequence, whether an unknown Markov chain is identical to or -far from a given reference chain. We obtain nearly matching (up to logarithmic factors) upper and lower sample complexity bounds for our notion of distance, which is bas…
It is often stated in papers tackling the task of inferring Bayesian network structures from data that there are these two distinct approaches: (i) Apply conditional independence tests when testing for the presence or otherwise of edges; (ii) Search the model space using a scoring metric. Here I argue that for complete…
Paper proposes a method to encrypt faces while maintaining visual similarity.
We consider testing and learning problems on causal Bayesian networks as defined by Pearl (Pearl, 2009). Given a causal Bayesian network on a graph with discrete variables and bounded in-degree and bounded `confounded components', we show that interventions on an unknown causal Bayesian ne…
Robust covariance testing requires significantly more samples in contaminated data.
Market portfolio decomposed into body and tail legs
The paper analyzes ridge regression with random features for non-identically distributed data.
We derive a new discrepancy statistic for measuring differences between two probability distributions based on combining Stein's identity with the reproducing kernel Hilbert space theory. We apply our result to test how well a probabilistic model fits a set of observations, and derive a new class of powerful goodness-o…
New method identifies whether equity return predictability is due to magnitude shrinkage or directional reversal.
Real-world machine learning applications often have complex test metrics, and may have training and test data that are not identically distributed. Motivated by known connections between complex test metrics and cost-weighted learning, we propose addressing these issues by using a weighted loss function with a standard…
Study decomposes market portfolio into body and tail legs, revealing systematic differences.
We evaluate the impact of probabilistically-constructed digital identity data collected from Sep. to Dec. 2017 (approx.), in the context of Lookalike-targeted campaigns. The backbone of this study is a large set of probabilistically-constructed "identities", represented as small bags of cookies and mobile ad identifier…
We propose a new setting for testing properties of distributions while receiving samples from several distributions, but few samples per distribution. Given samples from distributions, , we design testers for the following problems: (1) Uniformity Testing: Testing whether all the 's are …
We study distribution testing with communication and memory constraints in the following computational models: (1) The {\em one-pass streaming model} where the goal is to minimize the sample complexity of the protocol subject to a memory constraint, and (2) A {\em distributed model} where the data samples reside at mul…
The paper finds a pervasive and severe bias in accounting semi-identity models.
A single algebraic identity unifies information-theoretic variational results.
A machine learning configuration refers to a combination of preprocessor, learner, and hyperparameters. Given a set of configurations and a large dataset randomly split into training and testing set, we study how to efficiently select the best configuration with approximately the highest testing accuracy when trained f…
This work develops a non-parametric test for relational independence in non-i.i.d. data.
We study the problem of independence testing given independent and identically distributed pairs taking values in a -finite, separable measure space. Defining a natural measure of dependence as the squared -distance between a joint density and the product of its marginals, we first show that there is…
A new test detects noise in graph data, useful for forecasting.
We revisit the Kolmogorov-Smirnov and Cramér-von Mises goodness-of-fit (GoF) tests and propose a generalisation to identically distributed, but dependent univariate random variables. We show that the dependence leads to a reduction of the "effective" number of independent observations. The generalised GoF tests are not…
A new permutation method improves two-sample testing power.
Kernel test evaluates dynamical system data streams.
Nowadays, machine learning methods have been widely used in stock prediction. Traditional approaches assume an identical data distribution, under which a learned model on the training data is fixed and applied directly in the test data. Although such assumption has made traditional machine learning techniques succeed i…
TAME learns new tasks without knowing them, outperforming existing methods.
Sequential tests for two-sample and independence testing using betting strategies.
A new test assesses text similarity between two groups of documents.
Methods for combining predictions from different models in a supervised learning setting must somehow estimate/predict the quality of a model's predictions at unknown future inputs. Many of these methods (often implicitly) make the assumption that the test inputs are identical to the training inputs, which is seldom re…
Hypothesis testing in the linear regression model is a fundamental statistical problem. We consider linear regression in the high-dimensional regime where the number of parameters exceeds the number of samples (). In order to make informative inference, we assume that the model is approximately sparse, that is th…
New test detects differences in heterogeneous datasets.
Improved covariance matrix estimation for portfolio optimization with guaranteed PSD and controlled conditioning.
Sample efficiency and scalability to a large number of agents are two important goals for multi-agent reinforcement learning systems. Recent works got us closer to those goals, addressing non-stationarity of the environment from a single agent's perspective by utilizing a deep net critic which depends on all observatio…
We present the expected values from p-value hacking as a choice of the minimum p-value among independents tests, which can be considerably lower than the "true" p-value, even with a single trial, owing to the extreme skewness of the meta-distribution. We first present an exact probability distribution (meta-distrib…
Develops non-parametric tests for group symmetry in data.
Study improves sample complexity for distinguishing continuous distributions and causal relationships.
We tackle the Multi-task Batch Reinforcement Learning problem. Given multiple datasets collected from different tasks, we train a multi-task policy to perform well in unseen tasks sampled from the same distribution. The task identities of the unseen tasks are not provided. To perform well, the policy must infer the tas…
The Trouvé group from image analysis consists of the flows at a fixed time of all time-dependent vectors fields of a given regularity . For a multitude of regularity classes , we prove that the Trouvé group coincides wi…
Unified score and distance-based GoF tests for model adequacy.
Many-to-Many VTN improves voice conversion across multiple speakers.
A new method simulates a lazy version of a Markov chain for empirical inference.
Near-optimal tests and confidence sequences for non-parametric data.
Unified benchmarks assess data poisoning and backdoor attacks.