Random forests classify Pokemon names based on evolutionary status.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
CTRF combines logged data and randomized experiments for robust prediction.
An algorithm finds optimal covariates for blocking in randomized experiments.
New framework for choosing optimal proxy metrics from past experiments.
Method improves treatment effect estimation in randomized experiments.
Recent theoretical work has identified random projection as a promising dimensionality reduction technique for learning mixtures of Gausians. Here we summarize these results and illustrate them by a wide variety of experiments on synthetic and real data.
New confidence intervals improve treatment effect estimation in randomized experiments.
Estimates sample size for subgroup analysis in randomized experiments.
Paper proposes a method to use in silico experiments with foundation models to reduce sample size.
New framework minimizes interference and selection bias in network A/B testing.
Study finds no significant difference in neural network weights with quantum random numbers.
We propose an algorithm named best-scored random forest for binary classification problems. The terminology "best-scored" means to select the one with the best empirical performance out of a certain number of purely random tree candidates as each single tree in the forest. In this way, the resulting forest can be more …
Estimates treatment effects in randomized experiments with non-compliance.
Paper uses ML to improve A/B testing for complex treatment effects.
We propose a nonparametric sequential test that aims to address two practical problems pertinent to online randomized experiments: (i) how to do a hypothesis test for complex metrics; (ii) how to prevent type error inflation under continuous monitoring. The proposed test does not require knowledge of the underlying…
New method aggregates GDS analyses of randomly selected interaction models to identify important factors in screening experiments.
T-Rex selector selects variables fast and controls FDR in high-dimensional data.
New method combines randomization tests and flexible models for valid inference without splitting data.
This is a technical report which explores the estimation methodologies on hyper-parameters in Markov Random Field and Gaussian Hidden Markov Random Field. In first section, we briefly investigate a theoretical framework on Metropolis-Hastings algorithm. Next, by using MH algorithm, we simulate the data from Ising model…
Our work is a simple extension of the paper "Exploration by Random Network Distillation". More in detail, we show how to efficiently combine Intrinsic Rewards with Experience Replay in order to achieve more efficient and robust exploration (with respect to PPO/RND) and consequently better results in terms of agent perf…
Study uses three sources to evaluate language models fairly.
Cluster-DP improves differential privacy in randomized experiments by clustering data.
Efficient search methods can outperform random search on challenging tasks.
The paper develops a method for self-normalized inference in adaptive experiments.
We construct a financial "Turing test" to determine whether human subjects can differentiate between actual vs. randomized financial returns. The experiment consists of an online video-game (http://arora.ccs.neu.edu) where players are challenged to distinguish actual financial market returns from random temporal permut…
Develops statistical inference for ML-discovered heterogeneous treatment effects.
New study shows limits to classifying brain activity from randomized EEG trials.
Despite widespread interest and practical use, the theoretical properties of random forests are still not well understood. In this paper we contribute to this understanding in two ways. We present a new theoretically tractable variant of random regression forests and prove that our algorithm is consistent. We also prov…
New method for ancestral inference in branching processes with random environments.
Randomized experiments are the gold standard for evaluating the effects of changes to real-world systems. Data in these tests may be difficult to collect and outcomes may have high variance, resulting in potentially large measurement error. Bayesian optimization is a promising technique for efficiently optimizing multi…
We propose random hinge forests, a simple, efficient, and novel variant of decision forests. Importantly, random hinge forests can be readily incorporated as a general component within arbitrary computation graphs that are optimized end-to-end with stochastic gradient descent or variants thereof. We derive random hinge…
Machine learning boosts RCT efficiency by controlling type I error and improving statistical power.
We describe a model of random links based on random 4-valent maps, which can be sampled due to the work of Schaeffer. We will look at the relationship between the combinatorial information in the diagram and the hyperbolic volume. Specifically, we show that for random alternating diagrams, the expected hyperbolic volum…
Cyclic and randomized stepsizes can lead to heavier tails in SGD, improving generalization.
In network embedding, random walks play a fundamental role in preserving network structures. However, random walk based embedding methods have two limitations. First, random walk methods are fragile when the sampling frequency or the number of node sequences changes. Second, in disequilibrium networks such as highly bi…
We study the problem of treatment effect estimation in randomized experiments with high-dimensional covariate information, and show that essentially any risk-consistent regression adjustment can be used to obtain efficient estimates of the average treatment effect. Our results considerably extend the range of settings …
We consider the problem of how to assign treatment in a randomized experiment, in which the correlation among the outcomes is informed by a network available pre-intervention. Working within the potential outcome causal framework, we develop a class of models that posit such a correlation structure among the outcomes. …
It has been shown recently that graph signals with small total variation can be accurately recovered from only few samples if the sampling set satisfies a certain condition, referred to as the network nullspace property. Based on this recovery condition, we propose a sampling strategy for smooth graph signals based on …
Customer scoring models are the core of scalable direct marketing. Uplift models provide an estimate of the incremental benefit from a treatment that is used for operational decision-making. Training and monitoring of uplift models require experimental data. However, the collection of data under randomized treatment as…
Scientific and business practices are increasingly resulting in large collections of randomized experiments. Analyzed together, these collections can tell us things that individual experiments in the collection cannot. We study how to learn causal relationships between variables from the kinds of collections faced by m…
Extends RL to random stopping times, improving optimization.
The paper tackles robust design selection for online experiments under uncertain interference mechanisms.
Spectral Method is a commonly used scheme to cluster data points lying close to Union of Subspaces by first constructing a Random Geometry Graph, called Subspace Clustering. This paper establishes a theory to analyze this method. Based on this theory, we demonstrate the efficiency of Subspace Clustering in fairly broad…
A faster graph kernel using optical random features.
Proposes a new method for reliable treatment effect interval estimates.
In a wide variety of applications, including personalization, we want to measure the difference in outcome due to an intervention and thus have to deal with counterfactual inference. The feedback from a customer in any of these situations is only 'bandit feedback' - that is, a partial feedback based on whether we chose…
Quantum models generalize well with little data, challenging traditional generalization theories.
We discuss the relative merits of optimistic and randomized approaches to exploration in reinforcement learning. Optimistic approaches presented in the literature apply an optimistic boost to the value estimate at each state-action pair and select actions that are greedy with respect to the resulting optimistic value f…