Contrastive learning simplifies statistical inference for complex models.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Experimental evidence for curve ratios on genus two surfaces.
Paper proposes a method to use in silico experiments with foundation models to reduce sample size.
Introduces a new geometric method for optimal experimental design.
Improves A/B testing by detecting minor treatment effects.
Paper optimizes experimental design for estimating treatment effect.
Optimizes experimental design using synthetic controls for better outcomes.
This paper proposes a new AED framework for multi-metric experiments with fixed budget.
We define a novel class of distances between statistical multivariate distributions by modeling an optimal transport problem on their marginals with respect to a ground distance defined on their conditionals. These new distances are metrics whenever the ground distance between the marginals is a metric, generalize both…
This work builds a hedging mechanism for experimental risk.
Novel neural architecture improves Bayesian experimental design efficiency.
Develops a flexible batched experimentation framework for limited adaptivity.
Certain fibered hyperbolic 3-manifolds admit a , which can be constructed algorithmically given the stable lamination of the monodromy. These triangulations were introduced by Agol in 2011, and have been further studied by several others in the years since. We obtain exper…
Bayesian optimization simplifies bioprocess engineering experiments.
Theory predicts neural scaling exponents from language statistics.
New method uses approximate KLD for intractable likelihood models.
The study examines how experimental design choices affect machine learning model performance.
We present a system that enables rapid model experimentation for tera-scale machine learning with trillions of non-zero features, billions of training examples, and millions of parameters. Our contribution to the literature is a new method (SA L-BFGS) for changing batch L-BFGS to perform in near real-time by using stat…
Bayesian DOE accelerates experimental design with improved efficiency.
Enhances robustness in experimental design through Generalised Bayesian inference.
Consistently checking the statistical significance of experimental results is one of the mandatory methodological steps to address the so-called "reproducibility crisis" in deep reinforcement learning. In this tutorial paper, we explain how the number of random seeds relates to the probabilities of statistical errors. …
Surrogate testing techniques have been used widely to investigate the presence of dynamical nonlinearities, an essential ingredient of deterministic chaotic processes. Traditional surrogate testing subscribes to statistical hypothesis testing and investigates potential differences in discriminant statistics between the…
Study evaluates reinforcement learning algorithms for sequential experimental design.
New algorithm balances user reward and statistical inference by mixing TS with UR based on difference size.
ALMAB-DC optimizes expensive black-box experiments using active learning and distributed computing.
The experimental design problem concerns the selection of k points from a potentially large design pool of p-dimensional vectors, so as to maximize the statistical efficiency regressed on the selected k design points. Statistical efficiency is measured by optimality criteria, including A(verage), D(eterminant), T(race)…
Paper introduces a new gradient statistic to improve deep learning convergence.
This paper connects functional data analysis with machine learning techniques.
In this paper, we propose a new framework for designing fast parallel algorithms for fundamental statistical subset selection tasks that include feature selection and experimental design. Such tasks are known to be weakly submodular and are amenable to optimization via the standard greedy algorithm. Despite its desirab…
Artificial intelligence (AI) is intrinsically data-driven. It calls for the application of statistical concepts through human-machine collaboration during generation of data, development of algorithms, and evaluation of results. This paper discusses how such human-machine collaboration can be approached through the sta…
A new method connects GLM and MLE for neuroimaging analysis.
Kulldorff's (1997) seminal paper on spatial scan statistics (SSS) has led to many methods considering different regions of interest, different statistical models, and different approximations while also having numerous applications in epidemiology, environmental monitoring, and homeland security. SSS provides a way to …
Identifying statistical dependence between the features and the label is a fundamental problem in supervised learning. This paper presents a framework for estimating dependence between numerical features and a categorical label using generalized Gini distance, an energy distance in reproducing kernel Hilbert spaces (RK…
New metrics boost A/B-test power by up to 210%.
Bayesian inference learns free energy landscapes from experimental data.
We propose to investigate test statistics for testing homogeneity in reproducing kernel Hilbert spaces. Asymptotic null distributions under null hypothesis are derived, and consistency against fixed and local alternatives is assessed. Finally, experimental evidence of the performance of the proposed approach on both ar…
Automated rock fragmentation assessment using deep learning and spatial statistics.
For many important problems the quantity of interest is an unknown function of the parameters, which is a random vector with known statistics. Since the dependence of the output on this random vector is unknown, the challenge is to identify its statistics, using the minimum number of function evaluations. This problem …
Optimal probing framework for scalable network monitoring.
This study measures price risk aversion using indirect utility functions in a lab experiment.
Statistical field theory aids in understanding deep learning complexities.
Detecting a change point is a crucial task in statistics that has been recently extended to the quantum realm. A source state generator that emits a series of single photons in a default state suffers an alteration at some point and starts to emit photons in a mutated state. The problem consists in identifying the poin…
In experimental design, we are given a large collection of vectors, each with a hidden response value that we assume derives from an underlying linear model, and we wish to pick a small subset of the vectors such that querying the corresponding responses will lead to a good estimator of the model. A classical approach …
Estimating statistical models within sensor networks requires distributed algorithms, in which both data and computation are distributed across the nodes of the network. We propose a general approach for distributed learning based on combining local estimators defined by pseudo-likelihood components, encompassing a num…
We offer a novel view of AdaBoost in a statistical setting. We propose a Bayesian model for binary classification in which label noise is modeled hierarchically. Using variational inference to optimize a dynamic evidence lower bound, we derive a new boosting-like algorithm called VIBoost. We show its close connections …
We consider 1-qubit mixed quantum state estimation by adaptively updating measurements according to previously obtained outcomes and measurement settings. Updates are determined by the average-variance-optimality (A-optimality) criterion, known in the classical theory of experimental design and applied here to quantum …
Estimates effect sizes and power from a pilot experiment.
Study examines mean estimation in high dimensions with small data.