Targeted Learning uses robust statistics for reproducible research.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This study connects Gaussian processes and RKHS, bridging two machine learning communities.
What makes a paper independently reproducible? Debates on reproducibility center around intuition or assumptions but lack empirical results. Our field focuses on releasing code, which is important, but is not sufficient for determining reproducibility. We take the first step toward a quantifiable answer by manually att…
We describe convolutional networks using harmonic functions.
Earlier we proposed the stochastic point process model, which reproduces a variety of self-affine time series exhibiting power spectral density S(f) scaling as power of the frequency f and derived a stochastic differential equation with the same long range memory properties. Here we present a stochastic differential eq…
AdaStop improves statistical testing for Deep RL algorithm comparisons.
Teaches reproducible research to medical students and postgrads.
The paper develops methods to handle missing data using regularized M-estimation in reproducing kernel Hilbert space.
ERICA assesses reproducibility in cluster analysis.
Paper reproduces a kernel-based scan B-statistic for online change-point detection.
Framework generates precise synthetic populations for scalable modeling.
We derive a system of stochastic differential equations simulating the dynamics of the three agent groups with herding interaction. Proposed approach can be valuable in the modeling of the complex socio-economic systems with similar composition of the agents. We demonstrate how the sophisticated statistical features of…
The paper develops a uniform function estimator in RKHS for regression.
New budget quantifies drift in closed-loop learning, improving reproducibility.
The role of kernels is central to machine learning. Motivated by the importance of power-law distributions in statistical modeling, in this paper, we propose the notion of power-law kernels to investigate power-laws in learning problem. We propose two power-law kernels by generalizing Gaussian and Laplacian kernels. Th…
HECT tests climate model outputs for reproducibility.
We propose a model of fractal point process driven by the nonlinear stochastic differential equation. The model is adjusted to the empirical data of trading activity in financial markets. This reproduces the probability distribution function and power spectral density of trading activity observed in the stock markets. …
Proposes incorporating noise sources in machine learning evaluation for more reliable conclusions.
Subsampling reduces computational cost in supervised learning in reproducing kernel Hilbert spaces.
As reinforcement learning (RL) achieves more success in solving complex tasks, more care is needed to ensure that RL research is reproducible and that algorithms herein can be compared easily and fairly with minimal bias. RL results are, however, notoriously hard to reproduce due to the algorithms' intrinsic variance, …
L-ARC improves model fairness by localizing risk guarantees.
This study assesses the reproducibility of 1H-MRS scans across different vendors and sessions.
Kernelized cumulants improve statistical analysis in high-dimensional spaces.
One of the main challenges in the parametrization of geological models is the ability to capture complex geological structures often observed in the subsurface. In recent years, generative adversarial networks (GAN) were proposed as an efficient method for the generation and parametrization of complex data, showing sta…
Study on kernel methods in large-scale machine learning problems.
The study maps ML quality dimensions to fairness, enhancing the QF4SA framework.
Time reversal invariance can be summarized as follows: no difference can be measured if a sequence of events is run forward or backward in time. Because price time series are dominated by a randomness that hides possible structures and orders, the existence of time reversal invariance requires care to be investigated. …
Auto-regressive conditionally heteroskedastic (ARCH) family models are still used, by practitioners in business and economic policy making, as a conditional volatility forecasting models. Furthermore ARCH models still are attracting an interest of the researchers. In this contribution we consider the well known GARCH(1…
This paper improves computational efficiency in kernel ridge regression under covariate shift.
We introduce kernel nonparametric tests for Lancaster three-variable interaction and for total independence, using embeddings of signed measures into a reproducing kernel Hilbert space. The resulting test statistics are straightforward to compute, and are used in powerful interaction tests, which are consistent against…
We introduce a novel data-driven order reduction method for nonlinear control systems, drawing on recent progress in machine learning and statistical dimensionality reduction. The method rests on the assumption that the nonlinear system behaves linearly when lifted into a high (or infinite) dimensional feature space wh…
We present eigenvalue decay estimates of integral operators associated with compositional dot-product kernels. The estimates improve on previous ones established for power series kernels on spheres. This allows us to obtain the volumes of balls in the corresponding reproducing kernel Hilbert spaces. We discuss the cons…
Develops robust persistence diagrams using kernel methods.
CPME embeds counterfactual outcomes in RKHS for flexible policy evaluation.
Study assesses consistency and reproducibility of LLMs in finance and accounting tasks.
Consistently checking the statistical significance of experimental results is one of the mandatory methodological steps to address the so-called "reproducibility crisis" in deep reinforcement learning. In this tutorial paper, we explain how the number of random seeds relates to the probabilities of statistical errors. …
We propose to investigate test statistics for testing homogeneity in reproducing kernel Hilbert spaces. Asymptotic null distributions under null hypothesis are derived, and consistency against fixed and local alternatives is assessed. Finally, experimental evidence of the performance of the proposed approach on both ar…
The distribution of money is analysed in connection with the Boltzmann distribution of energy in the degenerate states of molecules. Plots of the population density of income distribution for various countries are well reproduced by a Gamma function, confirming the validity of the statistical distribution at equilibriu…
Signals consisting of a sequence of pulses show that inherent origin of the 1/f noise is a Brownian fluctuation of the average interevent time between subsequent pulses of the pulse sequence. In this paper we generalize the model of interevent time to reproduce a variety of self-affine time series exhibiting power spec…
The paper revisits and improves on a Bayesian relevance vector machine method for small sample sizes.
Deep learning has become increasingly popular in both supervised and unsupervised machine learning thanks to its outstanding empirical performance. However, because of their intrinsic complexity, most deep learning methods are largely treated as black box tools with little interpretability. Even though recent attempts …
A new depth measure for non-convex data supports, faster than halfspace depth.
Additive models play an important role in semiparametric statistics. This paper gives learning rates for regularized kernel based methods for additive models. These learning rates compare favourably in particular in high dimensions to recent results on optimal learning rates for purely nonparametric regularized kernel …
New method uses machine learning to improve statistical inference.
Research develops a generic method for evaluating trading platform components.
Artificial intelligence (AI) is intrinsically data-driven. It calls for the application of statistical concepts through human-machine collaboration during generation of data, development of algorithms, and evaluation of results. This paper discusses how such human-machine collaboration can be approached through the sta…
New method learns policies from offline data using operator models.
Survey of determinism issues in financial AI systems.