Improved outlier detection in hierarchical Gaussian Processes using Wasserstein-2 kernels.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We study the problem of empirical minimization for variance-type functionals over functional classes. Sharp non-asymptotic bounds for the excess variance are derived under mild conditions. In particular, it is shown that under some restrictions imposed on the functional class fast convergence rates can be achieved incl…
The method of covariate adjustment is often used for estimation of population average treatment effects in observational studies. Graphical rules for determining all valid covariate adjustment sets from an assumed causal graphical model are well known. Restricting attention to causal linear models, a recent article der…
Bayesian model captures mean and variance of response variables.
Study optimal adjustment sets for causal policies with hidden variables.
When randomized ensembles such as bagging or random forests are used for binary classification, the prediction error of the ensemble tends to decrease and stabilize as the number of classifiers increases. However, the precise relationship between prediction error and ensemble size is unknown in practice. In the standar…
The latest generation of volatility derivatives goes beyond variance and volatility swaps and probes our ability to price realized variance and sojourn times along bridges for the underlying stock price process. In this paper, we give an operator algebraic treatment of this problem based on Dyson expansions and moment …
A novel k-NN method estimates conditional mean and variance efficiently.
Time-subordinated Brownian motion models improve financial market stochastic distribution.
Existing strategies for finite-armed stochastic bandits mostly depend on a parameter of scale that must be known in advance. Sometimes this is in the form of a bound on the payoffs, or the knowledge of a variance or subgaussian parameter. The notable exceptions are the analysis of Gaussian bandits with unknown mean and…
This paper tackles the problem of selecting among several linear estimators in non-parametric regression; this includes model selection for linear regression, the choice of a regularization parameter in kernel ridge regression, spline smoothing or locally weighted regression, and the choice of a kernel in multiple kern…
Flexible model captures commodity skews with maturity effects.
Subagging improves regression tree performance, especially with many splits.
The paper explains why estimating a history-dependent policy can reduce MSE in reinforcement learning.
This paper solves the intractability barrier in non-parametric information geometry by introducing a novel framework.
We propose generalized random forests, a method for non-parametric statistical estimation based on random forests (Breiman, 2001) that can be used to fit any quantity of interest identified as the solution to a set of local moment equations. Following the literature on local maximum likelihood estimation, our method co…
Enforces physical constraints in GP regression models.
Statistical physics approaches can be used to derive accurate predictions for the performance of inference methods learning from potentially noisy data, as quantified by the learning curve defined as the average error versus number of training examples. We analyse a challenging problem in the area of non-parametric inf…
Accounting for the non-normality of asset returns remains challenging in robust portfolio optimization. In this article, we tackle this problem by assessing the risk of the portfolio through the "amount of randomness" conveyed by its returns. We achieve this by using an objective function that relies on the exponential…
A new realized conditional autoregressive Value-at-Risk (VaR) framework is proposed, through incorporating a measurement equation into the original quantile regression model. The framework is further extended by employing various Expected Shortfall (ES) components, to jointly estimate and forecast VaR and ES. The measu…
Proposes a general method to derive regret bounds for multi-armed bandit algorithms.
Efficient adjustment sets found for cost-minimized causal estimations.
We propose a Bayesian non-parametric approach for modeling the distribution of multiple returns. In particular, we use an asymmetric dynamic conditional correlation (ADCC) model to estimate the time-varying correlations of financial returns where the individual volatilities are driven by GJR-GARCH models. The ADCC-GJR-…
Proposes method for eliciting non-parametric joint priors using normalizing flows.
Paper introduces new importance metrics for machine learning models, linking them to CATE.
This work creates a CS for non-negative heavy-tailed data with bounded mean.
Adversarially robust machine learning has received much recent attention. However, prior attacks and defenses for non-parametric classifiers have been developed in an ad-hoc or classifier-specific basis. In this work, we take a holistic look at adversarial examples for non-parametric classifiers, including nearest neig…
The paper develops asymptotic theory for QRF variable importance, revealing a bias-variance trade-off.
Study shows double descent curve in high-dimensional linear regression with random projections.
The weighted k-nearest neighbors algorithm is one of the most fundamental non-parametric methods in pattern recognition and machine learning. The question of setting the optimal number of neighbors as well as the optimal weights has received much attention throughout the years, nevertheless this problem seems to have r…
NPOD algorithm improves efficiency in estimating pharmacokinetic parameters.
This paper improves bandwidth selectors for SPBNs to enhance their performance.
The study optimizes sampling in complex systems with probabilistic response distributions.
Modeling structure in complex networks using Bayesian non-parametrics makes it possible to specify flexible model structures and infer the adequate model complexity from the observed data. This paper provides a gentle introduction to non-parametric Bayesian modeling of complex networks: Using an infinite mixture model …
We consider the optimization of a quadratic objective function whose gradients are only accessible through a stochastic oracle that returns the gradient at any given point plus a zero-mean finite variance random error. We present the first algorithm that achieves jointly the optimal prediction error rates for least-squ…
This study examines when non-parametric methods are robust to adversarial examples.
Synthetic augmentation improves financial machine learning performance in variance-dominant regimes.
Paper introduces a new power-dominance axis in estimator design.
The latent feature relational model (LFRM) is a generative model for graph-structured data to learn a binary vector representation for each node in the graph. The binary vector denotes the node's membership in one or more communities. At its core, the LFRM miller2009nonparametric is an overlapping stochastic blockmodel…
Simpler GNNs with low-rank non-parametric aggregators perform well on graph benchmarks.
Dirichlet Process(DP) is a Bayesian non-parametric prior for infinite mixture modeling, where the number of mixture components grows with the number of data items. The Hierarchical Dirichlet Process (HDP), is an extension of DP for grouped data, often used for non-parametric topic modeling, where each group is a mixtur…
We propose a representation of Gaussian processes (GPs) based on powers of the integral operator defined by a kernel function, we call these stochastic processes integral Gaussian processes (IGPs). Sample paths from IGPs are functions contained within the reproducing kernel Hilbert space (RKHS) defined by the kernel fu…
Study provides guarantees for kernel clustering under non-parametric mixtures.
Develops flexible non-parametric ACFs using B-spline kernels.
A novel MCMC method clusters data faster and more accurately.
One of the fundamental problems in supervised classification and in machine learning in general, is the modelling of non-parametric invariances that exist in data. Most prior art has focused on enforcing priors in the form of invariances to parametric nuisance transformations that are expected to be present in data. Le…
Study evaluates policies in partially observable environments without full model specification.
Non-parametric estimators improve quickest changepoint detection under irregular sequence lengths.