Proves bounds on spanning two-forests and random cut sizes.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Differentiable clustering method using perturbed spanning forests.
We study the asymptotic expansion of the determinant of the graph Laplacian associated to discretizations of a half-translation surface endowed with a flat unitary vector bundle. By doing so, over the discretizations, we relate the asymptotic expansion of the number of spanning trees and the sum of cycle-rooted spannin…
We establish an asymptotic relation between the spectrum of the discrete Laplacian associated to discretizations of a half-translation surface with a flat unitary vector bundle and the spectrum of the Friedrichs extension of the Laplacian with von Neumann boundary conditions. As an interesting byproduct of our study, w…
A result about spanning forests for graphs yields a short proof of Krebes's theorem concerning embedded tangles in links.
The paper develops a method to sparsify magnetic Laplacians using multi-type spanning forests.
We study graph estimation and density estimation in high dimensions, using a family of density estimators based on forest structured undirected graphical models. For density estimation, we do not assume the true distribution corresponds to a forest; rather, we form kernel density estimates of the bivariate and univaria…
Let X and Y be infinite graphs, such that the automorphism group of X is nonamenable, and the automorphism group of Y has an infinite orbit. We prove that there is no automorphism-invariant measure on the set of spanning trees in the direct product X times Y. This implies that the minimal spanning forest corresponding …
Stochastic networks based on random point sets as nodes have attracted considerable interest in many applications, particularly in communication networks, including wireless sensor networks, peer-to-peer networks and so on. The study of such networks generally requires the nodes to be independently and uniformly distri…
SHAKE-GNN scales GNNs for large graphs with multi-scale representations.
We present a framework for incorporating prior information into nonparametric estimation of graphical models. To avoid distributional assumptions, we restrict the graph to be a forest and build on the work of forest density estimation (FDE). We reformulate the FDE approach from a Bayesian perspective, and introduce pri…
Paper offers a simple CDS approximation formula with high accuracy.
Predicts stock volatility using Twitter data and random forests.
We propose a topological learning algorithm for the estimation of the conditional dependency structure of large sets of random variables from sparse and noisy data. The algorithm, named Maximally Filtered Clique Forest (MFCF), produces a clique forest and an associated Markov Random Field (MRF) by generalising Prim's m…
DOFEN improves DNN performance on tabular data benchmarks.
The paper develops a neural network model for SPX option pricing.
Classification outperforms regression in portfolio construction, yielding higher Sharpe ratios.
Paper studies gradient fields from discrete Morse functions for watershed-cut computation.
Paper introduces MVS to detect non-Markovian observations in reinforcement learning.
iMondrian forest combines isolation forest and Mondrian forest for better anomaly detection.
Ensembles of randomized decision trees, usually referred to as random forests, are widely used for classification and regression tasks in machine learning and statistics. Random forests achieve competitive predictive performance and are computationally efficient to train and test, making them excellent candidates for r…
New random forest method provides optimal rates and confidence bands.
RFpredInterval package builds prediction intervals for random forests and boosted forests.
We propose random hinge forests, a simple, efficient, and novel variant of decision forests. Importantly, random hinge forests can be readily incorporated as a general component within arbitrary computation graphs that are optimized end-to-end with stochastic gradient descent or variants thereof. We derive random hinge…
Improved random forest proximities capture data geometry.
Deep forests enhance expressiveness exponentially with depth, not width or tree size.
This paper improves forest pruning to balance accuracy and interpretability.
This paper is a comment on the survey paper by Biau and Scornet (2016) about random forests. We focus on the problem of quantifying the impact of each ingredient of random forests on their performance. We show that such a quantification is possible for a simple pure forest , leading to conclusions that could apply more…
Forest-guided smoothing uses random forest outputs for interpretable local smoothers.
In this paper we propose using the principle of boosting to reduce the bias of a random forest prediction in the regression setting. From the original random forest fit we extract the residuals and then fit another random forest to these residuals. We call the sum of these two random forests a \textit{one-step boosted …
A framework for the generation of bridge-specific fragility utilizing the capabilities of machine learning and stripe-based approach is presented in this paper. The proposed methodology using random forests helps to generate or update fragility curves for a new set of input parameters with less computational effort and…
By seeking the narrowest prediction intervals (PIs) that satisfy the specified coverage probability requirements, the recently proposed quality-based PI learning principle can extract high-quality PIs that better summarize the predictive certainty in regression tasks, and has been widely applied to solve many practical…
Improves time series classification with forest proximities.
Random forests reduce bias and variance, especially in low SNR settings.
New random forest variants achieve optimal performance in high dimensions.
Enhances random forest consistency and introduces DMRF for improved performance.
Weighted SVM (or fuzzy SVM) is the most widely used SVM variant owning its effectiveness to the use of instance weights. Proper selection of the instance weights can lead to increased generalization performance. In this work, we extend the span error bound theory to weighted SVM and we introduce effective hyperparamete…
Online random forests improve Q-learning performance in specific gym environments.
New method learns representations for decision forests using input perturbation.
Improved random forest models enhance machine learning predictions.
Introduced by Breiman, Random Forests are widely used classification and regression algorithms. While being initially designed as batch algorithms, several variants have been proposed to handle online learning. One particular instance of such forests is the \emph{Mondrian Forest}, whose trees are built using the so-cal…
We investigate the effect of the proportional hazards assumption on prognostic and predictive models of the survival time of patients suffering from amyotrophic lateral sclerosis (ALS). We theoretically compare the underlying model formulations of several variants of survival forests and implementations thereof, includ…
Many scientific and engineering challenges -- ranging from personalized medicine to customized marketing recommendations -- require an understanding of treatment effect heterogeneity. In this paper, we develop a non-parametric causal forest for estimating heterogeneous treatment effects that extends Breiman's widely us…
The study uses machine learning to predict cryptocurrency market trends and design profitable trading strategies.
FAST-DAD distills complex ensemble models into faster, more accurate individual models.
Decision forests, including Random Forests and Gradient Boosting Trees, have recently demonstrated state-of-the-art performance in a variety of machine learning settings. Decision forests are typically ensembles of axis-aligned decision trees; that is, trees that split only along feature dimensions. In contrast, many r…
A fast method for finding counterfactual explanations for decision forests.
We improve random forest consistency and performance with DMRF, a new variant.