We propose a tree regularization framework, which enables many tree models to perform feature selection efficiently. The key idea of the regularization framework is to penalize selecting a new feature for splitting when its gain (e.g. information gain) is similar to the features used in previous splits. The regularizat…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Proposes regional tree regularization for interpretable deep models.
Tree regularization makes deep models interpretable by approximating them with simple decision trees.
Piecewise-linear regression trees improve tree-based regression with theoretical and practical benefits.
A new algorithm, Regular Tree Search, tackles non-convex simulation optimization problems.
New method for comparing different mass measures on tree structures using entropy partial transport.
Proposes new attribution methods for trees with regularization.
HS improves tree-based models' accuracy and interpretability without changing their structure.
We study the set of critical exponents of discrete groups acting on regular trees. We prove that for every real number between and , there is a discrete subgroup acting without inversion on a -regular tree whose critical exponent is equal to . Explicit construction of edge-index…
One obstacle that so far prevents the introduction of machine learning models primarily in critical areas is the lack of explainability. In this work, a practicable approach of gaining explainability of deep artificial neural networks (NN) using an interpretable surrogate model based on decision trees is presented. Sim…
Enhances GBDT robustness with one-hot encoding and regularization.
Dynamic CBDT improves treatment effect estimation in clinical data.
This study argues for pruning trees in random forests to improve performance in low signal-to-noise scenarios.
Paper proposes a novel SVM method for creating survival trees.
This paper examines a novel gradient boosting framework for regression. We regularize gradient boosted trees by introducing subsampling and employ a modified shrinkage algorithm so that at every boosting stage the estimate is given by an average of trees. The resulting algorithm, titled Boulevard, is shown to converge …
New framework optimizes classification trees with logistic loss and regularization.
Regularizes decision trees to reduce inference time by up to 4x with minimal accuracy loss.
New method improves feature selection in tree-based models.
The lack of interpretability remains a key barrier to the adoption of deep models in many applications. In this work, we explicitly regularize deep models so human users might step through the process behind their predictions in little time. Specifically, we train deep time-series models so their class-probability pred…
We consider certain groups of tree automorphisms as so-called diffeological groups. The notion of diffeology, due to Souriau, allows to endow non-manifold topological spaces, such as regular trees that we look at, with a kind of a differentiable structure that in many ways is close to that of a smooth manifold; a suita…
Let G be a finitely presented group. Scott and Swarup have constructed a canonical splitting of G which encloses all almost invariant sets over virtually polycyclic subgroups of a given length. We give an alternative construction of this regular neighbourhood, by showing that it is the tree of cylinders of a JSJ splitt…
Symmetric TSP is structurally equivalent to a constrained Group Steiner Tree Problem.
Paper explains how tree ensembles improve predictions by smoothing and regulating smoothness.
Quasi-Sturmian words, which are infinite words with factor complexity eventually share many properties with Sturmian words. In this paper, we study the quasi-Sturmian colorings on regular trees. There are two different types, bounded and unbounded, of quasi-Sturmian colorings. We obtain an induction algorithm sim…
Phylogenetic tree inference using deep DNA sequencing is reshaping our understanding of rapidly evolving systems, such as the within-host battle between viruses and the immune system. Densely sampled phylogenetic trees can contain special features, including "sampled ancestors" in which we sequence a genotype along wit…
Tree ensembles such as random forests and boosted trees are accurate but difficult to understand, debug and deploy. In this work, we provide the inTrees (interpretable trees) framework that extracts, measures, prunes and selects rules from a tree ensemble, and calculates frequent variable interactions. An rule-based le…
New method aggregates nodes in sparse graphical models.
Paper shows MCTS approximates policy optimization, proposing an improved variant.
We study the problem of identifying the source of a diffusion spreading over a regular tree. When the degree of each node is at least three, we show that it is possible to construct confidence sets for the diffusion source with size independent of the number of infected nodes. Our estimators are motivated by analogous …
Paper introduces XBART for nonlinear regression, outperforming XGBoost.
Sharp bounds for spanning tree entropy in planar lattices.
Variable selection for high-dimensional linear models has received a lot of attention lately, mostly in the context of l1-regularization. Part of the attraction is the variable selection effect: parsimonious models are obtained, which are very suitable for interpretation. In terms of predictive power, however, these re…
This paper approximates 1-Wasserstein distance using tree-based embedding.
Decision trees are an extremely popular machine learning technique. Unfortunately, overfitting in decision trees still remains an open issue that sometimes prevents achieving good performance. In this work, we present a novel approach for the construction of decision trees that avoids the overfitting by design, without…
New algorithm improves convergence of gradient boosting trees.
Groups acting on product trees are boundary rigid.
A new gradient tree boosting framework reduces variance and accelerates performance.
TF Boosted Trees (TFBT) is a new open-sourced frame-work for the distributed training of gradient boosted trees. It is based on TensorFlow, and its distinguishing features include a novel architecture, automatic loss differentiation, layer-by-layer boosting that results in smaller ensembles and faster prediction, princ…
We introduce a novel boosting algorithm called `KTBoost' which combines kernel boosting and tree boosting. In each boosting iteration, the algorithm adds either a regression tree or reproducing kernel Hilbert space (RKHS) regression function to the ensemble of base learners. Intuitively, the idea is that discontinuous …
SVR-Tree improves classification trees for imbalanced and sparse data.
Invariant detects triple points in sphere immersions.
We consider the problem of learning a forest of nonlinear decision rules with general loss functions. The standard methods employ boosted decision trees such as Adaboost for exponential loss and Friedman's gradient boosting for general loss. In contrast to these traditional boosting algorithms that treat a tree learner…
ControlBurn selects few features from tree ensembles for better model interpretability.
Estimates tree-based density from random vectors.
Universal inequalities for Laplacian eigenvalues on discrete groups.
We present a construction, called the limit of a tree system of spaces (or, less formally, a tree of spaces). The construction is designed to produce compact metric spaces that resemble fractals, out of more regular spaces, such as closed manifolds, compact polyhedra, compact Menger manifolds, etc. Such spaces are pote…
Multivariate boosted trees improve forecasting and control by capturing correlated predictions.
This study investigates self-supervised learning with Wasserstein distance on tree structures.