Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

87174261348 · Jun 202019922001200920182026
48 results for multiple trees

This paper describes experiments, on two domains, to investigate the effect of averaging over predictions of multiple decision trees, instead of using a single tree. Other authors have pointed out theoretical and commonsense reasons for preferring the multiple tree approach. Ideally, we would like to consider predictio…

2013-03-27abs ↗pdf ↗

We define the beta diffusion tree, a random tree structure with a set of leaves that defines a collection of overlapping subsets of objects, known as a feature allocation. A generative process for the tree structure is defined in terms of particles (representing the objects) diffusing in some continuous space, analogou…

2014-08-14abs ↗pdf ↗

The paper introduces new measures to quantify variability in decision tree models due to observational multiplicity.

problem The variability in decision tree models due to observational multiplicity.
method Introduces leaf regret and structural regret to decompose observational multiplicity.
result Structural regret is the primary driver of observational multiplicity, accounting for over 15 times the variability of leaf regret in some datasets.

Recently proposed budding tree is a decision tree algorithm in which every node is part internal node and part leaf. This allows representing every decision tree in a continuous parameter space, and therefore a budding tree can be jointly trained with backpropagation, like a neural network. Even though this continuity …

2014-12-19abs ↗pdf ↗

Proposes a generalized causal tree for handling multiple treatments in uplift modeling.

problem Handling multiple treatments in uplift modeling.
method Generalizes causal tree algorithm to handle multiple discrete and continuous-valued treatments.
result Demonstrates improved performance over existing methods in experiments and real data examples.

This work addresses fairness constraints for multiple subpopulations in machine learning models.

problem Fairness constraints for multiple subpopulations in machine learning models.
method Constraining the expected outcome of subpopulations in kernel regression and decision tree regression, specifically random forests and boosted trees.
result The proposed solution does not affect the computational or memory complexity of decision trees and can be easily integrated post training.

Tree-based algorithm for functional data analysis reduces generalization error.

problem Classification and regression problems with functional data.
method Constrained convex optimization for weighted functional L2L^{2} space, multiple splitting rules, and weighted integral features.
result Reduces generalization error while maintaining interpretability.

Adaptive Bayesian model for covariate-dependent power spectra analysis.

problem Estimating complex relationships and interactions between covariates and power spectra.
method Bayesian sum of trees model with local power spectrum estimation and reversible-jump MCMC for tree modifications.
result The method can accurately recover both smooth and abrupt changes in power spectra across multiple covariates.

Enhances tree search methods in reinforcement learning for better convergence.

problem Non-contractive nature of standard tree search methods in reinforcement learning.
method Proposes a new method to back up values at the root using the optimal tree path return.
result Establishes a γhγ^h-contracting procedure leading to better convergence rates.

Proposes ReDT for interpretable, compressed, and robust decision trees.

problem Improving interpretability and performance of decision trees.
method Knowledge distillation with soft labels and multiple cross-validation.
result ReDT achieves fewer nodes than classical decision trees while maintaining good performance and interpretability.

Technology and collaboration enable dramatic increases in the size of psychological and psychiatric data collections, but finding structure in these large data sets with many collected variables is challenging. Decision tree ensembles like random forests (Strobl, Malley, and Tutz, 2009) are a useful tool for finding st…

2015-11-06abs ↗pdf ↗

LdSM builds efficient multi-label decision trees with logarithmic depth.

problem Efficiently annotate data points with relevant subsets of labels from a large label set.
method Develops LdSM algorithm for multi-label decision trees with logarithmic depth, optimizing a novel objective function for balanced splits and high class purity.
result Minimizing the proposed objective function leads to pure and balanced data splits, achieving high prediction accuracy and low prediction time.

The paper introduces a method to control false splits in tree-based data aggregation.

problem Identifying the correct subgroups to treat as a single entity in tree-based data.
method Introduces the 'false split rate' and proposes a multiple hypothesis testing algorithm for tree-based aggregation.
result The proposed algorithm controls the false split rate, demonstrating its effectiveness on stock volatility and taxi fare data.

CART can bias propensity score estimates with missing data, but multiple imputation is better.

problem Bias in propensity score estimation with CART and missing data.
method Examined CART performance with different approaches to missing data: direct CART, complete case analysis, and multiple imputation.
result Multiple imputation followed by CART outperformed direct CART with missing data.

Kauri is a novel unsupervised binary tree for clustering that outperforms existing methods.

problem Learning a tree end-to-end for clustering without labels is an open challenge.
method Greedy maximization of the kernel KMeans objective without centroids.
result Kauri often outperforms existing unsupervised clustering methods, especially with non-linear kernels.

This paper uses supervised learning to predict optimal chunk-size for parallel linear algebra operations.

problem Finding the optimal chunk-size for parallel linear algebra operations.
method The paper uses supervised learning models (logistic regression, neural networks, decision trees) to predict the optimal chunk-size for multiple linear algebra operations.
result The custom decision tree model outperforms classical decision trees and other models in predicting optimal chunk-size for linear algebra operations.

We propose a new algorithm called PLUTO for building logistic regression trees to binary response data. PLUTO can capture the nonlinear and interaction patterns in messy data by recursively partitioning the sample space. It fits a simple or a multiple linear logistic regression model in each partition. PLUTO employs th…

2014-11-25abs ↗pdf ↗

A new method creates simpler, more interpretable decision trees from complex ensembles.

problem Complex tree ensembles reduce interpretability and control over machine learning models.
method Dynamic-programming based algorithm for finding a minimum-size decision tree.
result Optimal born-again trees are simpler and more interpretable than original ensembles.

Piecewise-linear regression trees improve tree-based regression with theoretical and practical benefits.

problem Improving tree-based regression models with theoretical guarantees and practical tractability.
method Regularized piecewise-linear node-splitting criterion, LASSO-type and 2\ell_{2} regularization, variable selection procedure.
result New high-probability generalization error bounds for piecewise-linear regression trees.

New algorithms improve reinforcement learning with multi-step greedy policies.

problem Difficulty in monotonic policy improvement with soft-policy updates.
method Formulated and analyzed online and approximate algorithms using multi-step greedy operators.
result Guaranteed monotonic policy improvement with sufficiently large update stepsize.

The Farey tree helps embed rational balls and lens spaces into complex projective space.

problem Embedding rational homology balls and lens spaces into complex projective space.
method Recursive Kirby calculus argument using the Farey tree.
result Explicit constructions of embeddings of triples of rational homology balls into homotopy CP2\mathbb{CP}^2.

Three methods combine one-class classifiers with MST-CD and N-ary Trees for binary classification.

problem Binary classification with overlapping and imbalanced classes.
method Combining one-class classifiers with MST-CD and N-ary Trees to handle inconsistencies and spurious connections.
result The proposed methods are feasible and comparable to state-of-the-art algorithms.

flexBART improves BART for categorical predictors by creating flexible tree partitions.

problem Limitation of BART in handling categorical predictors with one-hot encoding.
method flexBART re-implements BART with regression trees that can assign multiple levels to both branches of a decision tree node, and proposes a new decision rule prior for spatial data.
result flexBART often yields improved predictive performance and scales better to larger datasets than existing BART implementations.

Gradient boosting adapted for multi-label and multi-output tasks.

problem Joint prediction of multiple classification or regression outputs.
method Gradient tree boosting with random output projections.
result Random projection improves adaptation to different output correlation patterns.

Max-Cut decision tree improves classification accuracy and reduces computation time.

problem Improving decision tree accuracy and efficiency for complex classification tasks.
method Alternative splitting metric (max cut) and PCA-based feature selection at each node.
result 49% improvement in accuracy with 94% reduction in CPU time on CIFAR-100 data.

The paper proposes a new algorithm to select subsets of training data for better accuracy and explainability.

problem Tackles the challenge of balancing accuracy and explainability in pattern recognition.
method Identifies multiple subsets with simple local patterns by clustering similar instances.
result The sub-setting algorithm outperformed traditional decision trees by 15% on the international stroke dataset.

Learning structured outputs with general structures is computationally challenging, except for tree-structured models. Thus we propose an efficient boosting-based algorithm AdaBoost.MRF for this task. The idea is based on the realization that a graph is a superimposition of trees. Different from most existing work, our…

2014-07-24abs ↗pdf ↗

Energy trees handle complex data structures with multiple variable types.

problem Handling intricate data structures with various types of covariates.
method Energy trees, a regression and classification model, use energy statistics to accommodate structured covariates of different types.
result Energy trees maintain statistical foundations, interpretability, and robustness to overfitting.