Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

326395126 · Jun 202019922001200920182026
48 results for Tree-based clustering

nTreeClus clusters categorical sequences using tree-based learners and k-mers.

problem Challenges in clustering categorical and sequential data.
method nTreeClus uses Tree-based Learners, k-mers, and autoregressive models for categorical time series.
result nTreeClus outperformed baseline methods in various validation metrics.

Simulation study evaluates tree-based imputation methods for multi-level data.

problem Ignoring dependencies in hierarchical data can compromise imputation accuracy.
method Chained Random Forests and Extreme Gradient Boosting (mixgb) adapted for multi-level data.
result Adapted boosting methods outperform traditional MICE for Level-1 variables at higher missingness rates.

FREEtree improves tree-based methods for correlated longitudinal data.

problem Poor performance of Random Forests in high dimensional longitudinal data with correlated features.
method FREEtree uses a piecewise random effects model and clustering with WGCNA to select features and maintain interpretability.
result FREEtree outperforms other tree-based methods in prediction and feature selection accuracy.

A new hybrid MCMC method guides MCMC with tree-based clustering for faster and more efficient inference.

problem Slow convergence of MCMC methods in posterior inference for NRM mixture models.
method Tree-guided MCMC (tgMCMC) that combines MCMC's convergence guarantees with IBHC's efficiency.
result tgMCMC provides faster convergence and better performance compared to MCMC and IBHC alone.

Efficient Maximum Inner Product Search (MIPS) is an important task that has a wide applicability in recommendation systems and classification with a large number of classes. Solutions based on locality-sensitive hashing (LSH) as well as tree-based solutions have been investigated in the recent literature, to perform ap…

2015-07-21abs ↗pdf ↗

This work improves model estimation efficiency and subgroup identification in networked systems.

problem Improving model estimation efficiency and subgroup identification in networked systems.
method A tree-based l1l_1 penalty and decentralized ADMM algorithm are used to solve the objective function in parallel.
result The approach outperforms in estimation accuracy, computation speed, and communication cost.

In this paper we analyzed dependencies in commodity markets investigating correlations of future contracts for commodities over the period 1998.09.01 - 2007.12.14. We constructed a minimal spanning tree based on the correlation matrix. The tree provides evidence for sector clusterization of investigated contracts. We a…

2008-03-27abs ↗pdf ↗

Study reviews tree-based methods and introduces new ensemble strategies.

problem Improving the efficiency and performance of tree-based machine learning models.
method Review of tree-based methods, introduction of ISLE framework, ARM model combination strategy, and modified ISLEs.
result Performance evaluation of modified ISLEs on real data sets.

The main goal of this paper is to study the geometric structures associated with the representation of tensors in subspace based formats. To do this we use a property of the so-called minimal subspaces which allows us to describe the tensor representation by means of a rooted tree. By using the tree structure and the d…

2015-05-12abs ↗pdf ↗

New method improves feature selection in tree-based models.

problem Previous feature selection methods in tree-based models lack sufficient regularization and sub-optimal performance.
method Developed a new gain penalization approach for tree-based models that allows for flexible feature-specific importance weights.
result The new method improves out-of-sample performance, especially with correlated features.

Enhanced tree-based classifiers use derivatives and geometry for better function classification.

problem Improving classification of high-dimensional time series data.
method Integrates Functional Data Analysis with tree-based ensemble techniques, leveraging derivative and geometric features.
result Significant improvements over traditional approaches in function classification.

A new tree-based model for multivariate responses interprets piecewise linear regimes.

problem Recovering piecewise multivariate linear regimes in complex data.
method Twoblock clustering trees with coskewness-based dimension reduction.
result Recovery of piecewise linear regimes in data.

The paper introduces a method to control false splits in tree-based data aggregation.

problem Identifying the correct subgroups to treat as a single entity in tree-based data.
method Introduces the 'false split rate' and proposes a multiple hypothesis testing algorithm for tree-based aggregation.
result The proposed algorithm controls the false split rate, demonstrating its effectiveness on stock volatility and taxi fare data.

T-LoHo model detects structured sparsity and smoothness on graph data.

problem Detecting structured sparsity and smoothness in graph-structured data.
method Tree-based Low-rank Horseshoe (T-LoHo) prior for multivariate parameters.
result Improves anomaly detection on road networks compared to other methods.

Tree-based models outperform deep learning on tabular data, especially for medium-sized datasets.

problem Understanding why tree-based models outperform deep learning on tabular data.
method Extensive benchmarks of tree-based and deep learning models on 45 datasets, accounting for hyperparameters.
result Tree-based models remain state-of-the-art on medium-sized tabular data, even without hyperparameter optimization.

Develops fair feature importance scores for tree-based models to interpret fairness.

problem Ensuring fairness in machine learning models, especially tree-based ones.
method Inspired by decision trees, proposes a novel fair feature importance score based on mean decrease in group bias.
result Valid interpretations of fairness for tree-based ensembles and surrogates of other ML systems.

Tree-based model averaging improves CATE estimation from diverse sites.

problem Limited sample size and privacy concerns prevent accurate personalized treatment effect estimation.
method Tree-based model averaging approach to estimate CATEs from multiple heterogeneous sites.
result Improved accuracy in estimating conditional average treatment effects (CATEs) across sites.

EmDT generates synthetic fraud data to improve detection accuracy.

problem Imbalanced datasets in fraud detection lead to poor performance on rare fraudulent transactions.
method EmDT uses UMAP clustering to identify fraudulent patterns and a Transformer denoising network to generate synthetic data.
result EmDT significantly improves classification performance compared to existing methods.

Improved Shapley Values for tree-based models, more accurate than existing methods.

problem Inaccurate Shapley Values in tree-based models leading to poor explanations.
method Introduced two new estimators exploiting tree structure, derived correct approach for categorical variables.
result More accurate Shapley Values for tree-based models, demonstrated through simulations.

The paper tackles high-dimensional function approximation using tree-based tensor formats.

problem Approximating high-dimensional functions in statistical learning.
method Empirical risk minimization over tree-based tensor formats, exploiting multilinear models and sparsity.
result Numerical stability and reliability of the proposed algorithms for learning.

A new tree-based model for varying coefficients using CGBM.

problem Modeling varying coefficients with high dimensionality and complex interactions.
method Tree-based varying coefficient model with CGBM for varying coefficients, dimension-wise early stopping, and feature importance scores.
result The model produces comparable out-of-sample loss to neural networks, demonstrating effectiveness.

This work analyzes tree-based methods from a ranking perspective, providing insights and new statistics.

problem Understanding the effectiveness of tree-based methods in finite-sample settings, especially symbolic feature selection.
method Local ranking perspective, finite-sample analysis, oracle bounds, posterior contraction results, concordant divergence statistics.
result New insights and statistics for evaluating symbolic feature mappings.

The paper introduces a method to incorporate feedback into tree-based anomaly detection to reduce false positives.

problem Difficulty in human analysts examining high-ranking anomalies due to false positives.
method Incorporates simple binary feedback into tree-based anomaly detectors, focusing on the Isolation Forest algorithm.
result Significantly improves the performance of the Isolation Forest algorithm by reducing false positives.

A new approach uses partial likelihood to improve tree-based density estimation and inference.

problem Inference on tree-based models suffers from overfitting and reduced efficiency due to data-independent partitioning.
method Proposes a partial likelihood approach to data-dependent partitioning of tree-based models.
result Significant gains in estimation accuracy and computational efficiency from adopting partial likelihood.

A new tree-based model improves uncertainty estimation in sequential optimization.

problem Improving uncertainty estimation in sequential model-based optimization.
method Proposed a new ensemble of randomized trees (BwO forest) with bagging and oversampling.
result BwO forest outperforms existing tree-based models in various optimization scenarios.

Improves tree-based models' interpretability for medical applications.

problem Lack of explainability in tree-based models.
method Developed new algorithms and tools for local and global model understanding.
result Combining local explanations reveals global model structure and identifies non-linear interactions.

We generate counterfactual explanations for tree-based boosting ensembles.

problem Understanding how tree-based models make predictions.
method Extending a method for random forests to GBDTs, accounting for tree sequential dependency and negative gradients.
result A method to generate counterfactual explanations for GBDTs.

Paper improves anomaly detection using tree-based ensembles with active learning.

problem Configuring anomaly detectors with true labels to minimize false positives.
method Develops batch and streaming active learning algorithms for tree-based ensembles.
result Significantly more anomalies discovered with active learning compared to baselines.

The paper creates a language evolution tree using word vectors from historical novels.

problem Exploring the evolution of language through historical texts.
method Constructed word vectors from novels, combined them, and used hierarchical clustering.
result Discovered a specific language evolution tree that reflects the year of the corpus.

The paper predicts travel times using tree-based ensembles.

problem Predicting travel times between urban points over short and long horizons.
method Tree-based ensemble methods trained on taxi trip records with additional features from weather and routing data.
result Adding routing data improves model performance and short-term predictions require less data.