Binary choice forests model customer choices in retailing.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Develops RF-GLS for binary geospatial data.
Customer behavior is often assumed to follow weak rationality, which implies that adding a product to an assortment will not increase the choice probability of another product in that assortment. However, an increasing amount of research has revealed that customers are not necessarily rational when making decisions. In…
Random forest models predict CLABSI risk in hospital admissions, with static models performing similarly to dynamic ones.
In this paper I present an extended implementation of the Random ferns algorithm contained in the R package rFerns. It differs from the original by the ability of consuming categorical and numerical attributes instead of only binary ones. Also, instead of using simple attribute subspace ensemble it employs bagging and …
We introduce canonical correlation forests (CCFs), a new decision tree ensemble method for classification and regression. Individual canonical correlation trees are binary decision trees with hyperplane splits based on local canonical correlation coefficients calculated during training. Unlike axis-aligned alternatives…
Online BSP-Forest improves space partitioning for large-scale classification and regression.
Enhanced Random Forests outperform XGBoost across binary classification datasets.
We characterize and study variable importance (VIMP) and pairwise variable associations in binary regression trees. A key component involves the node mean squared error for a quantity we refer to as a maximal subtree. The theory naturally extends from single trees to ensembles of trees and applies to methods like rando…
Paper improves anomaly detection by using non-uniform random choices in isolation forests.
Study binary choice with asymmetric loss, offering simple solutions.
Excellent ranking power along with well calibrated probability estimates are needed in many classification tasks. In this paper, we introduce a technique, Calibrated Boosting-Forest that captures both. This novel technique is an ensemble of gradient boosting machines that can support both continuous and binary labels. …
Approximate Bayesian computation (ABC) methods provide an elaborate approach to Bayesian inference on complex models, including model choice. Both theoretical arguments and simulation experiments indicate, however, that model posterior probabilities may be poorly evaluated by standard ABC techniques. We propose a novel…
Fuzzy Forests reduces feature space in high-dimensional survey data.
A new deep learning model for tabular data improves accuracy over GBDT.
The Binary Space Partitioning~(BSP)-Tree process is proposed to produce flexible 2-D partition structures which are originally used as a Bayesian nonparametric prior for relational modelling. It can hardly be applied to other learning tasks such as regression trees because extending the BSP-Tree process to a higher dim…
Signature Isolation Forest removes constraints from FIF by using rough path theory's signature transform.
This research sets limits on how complex multi-class learning problems can be.
Enhances preference learning by incorporating response times into binary choices.
We propose an algorithm named best-scored random forest for binary classification problems. The terminology "best-scored" means to select the one with the best empirical performance out of a certain number of purely random tree candidates as each single tree in the forest. In this way, the resulting forest can be more …
Space partitioning methods such as random forests and the Mondrian process are powerful machine learning methods for multi-dimensional and relational data, and are based on recursively cutting a domain. The flexibility of these methods is often limited by the requirement that the cuts be axis aligned. The Ostomachion p…
Probabilistic-driven classification techniques extend the role of traditional approaches that output labels (usually integer numbers) only. Such techniques are more fruitful when dealing with problems where one is not interested in recognition/identification only, but also into monitoring the behavior of consumers and/…
This paper proposes a novel type of random forests called a denoising random forests that are robust against noises contained in test samples. Such noise-corrupted samples cause serious damage to the estimation performances of random forests, since unexpected child nodes are often selected and the leaf nodes that the i…
The paper examines the stability of binary choice models using Gini index and scoring indicators.
We have applied a little-known data transformation to subsets of the Surveillance, Epidemiology, and End Results (SEER) publically available data of the National Cancer Institute (NCI) to make it suitable input to standard machine learning classifiers. This transformation properly treats the right-censored data in the …
A new method identifies class-specific covariates in multi-class prediction tasks.
AF improves classification models by adaptively weighting trees.
In this work, we propose a novel node splitting method for regression trees and incorporate it into the regression forest framework. Unlike traditional binary splitting, where the splitting rule is selected from a predefined set of binary splitting rules via trial-and-error, the proposed node splitting method first fin…
Random forests are a powerful method for non-parametric regression, but are limited in their ability to fit smooth signals, and can show poor predictive performance in the presence of strong, smooth effects. Taking the perspective of random forests as an adaptive kernel method, we pair the forest kernel with a local li…
We introduce a novel scheme to train binary convolutional neural networks (CNNs) -- CNNs with weights and activations constrained to {-1,+1} at run-time. It has been known that using binary weights and activations drastically reduce memory size and accesses, and can replace arithmetic operations with more efficient bit…
A novel unsupervised outlier detection method using Randomized PCA Forest.
We investigate a class of binary choice models with social interactions. We propose a unifying perspective that integrates economic models using a utility function and psychological models using an impact function. A general approach for analyzing the equilibrium structure of these models within mean-field approximatio…
SOAR generates rules for both positive and negative classes in binary classification.
Machine learning struggles to predict binary options movements due to randomness.
In statistical learning theory, convex surrogates of the 0-1 loss are highly preferred because of the computational and theoretical virtues that convexity brings in. This is of more importance if we consider smooth surrogates as witnessed by the fact that the smoothness is further beneficial both computationally- by at…
This document is an invited chapter covering the specificities of ABC model choice, intended for the incoming Handbook of ABC by Sisson, Fan, and Beaumont (2017). Beyond exposing the potential pitfalls of ABC based posterior probabilities, the review emphasizes mostly the solution proposed by Pudlo et al. (2016) on the…
Bayesian Causal Forest models estimate treatment effects with noncompliance.
Optimal weighted random forests improve prediction accuracy.
Random Forest kernels improve performance in various regression and survival tasks.
This study argues for pruning trees in random forests to improve performance in low signal-to-noise scenarios.
In short, our experiments suggest that yes, on average, rotation forest is better than the most common alternatives when all the attributes are real-valued. Rotation forest is a tree based ensemble that performs transforms on subsets of attributes prior to constructing each tree. We present an empirical comparison of c…
Random Planted Forest interprets tree-based models by keeping some splits, leading to more interpretable predictions.
Proposes modifications to model-based forests for HTE estimation in observational data.
Decision forests learn to model text by evaluating categorical-set conditions.
Random Machines improves SVM performance with free kernel choice.
Paper proposes a new method for supervised manifold learning using random forest proximities.
The study optimizes machine learning classifiers for variable stars using CRTS data.
This paper improves deep forest models with soft routing and topology learning.