Bregman perspective on CART provides a unified framework for impurity measures.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A hybrid impurity measure balances theoretical soundness and computational efficiency.
We investigate how asymmetrizing an impurity function affects the choice of optimal node splits when growing a decision tree for binary classification. In particular, we relax the usual axioms of an impurity function and show how skewing an impurity function biases the optimal splits to isolate points of a particular c…
Decision trees with binary splits are popularly constructed using Classification and Regression Trees (CART) methodology. For binary classification and regression models, this approach recursively divides the data into two near-homogenous daughter nodes according to a split point that maximizes the reduction in sum of …
New method classifies spinor orbits in dimensions up to 14.
New local MDI variable importances derived from global scores match Shapley values.
We consider supersymmetric gauge theories with impurities in various dimensions. These systems arise in the study of intersecting branes. Unlike conventional gauge theories, the Higgs branch of an impurity theory can have compact directions. For models with eight supercharges, the Higgs branch is a hyperKahler manifold…
A new method estimates uncertainty without explicit prediction models.
A persistent challenge in practical classification tasks is that labeled training sets are not always available. In particle physics, this challenge is surmounted by the use of simulations. These simulations accurately reproduce most features of data, but cannot be trusted to capture all of the complex correlations exp…
Machine learning methods are applied to finding the Green's function of the Anderson impurity model, a basic model system of quantum many-body condensed-matter physics. Different methods of parametrizing the Green's function are investigated; a representation in terms of Legendre polynomials is found to be superior due…
The paper analyzes the convergence of CART under a SID condition, improving previous results.
Develops RF-GLS for binary geospatial data.
Plasma tomography consists in reconstructing the 2D radiation profile in a poloidal cross-section of a fusion device, based on line-integrated measurements along several lines of sight. The reconstruction process is computationally intensive and, in practice, only a few reconstructions are usually computed per pulse. I…
Optimizes harvesting in biopharmaceutical fermentation with limited data.
According to the United Nations World Water Assessment Programme, every day, 2 million tons of sewage and industrial and agricultural waste are discharged into the worlds water. In order to address this pervasive issue of increasing water pollution, while ensuring that the global population has an efficient, accurate, …
Global stability bounds for matrix frames in phase retrieval problems.
Tree ensembles such as Random Forests have achieved impressive empirical success across a wide variety of applications. To understand how these models make predictions, people routinely turn to feature importance measures calculated from tree ensembles. It has long been known that Mean Decrease Impurity (MDI), one of t…
Novel loss functions improve decision tree learning from noisy data.
New algorithm reduces discrimination in predictions.
We present ADHM-Nahm data for instantons on the Taub-NUT space and encode these data in terms of Bow Diagrams. We study the moduli spaces of the instantons and present these spaces as finite hyperkahler quotients. As an example, we find an explicit expression for the metric on the moduli space of one SU(2) instanton. W…
This paper optimizes high-dimensional oblique splits for decision trees, enhancing performance and computational efficiency.
Tree ensemble methods such as random forests [Breiman, 2001] are very popular to handle high-dimensional tabular data sets, notably because of their good predictive accuracy. However, when machine learning is used for decision-making problems, settling for the best predictive procedures may not be reasonable since enli…
The paper analyzes ESS metrics and their connections to entropy families.
Combines RFs and GLMs for better accuracy and interpretable feature importance.
New model-independent compact representations of imaginary-time data are presented in terms of the intermediate representation (IR) of analytical continuation. This is motivated by a recent numerical finding by the authors [J. Otsuki et al., arXiv:1702.03056]. We demonstrate the efficiency of the IR through continuous-…
Recommender systems play an essential role in the modern business world. They recommend favorable items like books, movies, and search queries to users based on their past preferences. Applying similar ideas and techniques to Monte Carlo simulations of physical systems boosts their efficiency without sacrificing accura…
Paper explores how unsupervised learning can be understood through linear algebra concepts.
Improved decision tree learning guarantees for complex functions.
Study uses random forest to detect unlawful insider trading in financial data.
Novel approach for creating interpretable classifiers using bilevel optimization of split-rules in NLDTs.
We seek decision rules for prediction-time cost reduction, where complete data is available for training, but during prediction-time, each feature can only be acquired for an additional cost. We propose a novel random forest algorithm to minimize prediction error for a user-specified {\it average} feature acquisition b…
Data analysis and machine learning have become an integrative part of the modern scientific methodology, offering automated procedures for the prediction of a phenomenon based on past observations, unraveling underlying patterns in data and providing insights about the problem. Yet, caution should avoid using machine l…
Despite the impressive performance of random forests (RF), its theoretical properties have not been thoroughly understood. In this paper, we propose a novel RF framework, dubbed multinomial random forest (MRF), to analyze the \emph{consistency} and \emph{privacy-preservation}. Instead of deterministic greedy split rule…
Collaborative Trees model analyzes feature interactions and additive effects.
Regularizes decision trees to reduce inference time by up to 4x with minimal accuracy loss.
Review and benchmark 58 feature selection methods for ML applications.
Shapelet is a discriminative subsequence of time series. An advanced shapelet-based method is to embed shapelet into accurate and fast random forest. However, it shows several limitations. First, random shapelet forest requires a large training cost for split threshold searching. Second, a single shapelet provides limi…
Rapidly growing product lines and services require a finer-granularity forecast that considers geographic locales. However the open question remains, how to assess the quality of a spatio-temporal forecast? In this manuscript we introduce a metric to evaluate spatio-temporal forecasts. This metric is based on an Opti- …
We study both the continuous model and the discrete model of the integer quantum Hall effect on the hyperbolic plane in the presence of disorder, extending the results of an earlier paper [CHMM]. Here we model impurities, that is we consider the effect of a random or almost periodic potential as opposed to just periodi…
ALPE improves mid-price forecasting in HFT with real-time data.
Semi-automatic data annotation helps experts label unlabeled samples based on feature space projection.
How to obtain a model with good interpretability and performance has always been an important research topic. In this paper, we propose rectified decision trees (ReDT), a knowledge distillation based decision trees rectification with high interpretability, small model size, and empirical soundness. Specifically, we ext…
FedForest adapts RF for federated learning, improving performance and efficiency.
The paper studies statistical properties of CART regression trees.
Supervised learning algorithms are nowadays successfully scaling up to datasets that are very large in volume, leveraging the potential of in-memory cluster-computing Big Data frameworks. Still, massive datasets with a number of large-domain categorical features are a difficult challenge for any classifier. Most off-th…
Study automates feature selection and clustering for HFT stock price forecasting.
A novel hybrid data-driven approach is developed for forecasting power system parameters with the goal of increasing the efficiency of short-term forecasting studies for non-stationary time-series. The proposed approach is based on mode decomposition and a feature analysis of initial retrospective data using the Hilber…
Kernel methods are popular in clustering due to their generality and discriminating power. However, we show that many kernel clustering criteria have density biases theoretically explaining some practically significant artifacts empirically observed in the past. For example, we provide conditions and formally prove the…