The paper analyzes the convergence of CART under a SID condition, improving previous results.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We investigate how asymmetrizing an impurity function affects the choice of optimal node splits when growing a decision tree for binary classification. In particular, we relax the usual axioms of an impurity function and show how skewing an impurity function biases the optimal splits to isolate points of a particular c…
This paper optimizes high-dimensional oblique splits for decision trees, enhancing performance and computational efficiency.
Decision trees with binary splits are popularly constructed using Classification and Regression Trees (CART) methodology. For binary classification and regression models, this approach recursively divides the data into two near-homogenous daughter nodes according to a split point that maximizes the reduction in sum of …
New local MDI variable importances derived from global scores match Shapley values.
Bregman perspective on CART provides a unified framework for impurity measures.
New method classifies spinor orbits in dimensions up to 14.
We consider supersymmetric gauge theories with impurities in various dimensions. These systems arise in the study of intersecting branes. Unlike conventional gauge theories, the Higgs branch of an impurity theory can have compact directions. For models with eight supercharges, the Higgs branch is a hyperKahler manifold…
In this paper, we propose a novel sufficient decrease technique for stochastic variance reduced gradient descent methods such as SVRG and SAGA. In order to make sufficient decrease for stochastic optimization, we design a new sufficient decrease criterion, which yields sufficient decrease versions of stochastic varianc…
In this paper, we propose a novel sufficient decrease technique for variance reduced stochastic gradient descent methods such as SAG, SVRG and SAGA. In order to make sufficient decrease for stochastic optimization, we design a new sufficient decrease criterion, which yields sufficient decrease versions of variance redu…
Tree ensemble methods such as random forests [Breiman, 2001] are very popular to handle high-dimensional tabular data sets, notably because of their good predictive accuracy. However, when machine learning is used for decision-making problems, settling for the best predictive procedures may not be reasonable since enli…
A new method estimates uncertainty without explicit prediction models.
A hybrid impurity measure balances theoretical soundness and computational efficiency.
A persistent challenge in practical classification tasks is that labeled training sets are not always available. In particle physics, this challenge is surmounted by the use of simulations. These simulations accurately reproduce most features of data, but cannot be trusted to capture all of the complex correlations exp…
Shapelet is a discriminative subsequence of time series. An advanced shapelet-based method is to embed shapelet into accurate and fast random forest. However, it shows several limitations. First, random shapelet forest requires a large training cost for split threshold searching. Second, a single shapelet provides limi…
Machine learning methods are applied to finding the Green's function of the Anderson impurity model, a basic model system of quantum many-body condensed-matter physics. Different methods of parametrizing the Green's function are investigated; a representation in terms of Legendre polynomials is found to be superior due…
ALPE improves mid-price forecasting in HFT with real-time data.
Study connects curvature bounds to map existence and flow solutions.
Optimizes harvesting in biopharmaceutical fermentation with limited data.
Study automates feature selection and clustering for HFT stock price forecasting.
Numerical observations on martingale couplings are confirmed under certain conditions.
Tree ensembles such as Random Forests have achieved impressive empirical success across a wide variety of applications. To understand how these models make predictions, people routinely turn to feature importance measures calculated from tree ensembles. It has long been known that Mean Decrease Impurity (MDI), one of t…
Develops fair feature importance scores for tree-based models to interpret fairness.
According to the United Nations World Water Assessment Programme, every day, 2 million tons of sewage and industrial and agricultural waste are discharged into the worlds water. In order to address this pervasive issue of increasing water pollution, while ensuring that the global population has an efficient, accurate, …
A novel hybrid data-driven approach is developed for forecasting power system parameters with the goal of increasing the efficiency of short-term forecasting studies for non-stationary time-series. The proposed approach is based on mode decomposition and a feature analysis of initial retrospective data using the Hilber…
Global stability bounds for matrix frames in phase retrieval problems.
Topological entropy decreases strictly along Ricci flow near hyperbolic metrics.
Novel loss functions improve decision tree learning from noisy data.
Combines RFs and GLMs for better accuracy and interpretable feature importance.
Collaborative Trees model analyzes feature interactions and additive effects.
New conditions for weighted composition operators in group homomorphisms.
Novel evolutionary strategy solves stochastic constrained optimization problems.
Given a geometrically finite hyperbolic cone-manifold, with the cone singularity sufficiently short, we construct a one parameter family of cone-manifolds decreasing the cone angle to zero. We also control the geometry of this one parameter family via the Schwarzian derivative of the projective boundary and the length …
In this paper, we study the relation of the monotonicity of Hawking Mass and geometric flow problems. We show that along the Hamilton-DeTurck flow with bounded curvature coupled with the modified mean curvature flow, the Hawking mass of the hypersphere with a sufficiently large radius in Schwarzschild spaces is monoton…
TREGO improves EGO for global optimization of high-dimensional problems.
We present ADHM-Nahm data for instantons on the Taub-NUT space and encode these data in terms of Bow Diagrams. We study the moduli spaces of the instantons and present these spaces as finite hyperkahler quotients. As an example, we find an explicit expression for the metric on the moduli space of one SU(2) instanton. W…
The object of our investigation is a point that gives the maximum value of a potential with a strictly decreasing radially symmetric kernel. It defines a center of a body in Rm. When we choose the Riesz kernel or the Poisson kernel as the kernel, such centers are called a radial center or an illuminating center, respec…
Selective regression allows abstention to improve fairness criteria.
Many optimization algorithms converge to stationary points. When the underlying problem is nonconvex, they may get trapped at local minimizers and occasionally stagnate near saddle points. We propose the Run-and-Inspect Method, which adds an "inspect" phase to existing algorithms that helps escape from non-global stati…
Data analysis and machine learning have become an integrative part of the modern scientific methodology, offering automated procedures for the prediction of a phenomenon based on past observations, unraveling underlying patterns in data and providing insights about the problem. Yet, caution should avoid using machine l…
We study the use of knowledge distillation to compress the U-net architecture. We show that, while standard distillation is not sufficient to reliably train a compressed U-net, introducing other regularization methods, such as batch normalization and class re-weighting, in knowledge distillation significantly improves …
We propose a method for recognizing moving vehicles, using data from roadside audio sensors. This problem has applications ranging widely, from traffic analysis to surveillance. We extract a frequency signature from the audio signal using a short-time Fourier transform, and treat each time window as an individual data …
New model-independent compact representations of imaginary-time data are presented in terms of the intermediate representation (IR) of analytical continuation. This is motivated by a recent numerical finding by the authors [J. Otsuki et al., arXiv:1702.03056]. We demonstrate the efficiency of the IR through continuous-…
Recommender systems play an essential role in the modern business world. They recommend favorable items like books, movies, and search queries to users based on their past preferences. Applying similar ideas and techniques to Monte Carlo simulations of physical systems boosts their efficiency without sacrificing accura…
ATSM are widely applied for pricing of bonds and interest rate derivatives but the consistency of ATSM when the short rate, r, is unbounded from below remains essentially an open question. First, the standard approach to ATSM uses the Feynman-Kac theorem which is easily applicable only when r is bounded from below. Sec…
Decentralized Bayesian learning reduces KL-divergence exponentially.
Paper explores how unsupervised learning can be understood through linear algebra concepts.
Adam and RMSProp are two of the most influential adaptive stochastic algorithms for training deep neural networks, which have been pointed out to be divergent even in the convex setting via a few simple counterexamples. Many attempts, such as decreasing an adaptive learning rate, adopting a big batch size, incorporatin…