Improved portfolio optimization method yields better risk-adjusted returns.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A common problem machine learning developers are faced with is overfitting, that is, fitting a pipeline too closely to the training data that the performance degrades for unseen data. Automated machine learning aims to free (or at least ease) the developer from the burden of pipeline creation, but this overfitting prob…
In this tutorial paper, we first define mean squared error, variance, covariance, and bias of both random variables and classification/predictor models. Then, we formulate the true and generalization errors of the model for both training and validation/test instances where we make use of the Stein's Unbiased Risk Estim…
Study k-folding map-germs to understand surface geometry.
K-fold CV improves machine learning model selection but faces challenges with small datasets.
A new cross-validation method reduces redundancy and improves model performance.
K-fold cross validation (CV) is a popular method for estimating the true performance of machine learning models, allowing model selection and parameter tuning. However, the very process of CV requires random partitioning of the data and so our performance estimates are in fact stochastic, with variability that can be s…
We construct the full linearisation functor which takes a graded bundle of degree (a particular kind of graded manifold) and produces a -fold vector bundle. We fully characterise the image of the full linearisation functor and show that we obtain a subcategory of -fold vector bundles consisting of symmetric $…
A manifold is locally \emph{-fold symmetric}, if for any point and any -dimensional vector subspace tangent to this point there exists a local isometry such that this point is a fixed point and the differential of the isometry restricted to that -dimensional vector subspace is minus the identity. We show that …
Groups with specific curvature have a regular language of geodesics.
Proposes K-Fold Causal BART for improved CATE estimation.
Improves test set performance and reduces out-of-sample disappointment for unstable models.
New method improves model risk prediction using cross-audit projection.
Let be a smooth generic immersion. Then the set of points, that have at least preimages is an image of a (non-generic) immersion. If the manifolds and are oriented and is even, then the manifold of -fold points is also oriented. In this paper we compute the oriented b…
Proposes a method to create prediction intervals for neural networks using cross-validation.
K-fold Cross Validation is commonly used to evaluate classifiers and tune their hyperparameters. However, it assumes that data points are Independent and Identically Distributed (i.i.d.) so that samples used in the training and test sets can be selected randomly and uniformly. In Human Activity Recognition datasets, we…
We consider branched coverings which are simple in the sense that any point of the target has at most one singular preimage. The cobordism classes of -fold simple branched coverings between -manifolds form an abelian group . Moreover, is a module over…
The paper improves confidence intervals for test error using cross-validation.
In this article, we derive concentration inequalities for the cross-validation estimate of the generalization error for stable predictors in the context of risk assessment. The notion of stability has been first introduced by \cite{DEWA79} and extended by \cite{KEA95}, \cite{BE01} and \cite{KUNIY02} to characterize cla…
Study uses machine learning to predict stroke risk in China with improved accuracy.
We consider a priori generalization bounds developed in terms of cross-validation estimates and the stability of learners. In particular, we first derive an exponential Efron-Stein type tail inequality for the concentration of a general function of n independent random variables. Next, under some reasonable notion of s…
RandALO speeds up risk estimation for large datasets.
We construct new constant mean curvature surfaces in H2xR. They arise as sister surfaces of Plateau solutions. It is a family of MC 1/2 surfaces with k ends, genus 1 and k-fold dihedral symmetry, k greater 2. The surfaces are Alexandrov- embedded.
Echo State Networks (ESNs) are known for their fast and precise one-shot learning of time series. But they often need good hyper-parameter tuning for best performance. For this good validation is key, but usually, a single validation split is used. In this rather practical contribution we suggest several schemes for cr…
Many versions of cross-validation (CV) exist in the literature; and each version though has different variants. All are used interchangeably by many practitioners; yet, without explanation to the connection or difference among them. This article has three contributions. First, it starts by mathematical formalization of…
Given smooth manifolds and , an integer , and an immersion , we have constructed an obstruction for existence of regular homotopy of to an immersion without -fold points. This obstruction takes values in certain framed bordism group, and for $(k+1)(n+1)…
We develop a robust convex algorithm to select the regularization parameter in model selection. In practice this would be automated in order to save practitioners time from having to tune it manually. In particular, we implement and test the convex method for -fold cross validation on ridge regression, although the …
Most modern supervised statistical/machine learning (ML) methods are explicitly designed to solve prediction problems very well. Achieving this goal does not imply that these methods automatically deliver good estimators of causal parameters. Examples of such parameters include individual regression coefficients, avera…
Optimal data splitting improves covariance matrix estimation in large datasets.
SOAK assesses data subset similarity for better model training.
Given a multifunction from to the fold symmetric product , we use the Dold-Thom Theorem to establish a homological selection Theorem. This is used to establish existence of Nash equilibria. Cost functions in problems concerning the existence of Nash Equilibria are traditionally multilinear in the mixe…
In this note we prove that any monohedral tiling of the closed circular unit disc with topological discs as tiles has a -fold rotational symmetry. This result yields the first nontrivial estimate about the minimum number of tiles in a monohedral tiling of the circular disc in which not all tiles contain t…
One can formulate the classical Kepler problem on the Heisenberg group, the simplest sub-Riemannian manifold. We take the sub-Riemannian Hamiltonian as our kinetic energy, and our potential is the fundamental solution to the Heisenberg sub-Laplacian. The resulting dynamical system is known to contain a fundamental inte…
Fair MP-Boost improves fairness and interpretability in boosting methods.
MP-Boost boosts accuracy faster and more interpretable than AdaBoost.
Boost-R uses gradient boosted trees for analyzing recurrence data.
Boosted decision trees typically yield good accuracy, precision, and ROC area. However, because the outputs from boosting are not well calibrated posterior probabilities, boosting yields poor squared error and cross-entropy. We empirically demonstrate why AdaBoost predicts distorted probabilities and examine three cali…
Extends boosting to multiclass online agnostic classification.
The VC-dimension of a set system is a way to capture its complexity and has been a key parameter studied extensively in machine learning and geometry communities. In this paper, we resolve two longstanding open problems on bounding the VC-dimension of two fundamental set systems: -fold unions/intersections of half-s…
Boosting is one of the most significant developments in machine learning. This paper studies the rate of convergence of Boosting, which is tailored for regression, in a high-dimensional setting. Moreover, we introduce so-called \textquotedblleft post-Boosting\textquotedblright. This is a post-selection estimator w…
Robust boosting improves regression accuracy in noisy data.
We present a new boosting algorithm, motivated by the large margins theory for boosting. We give experimental evidence that the new algorithm is significantly more robust against label noise than existing boosting algorithm.
Improved agnostic boosting with better sample efficiency.
Excellent ranking power along with well calibrated probability estimates are needed in many classification tasks. In this paper, we introduce a technique, Calibrated Boosting-Forest that captures both. This novel technique is an ensemble of gradient boosting machines that can support both continuous and binary labels. …
A Sasakian structure on a manifold is called {\it positive} if its basic first Chern class can be represented by a positive (1,1)-form with respect to its transverse holomorphic CR-structure. We prove a theorem that says that every positive Sasakian structure can be deformed to a Sasakian structure whose metric has pos…
Boosting algorithms are frequently used in applied data science and in research. To date, the distinction between boosting with either gradient descent or second-order Newton updates is often not made in both applied and methodological research, and it is thus implicitly assumed that the difference is irrelevant. The g…
The study provides statistical guarantees for Bayesian variational boosting.
In this paper we propose using the principle of boosting to reduce the bias of a random forest prediction in the regression setting. From the original random forest fit we extract the residuals and then fit another random forest to these residuals. We call the sum of these two random forests a \textit{one-step boosted …