Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

14294357 · Jun 202019922001200920172026
48 results for k-fold boosting

Improved portfolio optimization method yields better risk-adjusted returns.

problem Optimizing global minimum variance portfolios with reduced risk.
method k-fold boosted kk-BAHC covariance cleaning procedure for correlation matrices.
result Our method outperforms other filtering methods in Sharpe ratios, despite higher turnover.

K-fold CV improves machine learning model selection but faces challenges with small datasets.

problem Challenges in validating machine learning models, especially with small datasets.
method K-fold cross-validation with K-fold CUBV (Upper Bound of the actual risk) and Upper Bound of the actual risk for linear classifiers.
result K-fold CUBV is a robust criterion for detecting effects and validating accuracy values from machine learning models.

We construct the full linearisation functor which takes a graded bundle of degree kk (a particular kind of graded manifold) and produces a kk-fold vector bundle. We fully characterise the image of the full linearisation functor and show that we obtain a subcategory of kk-fold vector bundles consisting of symmetric $…

2015-12-08abs ↗pdf ↗

A manifold is locally \emph{kk-fold symmetric}, if for any point and any kk-dimensional vector subspace tangent to this point there exists a local isometry such that this point is a fixed point and the differential of the isometry restricted to that kk-dimensional vector subspace is minus the identity. We show that …

2016-07-19abs ↗pdf ↗

Groups with specific curvature have a regular language of geodesics.

problem Understanding the language of geodesics in non-positively curved triangle groups.
method Proving finitely many cone types and regularity of geodesic languages.
result The language of lexicographically first geodesics is regular and satisfies the fellow traveller property.

Proposes K-Fold Causal BART for improved CATE estimation.

problem Improving estimation of Conditional Average Treatment Effects (CATE).
method K-Fold Causal Bayesian Additive Regression Trees (K-Fold Causal BART).
result K-Fold Causal BART is not state-of-the-art for ATE and CATE estimation in the IHDP dataset, but provides insights into model robustness and evaluation methods.

Improves test set performance and reduces out-of-sample disappointment for unstable models.

problem Ensuring strong test set performance via cross-validation for unstable models.
method Nested k-fold cross-validation with hyperparameter selection based on a weighted sum of cross-validation metric and model stability measure.
result Improves out-of-sample MSE for sparse ridge regression and CART by 4% and 2% respectively, compared to k-fold cross-validation.

Let f:VnMmf:V^n\looparrowright M^m be a smooth generic immersion. Then the set of points, that have at least kk preimages is an image of a (non-generic) immersion. If the manifolds VnV^n and MmM^m are oriented and mnm-n is even, then the manifold of kk-fold points is also oriented. In this paper we compute the oriented b…

2000-08-07abs ↗pdf ↗

Proposes a method to create prediction intervals for neural networks using cross-validation.

problem Lack of prediction intervals for neural networks.
method k-fold cross-validation to construct conformal prediction intervals.
result Proposed method produces narrower intervals with similar coverage compared to SC method.

K-fold Cross Validation is commonly used to evaluate classifiers and tune their hyperparameters. However, it assumes that data points are Independent and Identically Distributed (i.i.d.) so that samples used in the training and test sets can be selected randomly and uniformly. In Human Activity Recognition datasets, we…

2019-04-04abs ↗pdf ↗

We consider branched coverings which are simple in the sense that any point of the target has at most one singular preimage. The cobordism classes of kk-fold simple branched coverings between nn-manifolds form an abelian group Cob1(n,k)Cob^1(n,k). Moreover, Cob1(,k)=n=0Cob1(n,k)Cob^1(*,k) = \bigoplus_{n=0}^{\infty} Cob^1(n,k) is a module over…

2017-07-08abs ↗pdf ↗

Study uses machine learning to predict stroke risk in China with improved accuracy.

problem Predicting stroke risk in China using machine learning.
method Combines VAR and GNN for causal inference, compares multiple classification algorithms, uses SMOTE undersampling.
result Gradient Boosting model shows highest performance and stability in predicting stroke risk.

We consider a priori generalization bounds developed in terms of cross-validation estimates and the stability of learners. In particular, we first derive an exponential Efron-Stein type tail inequality for the concentration of a general function of n independent random variables. Next, under some reasonable notion of s…

2017-06-19abs ↗pdf ↗

Echo State Networks (ESNs) are known for their fast and precise one-shot learning of time series. But they often need good hyper-parameter tuning for best performance. For this good validation is key, but usually, a single validation split is used. In this rather practical contribution we suggest several schemes for cr…

2019-08-22abs ↗pdf ↗

Given smooth manifolds VnV^n and MmM^m, an integer kk, and an immersion f:VMf:V\looparrowright M, we have constructed an obstruction for existence of regular homotopy of ff to an immersion f:VMf':V\looparrowright M without kk-fold points. This obstruction takes values in certain framed bordism group, and for $(k+1)(n+1)…

2002-03-13abs ↗pdf ↗

We develop a robust convex algorithm to select the regularization parameter in model selection. In practice this would be automated in order to save practitioners time from having to tune it manually. In particular, we implement and test the convex method for KK-fold cross validation on ridge regression, although the …

2014-11-27abs ↗pdf ↗

Most modern supervised statistical/machine learning (ML) methods are explicitly designed to solve prediction problems very well. Achieving this goal does not imply that these methods automatically deliver good estimators of causal parameters. Examples of such parameters include individual regression coefficients, avera…

2016-07-30abs ↗pdf ↗

Optimal data splitting improves covariance matrix estimation in large datasets.

problem Improving large covariance matrix estimation in high-dimensional settings.
method Focus on holdout method, derive closed-form error expression, connect to eigenvalue variance.
result Optimal train-test split scales as square root of matrix dimension.

Given a multifunction from XX to the kk-fold symmetric product Symk(X)Sym_k(X), we use the Dold-Thom Theorem to establish a homological selection Theorem. This is used to establish existence of Nash equilibria. Cost functions in problems concerning the existence of Nash Equilibria are traditionally multilinear in the mixe…

2011-11-03abs ↗pdf ↗

In this note we prove that any monohedral tiling of the closed circular unit disc with k3k \leq 3 topological discs as tiles has a kk-fold rotational symmetry. This result yields the first nontrivial estimate about the minimum number of tiles in a monohedral tiling of the circular disc in which not all tiles contain t…

2019-10-09abs ↗pdf ↗

One can formulate the classical Kepler problem on the Heisenberg group, the simplest sub-Riemannian manifold. We take the sub-Riemannian Hamiltonian as our kinetic energy, and our potential is the fundamental solution to the Heisenberg sub-Laplacian. The resulting dynamical system is known to contain a fundamental inte…

2013-11-23abs ↗pdf ↗

Boosted decision trees typically yield good accuracy, precision, and ROC area. However, because the outputs from boosting are not well calibrated posterior probabilities, boosting yields poor squared error and cross-entropy. We empirically demonstrate why AdaBoost predicts distorted probabilities and examine three cali…

2012-07-04abs ↗pdf ↗

The VC-dimension of a set system is a way to capture its complexity and has been a key parameter studied extensively in machine learning and geometry communities. In this paper, we resolve two longstanding open problems on bounding the VC-dimension of two fundamental set systems: kk-fold unions/intersections of half-s…

2018-07-20abs ↗pdf ↗

Boosting is one of the most significant developments in machine learning. This paper studies the rate of convergence of L2L_2Boosting, which is tailored for regression, in a high-dimensional setting. Moreover, we introduce so-called \textquotedblleft post-Boosting\textquotedblright. This is a post-selection estimator w…

2016-02-29abs ↗pdf ↗

We present a new boosting algorithm, motivated by the large margins theory for boosting. We give experimental evidence that the new algorithm is significantly more robust against label noise than existing boosting algorithm.

2009-05-13abs ↗pdf ↗

Excellent ranking power along with well calibrated probability estimates are needed in many classification tasks. In this paper, we introduce a technique, Calibrated Boosting-Forest that captures both. This novel technique is an ensemble of gradient boosting machines that can support both continuous and binary labels. …

2017-10-16abs ↗pdf ↗

A Sasakian structure on a manifold is called {\it positive} if its basic first Chern class can be represented by a positive (1,1)-form with respect to its transverse holomorphic CR-structure. We prove a theorem that says that every positive Sasakian structure can be deformed to a Sasakian structure whose metric has pos…

2001-04-11abs ↗pdf ↗

Boosting algorithms are frequently used in applied data science and in research. To date, the distinction between boosting with either gradient descent or second-order Newton updates is often not made in both applied and methodological research, and it is thus implicitly assumed that the difference is irrelevant. The g…

2018-08-09abs ↗pdf ↗

The study provides statistical guarantees for Bayesian variational boosting.

problem Statistical and convergence issues in variational boosting.
method Proposed a novel variational family and a functional Frank-Wolfe optimization algorithm.
result Demonstrated stochastic boundedness and provided convergence rate for boosting iterates.