Recent studies have shown that tuning prediction models increases prediction accuracy and that Random Forest can be used to construct prediction intervals. However, to our best knowledge, no study has investigated the need to, and the manner in which one can, tune Random Forest for optimizing prediction intervals { thi…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
In dynamic selection (DS) techniques, only the most competent classifiers, for the classification of a specific test sample are selected to predict the sample's class labels. The more important step in DES techniques is estimating the competence of the base classifiers for the classification of each specific test sampl…
Deep learning techniques have been hugely successful for traditional supervised and unsupervised machine learning problems. In large part, these techniques solve continuous optimization problems. Recently however, discrete generative deep learning models have been successfully used to efficiently search high-dimensiona…
New validation method prevents privacy breaches and biases in federated learning.
The correct use of model evaluation, model selection, and algorithm selection techniques is vital in academic machine learning research as well as in many industrial settings. This article reviews different techniques that can be used for each of these three subtasks and discusses the main advantages and disadvantages …
Survey of algorithms for testing AI-driven CPS safety.
ECV method optimizes ensemble parameters for randomized ensembles.
Machine learning detects survey validity from user behavior.
Proposes approximating computationally expensive explainability techniques using conformal regression.
Interpretable ML helps discover insights from big data.
Neural Architecture Search remains a very challenging meta-learning problem. Several recent techniques based on parameter-sharing idea have focused on reducing the NAS running time by leveraging proxy models, leading to architectures with competitive performance compared to those with hand-crafted designs. In this pape…
The paper identifies overfitting as the main bottleneck in efficient deep reinforcement learning.
We provide a rigorous numerical computation method to validate periodic, homoclinic and heteroclinic orbits as the continuation of singular limit orbits for the fast-slow system with one-dimensional slow variable . Our validation procedure is based on topological tools called isolatin…
This paper elaborates on the validation requirements for rating systems and probabilities of default (PDs) which were introduced with the New Capital Standards (Basel II). We start in Section 2 with some introductory remarks on the topics and approaches that will be discussed later on. Then we have a view on the develo…
Study compares resampling methods for rare event prediction in longitudinal studies.
New method narrows prediction intervals for individual treatment effects.
Method selects valid IVs from a large set using clustering and test of overidentifying restrictions.
New IF method improves accuracy in deep neural networks with noisy data.
Proves and tests methods for learning time-series with breaks.
New approach uses dynamic programming to efficiently discover failures in autonomous vehicle simulations.
Linear principal component analysis (PCA) can be extended to a nonlinear PCA by using artificial neural networks. But the benefit of curved components requires a careful control of the model complexity. Moreover, standard techniques for model selection, including cross-validation and more generally the use of an indepe…
The paper develops a cross-validation method for improving signal denoising techniques.
Paper improves PAC-Bayes bounds for various loss types.
New approach uses interpolation models and error bounds for verifiable scientific machine learning.
We introduce a new normalization technique that exhibits the fast convergence properties of batch normalization using a transformation of layer weights instead of layer outputs. The proposed technique keeps the contribution of positive and negative weights to the layer output balanced. We validate our method on a set o…
A novel validation method improves feature importance analysis in subject-specific ML models.
We aim at developing and improving the imbalanced business risk modeling via jointly using proper evaluation criteria, resampling, cross-validation, classifier regularization, and ensembling techniques. Area Under the Receiver Operating Characteristic Curve (AUC of ROC) is used for model comparison based on 10-fold cro…
A fast bootstrap method estimates cross-validation standard error.
Causal ML methods failed to validate their personalized treatment effects in two large trials.
Method improves treatment effect estimation in randomized experiments.
Popular machine learning estimators involve regularization parameters that can be challenging to tune, and standard strategies rely on grid search for this task. In this paper, we revisit the techniques of approximating the regularization path up to predefined tolerance in a unified framework and show that its comp…
We extend the validity of the Penrose singularity theorem to spacetime metrics of regularity . The proof is based on regularisation techniques, combined with recent results in low regularity causality theory.
Early stopping is a widely used technique to prevent poor generalization performance when training an over-expressive model by means of gradient-based optimization. To find a good point to halt the optimizer, a common practice is to split the dataset into a training and a smaller validation set to obtain an ongoing est…
Cross-validation estimates model performance on unseen data, not training data.
Confounding bias, missing data, and selection bias are three common obstacles to valid causal inference in the data sciences. Covariate adjustment is the most pervasive technique for recovering casual effects from confounding bias. In this paper, we introduce a covariate adjustment formulation for controlling confoundi…
We consider the problem of decomposing a multivariate polynomial as the difference of two convex polynomials. We introduce algebraic techniques which reduce this task to linear, second order cone, and semidefinite programming. This allows us to optimize over subsets of valid difference of convex decompositions (dcds) a…
Paper presents a workflow for reliable unsupervised learning in science.
New insights into ML models' accuracy and generalization for scientific problems.
A new method controls risk for set predictors using cross-validation.
Model inference, such as model comparison, model checking, and model selection, is an important part of model development. Leave-one-out cross-validation (LOO) is a general approach for assessing the generalizability of a model, but unfortunately, LOO does not scale well to large datasets. We propose a combination of u…
Study finds unsupervised imputation before cross-validation can reduce computational costs without significantly degrading model performance.
A novel ML verification technique using manifold learning.
The defining equations for Killing vector fields and conformal Killing vector fields are overdetermined systems of PDE. This makes it difficult to solve the systems numerically. We propose an approach which reduces the computation to the solution of a symmetric eigenvalue problem. The eigenvalue problem is then solved …
Unified techniques improve stability and replicability in changing data.
Machine learning systems increasingly depend on pipelines of multiple algorithms to provide high quality and well structured predictions. This paper argues interaction effects between clustering and prediction (e.g. classification, regression) algorithms can cause subtle adverse behaviors during cross-validation that m…
Recently, new defense techniques have been developed to tolerate Byzantine failures for distributed machine learning. The Byzantine model captures workers that behave arbitrarily, including malicious and compromised workers. In this paper, we break two prevailing Byzantine-tolerant techniques. Specifically we show robu…
The paper validates statistical models for groundwater data.
Techniques for understanding the functioning of complex machine learning models are becoming increasingly popular, not only to improve the validation process, but also to extract new insights about the data via exploratory analysis. Though a large class of such tools currently exists, most assume that predictions are p…