A model order reduction framework reduces financial risk analysis models efficiently.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We present local discriminative Gaussian (LDG) dimensionality reduction, a supervised dimensionality reduction technique for classification. The LDG objective function is an approximation to the leave-one-out training error of a local quadratic discriminant analysis classifier, and thus acts locally to each training po…
Infinitesimal boosting converges to a deterministic process in large sample limit.
A new kernel test reduces noise in MMD by focusing on leading eigen-directions.
We accelerate CNF by reducing ODE truncation errors with polynomial regularization.
The paper reduces estimation error in predicting borrower repayment by accounting for lender's credit decisions.
A novel approach is put forth that utilizes data similarity, quantified on a graph, to improve upon the reconstruction performance of principal component analysis. The tasks of data dimensionality reduction and reconstruction are formulated as graph filtering operations, that enable the exploitation of data node connec…
Generalization of time series prediction remains an important open issue in machine learning, wherein earlier methods have either large generalization error or local minima. We develop an analytically solvable, unsupervised learning scheme that extracts the most informative components for predicting future inputs, term…
In the covariate shift learning scenario, the training and test covariate distributions differ, so that a predictor's average loss over the training and test distributions also differ. In this work, we explore the potential of extreme dimension reduction, i.e. to very low dimensions, in improving the performance of imp…
Semi-supervised learning improves with partial label information.
Proposes a fuzzy rule-based method for data visualization.
In this paper, we propose a way to combine two acceleration techniques for the -regularized least squares problem: safe screening tests, which allow to eliminate useless dictionary atoms; and the use of fast structured approximations of the dictionary matrix. To do so, we introduce a new family of screening …
Data-aware activation function customization reduces neural network error.
HSR reduces analyst earnings forecast errors by lowering travel friction.
Testing the implementation of deep learning systems and their training routines is crucial to maintain a reliable code base. Modern software development employs processes, such as Continuous Integration, in which changes to the software are frequently integrated and tested. However, testing the training routines requir…
Unified approach combines prediction-powered inference and variance reduction for semi-supervised optimization.
We study efficient deep learning training algorithms that process received wireless signals, if a test Signal to Noise Ratio (SNR) estimate is available. We focus on two tasks that facilitate source identification: 1- Identifying the modulation type, 2- Identifying the wireless technology and channel in the 2.4 GHz ISM…
Margin enlargement over training data has been an important strategy since perceptrons in machine learning for the purpose of boosting the robustness of classifiers toward a good generalization ability. Yet Breiman (1999) showed a dilemma that a uniform improvement on margin distribution does NOT necessarily reduces ge…
Redundancy in AI perception systems doesn't guarantee independent error occurrences.
New risk factors improve stress testing accuracy.
In this paper we propose strategies for estimating performance of a classifier when labels cannot be obtained for the whole test set. The number of test instances which can be labeled is very small compared to the whole test data size. The goal then is to obtain a precise estimate of classifier performance using as lit…
Communication overhead is a major bottleneck hampering the scalability of distributed machine learning systems. Recently, there has been a surge of interest in using gradient compression to improve the communication efficiency of distributed neural network training. Using 1-bit quantization, signSGD with majority vote …
New method reduces Monte Carlo error in option pricing and Greeks estimation.
Improves A/B testing by detecting minor treatment effects.
PEAK tests means of multiple data streams with sequential betting.
For many causal effect parameters of interest, doubly robust machine learning (DRML) estimators are the state-of-the-art, incorporating the good prediction performance of machine learning; the decreased bias of doubly robust estimators; and the analytic tractability and bias reduction of sample splitting wi…
This review tackles long horizon forecasting in time series analysis using deep learning.
Study reveals pervasive label errors in test sets, affecting machine learning benchmarks.
The paper improves confidence intervals for test error using cross-validation.
Study controls error rates of binary classifiers using hypothesis testing.
This work improves confidence intervals for Cox model test error using nested CV.
This paper tackles fairness in PCA by balancing it with reconstruction error.
New method uses machine learning to improve statistical inference.
Neuc-MDS extends MDS for non-Euclidean data.
We present a novel Monte Carlo based LSV calibration algorithm that applies to all stochastic volatility models, including the non-Markovian rough volatility family. Our framework overcomes the limitations of the particle method proposed by Guyon and Henry-Labordère (2012) and theoretically guarantees a variance reduct…
We characterize the asymptotic performance of nonparametric one- and two-sample testing. The exponential decay rate or error exponent of the type-II error probability is used as the asymptotic performance metric, and an optimal test achieves the maximum rate subject to a constant level constraint on the type-I error pr…
This paper examines how labeling error affects contrastive learning and proposes data dimensionality reduction methods to mitigate its impact.
A method for noise reduction in functional time series using FPCA.
Proposes a new test for validating multivariate dynamic regression models.
This paper introduces a new unsupervised method for dimensionality reduction via regression (DRR). The algorithm belongs to the family of invertible transforms that generalize Principal Component Analysis (PCA) by using curvilinear instead of linear features. DRR identifies the nonlinear features through multivariate r…
The repeated community-wide reuse of test sets in popular benchmark problems raises doubts about the credibility of reported test-error rates. Verifying whether a learned model is overfitted to a test set is challenging as independent test sets drawn from the same data distribution are usually unavailable, while other …
We propose to model the acoustic space of deep neural network (DNN) class-conditional posterior probabilities as a union of low-dimensional subspaces. To that end, the training posteriors are used for dictionary learning and sparse coding. Sparse representation of the test posteriors using this dictionary enables proje…
TCE measures calibration error with a test-based approach.
Study loop corrections in random feature models affecting training and test errors.
Overparameterized deep networks have the capacity to memorize training data with zero \emph{training error}. Even after memorization, the \emph{training loss} continues to approach zero, making the model overconfident and the test performance degraded. Since existing regularizers do not directly aim to avoid zero train…
DD algorithm tracks test error from train error without validation data.
Sparse model for noisy datasets using hierarchical regularization.
We use an adversarial expert based online learning algorithm to learn the optimal parameters required to maximise wealth trading zero-cost portfolio strategies. The learning algorithm is used to determine the relative population dynamics of technical trading strategies that can survive historical back-testing as well a…