Large language models correlate in errors, even with different architectures and providers.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper tackles fairness in CCA by minimizing correlation disparity error.
A possible data source for the estimation of asset correlations is default time series. This study investigates the systematic error that is made if the exposure pool underlying a default time series is assumed to be homogeneous when in reality it is not. We find that the asset correlation will always be underestimated…
The paper introduces a method to model error correlations in multivariate time series forecasting.
Overparameterized models can worsen minority group errors even when overall test error improves.
Bayesian method improves segmentation accuracy with noisy labels.
Paper introduces a new cost function to improve deep learning model generalization.
While the objective in traditional multi-armed bandit problems is to find the arm with the highest mean, in many settings, finding an arm that best captures information about other arms is of interest. This objective, however, requires learning the underlying correlation structure and not just the means of the arms. Se…
In their seminal work Carr and Lee (2008) show how to robustly price and replicate a variety of claims written on the quadratic variation of a risky asset under the assumption that the asset's volatility process is independent of the Brownian motion that drives the asset's price. Additionally, they propose a correlatio…
K-fold cross-validation (CV) with squared error loss is widely used for evaluating predictive models, especially when strong distributional assumptions cannot be taken. However, CV with squared error loss is not free from distributional assumptions, in particular in cases involving non-i.i.d. data. This paper analyzes …
We introduce a new approach to variable selection, called Predictive Correlation Screening, for predictor design. Predictive Correlation Screening (PCS) implements false positive control on the selected variables, is well suited to small sample sizes, and is scalable to high dimensions. We establish asymptotic bounds f…
Existing guarantees in terms of rigorous upper bounds on the generalization error for the original random forest algorithm, one of the most frequently used machine learning methods, are unsatisfying. We discuss and evaluate various PAC-Bayesian approaches to derive such bounds. The bounds do not require additional hold…
Autonomy and adaptation of machines requires that they be able to measure their own errors. We consider the advantages and limitations of such an approach when a machine has to measure the error in a regression task. How can a machine measure the error of regression sub-components when it does not have the ground truth…
New DP methods for estimating means and frequencies with varying privacy demands.
Sparse GCA finds linear relationships in multiple datasets, using gradient descent.
We study the relationship between catastrophic forgetting and properties of task sequences. In particular, given a sequence of tasks, we would like to understand which properties of this sequence influence the error rates of continual learning algorithms trained on the sequence. To this end, we propose a new procedure …
The risk of a credit portfolio depends crucially on correlations between the probability of default (PD) in different economic sectors. Often, PD correlations have to be estimated from relatively short time series of default rates, and the resulting estimation error hinders the detection of a signal. We present statist…
This work improves confidence intervals for Cox model test error using nested CV.
We focus on emergence of the power-law cross-correlations from processes with both short and long term memory properties. In the case of correlated error-terms, the power-law decay of the cross-correlation function comes automatically with the characteristics of separate processes. Bivariate Hurst exponent is then equa…
In an efficient stock market, the log-returns and their time-dependent variances are often jointly modelled by stochastic volatility models (SVMs). Many SVMs assume that errors in log-return and latent volatility process are uncorrelated, which is unrealistic. It turns out that if a non-zero correlation is included in …
The paper studies statistical properties of CART regression trees.
Random Forests adapted for dependent data using GLS.
The problem of learning tree-structured Gaussian graphical models from independent and identically distributed (i.i.d.) samples is considered. The influence of the tree structure and the parameters of the Gaussian distribution on the learning rate as the number of samples increases is discussed. Specifically, the error…
The study corrects measurement error in evaluating health effects of multiple pollutants.
A visualization aids in comparing regression models by highlighting errors and correlations.
Paper develops streaming algorithms to estimate classifier accuracy on unlabeled data.
Kernel methods linked to feature subspaces and maximal correlation kernels.
Redundancy in AI perception systems doesn't guarantee independent error occurrences.
We solve matrix denoising with both row and column correlations, setting limits and designing optimal methods.
Geostatistical learning faces unique challenges due to spatial correlation and covariate shifts.
In this work, we present a novel robust distributed beamforming (RDB) approach based on low-rank and cross-correlation techniques. The proposed RDB approach mitigates the effects of channel errors in wireless networks equipped with relays based on the exploitation of the cross-correlation between the received data from…
Improved deep probabilistic time series forecasting by learning error autocorrelation.
This paper studies ordered weighted L1 (OWL) norm regularization for sparse estimation problems with strongly correlated variables. We prove sufficient conditions for clustering based on the correlation/colinearity of variables using the OWL norm, of which the so-called OSCAR is a particular case. Our results extend pr…
We present a method to compensate statistical errors in the calculation of correlations on asynchronous time series. The method is based on the assumption of an underlying time series. We set up a model and apply it to financial data to examine the decrease of calculated correlations towards smaller return intervals (E…
Alpha-based performance evaluation may fail to capture correlated residuals due to model errors. This paper proposes using the Generalized Information Ratio (GIR) to measure performance under misspecified benchmarks. Motivated by the theoretical link between abnormal returns and residual covariance matrix, GIR is deriv…
Enhances robustness of MOGP regression for multiple correlated outputs.
Data mining techniques on the biological analysis are spreading for most of the areas including the health care and medical information. We have applied the data mining techniques, such as KNN, SVM, MLP or decision trees over a unique dataset, which is collected from 16,380 analysis results for a year. Furthermore we h…
A new method selects models for ensemble learning to maximize mutual information, outperforming existing approaches.
Efficient algorithm for graph matching in correlated stochastic block models.
The strength of association between a pair of data vectors is represented by a nonnegative real number, called matching weight. For dimensionality reduction, we consider a linear transformation of data vectors, and define a matching error as the weighted sum of squared distances between transformed vectors with respect…
This paper sets thresholds for recovering vertex correspondences in partially correlated graphs.
A new method for approximating softmax and Gaussian kernels with reduced error.
Proposes a method to select features for deep learning in noisy, high-dimensional data.
The Sharpe ratio, which is defined as the ratio of the excess expected return of an investment to its standard deviation, has been widely cited in the financial literature by researchers and practitioners. However, very little attention has been paid to the statistical properties of the estimation of the ratio. Lo (200…
We focus on power-law coherency as an alternative approach towards studying power-law cross-correlations between simultaneously recorded time series. To be able to study empirical data, we introduce three estimators of the power-law coherency parameter based on popular techniques usually utilized for studying pow…
Paper forecasts stock correlations using a hybrid model combining graph neural networks and transformers.
We demonstrate that the lowest possible price change (tick-size) has a large impact on the structure of financial return distributions. It induces a microstructure as well as it can alter the tail behavior. On small return intervals, the tick-size can distort the calculation of correlations. This especially occurs on s…
Locally private algorithm improves online federated learning with correlated noise.