Study uses random matrix test to find significant factors in cryptocurrency forecasts.
problem Determining the optimal number of factors in cryptocurrency forecast models.
method Applied a random matrix test to a forecast model of Reduced Rank Regression (RRR) on cryptocurrencies.
result Consistent results with visual inspection, minimal computational cost compared to cross-validation.
Paper uses Random Matrix Theory for optimal training-testing data split.
problem Finding ideal training-testing data split for linear regression.
method Random Matrix Theory applied to Gaussian multivariate data.
result Ideal training and test sizes derived for any model.
New test for latent block models to determine cluster numbers.
problem No statistical test for latent block models.
method Developed a goodness-of-fit test using random matrix theory.
result Demonstrated the effectiveness of the test method.
CovRegRF estimates covariance matrix from covariates using random forests.
problem Estimating conditional covariances or correlations among multivariate responses.
method Random forest trees with a custom splitting rule to maximize covariance difference.
result Accurate covariance matrix estimates and controlled Type-1 error.
Improves detection of low-rank signals from noisy data matrices.
problem Statistical detection of low-rank signals in noisy data matrices.
method Entrywise pre-transforming data matrix for non-Gaussian noise, sharp phase transition thresholds, central limit theorem for linear spectral statistics, hypothesis test.
result Improves detection of low-rank signals from noisy data matrices, generalizing known results.
U-Net trained to recover acoustic interference striations from distorted data.
problem Recovering acoustic interference striations from distorted signals.
method Training a U-Net using a random mode-coupling matrix model to generate training data.
result U-Net successfully recovers AISs under various conditions.
Geometric QHD tests improve hub detection in correlated data.
problem Detecting hubs in correlated data with evolving correlations.
method Geometric QHD tests combining QCD and QHD, clustering.
result Improved hub detection in correlated data.
Study proposes memory-efficient backpropagation for linear layers in neural networks.
problem Significant memory usage in backpropagation through linear layers in neural networks.
method Randomized matrix multiplications to reduce memory usage with a moderate decrease in test accuracy.
result Demonstrated benefits of the proposed method on fine-tuning pre-trained models.
Dimensional reduction of high dimensional data can be achieved by keeping only the relevant eigenmodes after principal component analysis. However, differentiating relevant eigenmodes from the random noise eigenmodes is problematic. A new method based on the random matrix theory and a statistical goodness-of-fit test i…
Study characterizes training and test risks for MAP regression with Gaussian priors.
problem Understanding high-dimensional behavior of regularized linear regression with informative priors.
method Maximum a posteriori (MAP) regression with Gaussian priors, using random matrix theory.
result Closed-form risk formulas reveal the bias-variance-prior tradeoff and explain double descent.
Meta-learning improves predictions with generalized ridge regression in high-dimensional settings.
problem Improving meta-learning performance in high-dimensional settings.
method Generalized ridge regression applied to high-dimensional multivariate random-effects linear models.
result Optimal predictive risk achieved when using the inverse of the covariance matrix of random coefficients.
We present a new method for estimating multivariate, second-order stationary Gaussian Random Field (GRF) models based on the Sparse Precision matrix Selection (SPS) algorithm, proposed by Davanloo et al. (2015) for estimating scalar GRF models. Theoretical convergence rates for the estimated between-response covariance…
The paper develops a test for EU portfolio efficiency in high dimensions.
problem Testing the efficiency of the EU portfolio in high-dimensional settings.
method Shrinkage-based approach for portfolio weights and random matrix theory.
result Asymptotic behavior of the test statistic under high-dimensional conditions.
This paper presents a sequential randomized lowrank matrix factorization approach for incrementally predicting values of an unknown function at test points using the Gaussian Processes framework. It is well-known that in the Gaussian processes framework, the computational bottlenecks are the inversion of the (regulariz…
Characterizes RFF regression in large n,p,N setting, providing precise learning phases and double descent curve.
problem Characterizes RFF regression in large n,p,N setting. method Characterizes the exact asymptotics of random Fourier feature (RFF) regression in the realistic setting of large n,p,N. result Characterizes two qualitatively different phases of learning and the corresponding double descent test error curve.
Investment diversification affects financial stability, depending on network connectivity.
problem Analyzing stability of financial networks with diversified portfolios.
method Random matrix dynamical model with portfolio rebalancing, considering heterogeneity and diversification effects.
result Stability/instability transition depends on the largest eigenvalue of the random matrix.
RF models implicitly regularize kernel methods as feature count increases.
problem Understanding implicit regularization in RF models.
method Random matrix theory applied to Gaussian RF models and KRR.
result The average RF predictor is close to a KRR predictor with an effective ridge.
A central problem of random matrix theory is to understand the eigenvalues of spiked random matrix models, introduced by Johnstone, in which a prominent eigenvector (or "spike") is planted into a random matrix. These distributions form natural statistical models for principal component analysis (PCA) problems throughou…
DBCL defends collaborative learning by sketching parameters to prevent gradient-based privacy inference attacks.
problem Privacy leaks in collaborative machine learning due to gradient-based attacks.
method Random matrix sketching applied to parameters, followed by re-generation of sketching after each iteration.
result DBCL prevents effective gradient-based privacy inference attacks without significant computational or accuracy costs.
We use matricial free energy to regularize autoencoders, producing Gaussian-like codes.
problem Generating Gaussian-like codes for autoencoders.
method Define a differentiable loss function based on singular values of the code matrix, minimizing matricial free energy.
result Minimizing matricial free energy results in Gaussian-like codes that generalize.
Study shows mixtures of nonlinearities can improve deep learning performance.
problem Improving deep learning performance with large datasets and complex models.
method Analyzed random feature regression with features F=f(WX+B) for a random weight matrix W and random bias vector B. result Mixture of nonlinearities can improve both training and test errors over a single nonlinearity.
We analyze cross-correlations between price fluctuations of different stocks using methods of random matrix theory (RMT). Using two large databases, we calculate cross-correlation matrices C of returns constructed from (i) 30-min returns of 1000 US stocks for the 2-yr period 1994--95 (ii) 30-min returns of 881 US stock…
Paper analyzes singular subspace estimation in noisy matrix models.
problem Estimating low-rank signals in noisy matrix data.
method Asymptotic distributional theory, extreme value theory, saddle point approximation, random matrix theory.
result Plug-in test statistic based on two-to-infinity norm has higher power for detecting structured alternatives.
The paper uses random matrix theory for multi-task regression, improving time series forecasting.
problem Improving time series forecasting using multi-task regression.
method Applying random matrix theory to multi-task regression problems, deriving closed-form solutions for optimization.
result Provides a robust foundation for hyperparameter optimization in multi-task regression scenarios.
Mixed membership factorization is a popular approach for analyzing data sets that have within-sample heterogeneity. In recent years, several algorithms have been developed for mixed membership matrix factorization, but they only guarantee estimates from a local optimum. Here, we derive a global optimization (GOP) algor…
New method allows generating independent data matrices from summary statistics.
problem Generating independent data matrices from summary statistics like mean and covariance.
method Thinning a Wishart random matrix based on sample mean and covariance.
result It is possible to generate two independent data matrices from summary statistics.
Paper develops statistical tests for covariance matrix regression on manifold.
problem Regression with random covariance matrices in Fréchet space.
method Develops Wasserstein F-tests for Bures-Wasserstein manifold.
result Asymptotic null distribution and power of the test.
We apply random matrix theory to compare correlation matrix estimators C obtained from emerging market data. The correlation matrices are constructed from 10 years of daily data for stocks listed on the Johannesburg Stock Exchange (JSE) from January 1993 to December 2002. We test the spectral properties of C against ra…
New method for inference on covariates in NMF with random effects.
problem Formal inference for covariate effects in NMF with non-negativity constraints.
method NMF-RE model with random effects, ridge updates, df-based cap, asymptotic linearization, wild bootstrap.
result Valid inference on covariates with non-negativity constraint, avoiding degeneracy.
In this study, we attempted to determine how eigenvalues change, according to random matrix theory (RMT), in stock market data as the number of stocks comprising the correlation matrix changes. Specifically, we tested for changes in the eigenvalue properties as a function of the number and type of stocks in the correla…
Quantum-assisted Gaussian process speeds up data regression.
problem High computational complexity of Gaussian process regression for large datasets.
method Quantum-assisted sparse Gaussian process regression using random Fourier features.
result Achieves polynomial-order computational speedup compared to classical methods.
The paper addresses statistical inference in matching markets with dependent missingness.
problem Statistical inference for two-sided matching markets with matching-induced dependence.
method Non-convex algorithm based on Grassmannian gradient descent, debiasing and projection framework.
result Near-optimal entrywise convergence rates for various matching mechanisms.
The paper examines how well node similarities are preserved by random projections in graph embeddings.
problem The preservation of node similarities under random projections in graph embeddings.
method Investigation of dot product and cosine similarity preservation by random projections over graph matrix rows.
result Random projections produce unreliable embeddings for dot product, especially for high-degree nodes.
The paper proposes a new portfolio allocation method combining RMT and machine learning.
problem Optimal allocation instability in high-dimensional portfolios.
method Combines Random Matrix Theory covariance estimators with Nested Clustered Optimization.
result The modified NCO algorithm achieves stable allocations without risky short positions.
New tool detects 'fleeting modes' causing excess risk in financial markets.
problem Detecting portfolios with statistically significant excess risk in financial markets.
method Random Matrix Theory to identify 'fleeting modes' independent of underlying correlation structure.
result Fleeting modes exist in both futures and equity markets, and momentum is a source of excess risk.
New tools in nonlinear random matrices improve understanding of the Sum of Squares hierarchy.
problem Improving the Sum of Squares (SoS) hierarchy's performance on average-case problems.
method Developed new tools in nonlinear random matrices and applied them to analyze the SoS hierarchy.
result Subexponential-time SoS lower bounds for various problems, offering evidence for the low-degree likelihood ratio hypothesis.
Noise-cleaning fMRI brain activity matrices for better precision estimation.
problem Denoise precision matrices of fMRI time series to estimate true matrices.
method Comparison of various noise-cleaning algorithms on synthetic and real fMRI data.
result Optimal Rotationally Invariant Estimator outperforms others in fMRI data.
Maximizes stock portfolio predictability using machine learning.
problem Improving stock portfolio performance through predictive modeling.
method Optimal constrained weights in the MPP constructed using Elastic Net, Random Forest, and Support Vector Regression models.
result MPP portfolios can outperform or underperform the index based on the time period.
Paper proposes a generalized precision matrix for t-Student distributions to improve portfolio optimization.
problem Limitations of inverse covariance matrix in non-Gaussian settings.
method Exploits local dependence function to define generalized precision matrix (GPM) for multivariate t-Student distribution.
result GPM leads to statistically significant lower out-of-sample variances in minimum-variance portfolios.
Analyzes bias-variance in overparameterized linear models using random features.
problem Understanding bias-variance trade-off in overparameterized models.
method Zero-temperature cavity method and random matrix theory.
result Three phase transitions in the linear random features model.
New spectral tests assess network model fits efficiently.
problem Determining if network models fit data well and extrapolate.
method Random matrix theory-derived goodness-of-fit tests.
result General approach simplifies parameter selection in network models.
This paper explores how random sampling and coding can speed up approximate matrix multiplication.
problem Efficiently computing large-scale matrix multiplications in distributed systems.
method Proposes two schemes: coding for recovery and random sampling for approximation.
result Investigates tradeoffs between recovery threshold and approximation error.
Study on eigenvalue distribution of correlated time series deforming the semi-circle law.
problem Eigenvalue distribution of correlated time series differs from the semi-circle law.
method Analysis of Wigner random matrix with temporal correlation.
result Eigenvalue distribution converges to a deformed semi-circle law with longer tail and higher peak.
Paper develops new method for detecting latent structure in large symmetric data matrices.
problem Testing for latent structure in large symmetric data matrices.
method Introduces Wilcoxon--Wigner random matrices based on normalized rank statistics.
result Establishes asymptotic Gaussian fluctuations for leading eigenvalue and eigenvector of Wilcoxon--Wigner matrices.
We present a general framework, the coupled compound Poisson factorization (CCPF), to capture the missing-data mechanism in extremely sparse data sets by coupling a hierarchical Poisson factorization with an arbitrary data-generating model. We derive a stochastic variational inference algorithm for the resulting model …
Paper develops a new test for high-dimensional matrix-valued data.
problem Hypothesis testing for mean of matrix-valued data in high-dimensional settings.
method Proposes a new test statistic for high-dimensional matrix rank testing.
result Develops a novel approach for sparse singular value decomposition (SVD) estimation.
New tail inequalities for sums of random matrices without matrix-dimension terms.
problem Tail behavior of matrix functions in high-dimensional settings.
method Developed new tail inequalities for matrix sums, independent of matrix dimension.
result Tail inequalities for various matrix functions without matrix-dimension terms.
Improved matrix approximation using randomized algorithms.
problem Finding better approximations of given matrices.
method Randomized algorithms to compute (HT) as an improved approximation. result Computed (HT) provides a better approximation than given F∗.