The statistical leverage scores of a complex matrix record the degree of alignment between col and the coordinate axes in . These score are used in random sampling algorithms for solving certain numerical linear algebra problems. In this paper we present a max-plus algebr…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Efficiently approximates statistical leverage scores for faster KRR.
Active learning aims to obtain a classifier of high accuracy by using fewer label requests in comparison to passive learning by selecting effective queries. Many active learning methods have been developed in the past two decades, which sample queries based on informativeness or representativeness of unlabeled data poi…
One popular method for dealing with large-scale data sets is sampling. For example, by using the empirical statistical leverage scores as an importance sampling distribution, the method of algorithmic leveraging samples and rescales rows/columns of data matrices to reduce the data size before performing computations on…
In many real-world machine learning applications, unlabeled data are abundant whereas class labels are expensive and scarce. An active learner aims to obtain a model of high accuracy with as few labeled instances as possible by effectively selecting useful examples for labeling. We propose a new selection criterion tha…
Bitcoin treasury companies leverage stock to grow, using advanced statistical methods.
Massively parallel architectures such as the GPU are becoming increasingly important due to the recent proliferation of data. In this paper, we propose a key class of hybrid parallel graphlet algorithms that leverages multiple CPUs and GPUs simultaneously for computing k-vertex induced subgraph statistics (called graph…
In this paper, we consider the problem of column subset selection. We present a novel analysis of the spectral norm reconstruction for a simple randomized algorithm and establish a new bound that depends explicitly on the sampling probabilities. The sampling dependent error bound (i) allows us to better understand the …
New algorithms estimate matrix leverage scores using rank revealing and randomization.
Study tests five popular trading signal families and finds four refuted, one inconclusive, and one not refuted.
Efficiently estimates private least squares with linear error growth.
Leverage score sampling provides an appealing way to perform approximate computations for large matrices. Indeed, it allows to derive faithful approximations with a complexity adapted to the problem at hand. Yet, performing leverage scores sampling is a challenge in its own right requiring further approximations. In th…
Improved GoF statistics using entropy-regularized optimal transport for multivariate rank.
Statistical leverage scores emerged as a fundamental tool for matrix sketching and column sampling with applications to low rank approximation, regression, random feature learning and quadrature. Yet, the very nature of this quantity is barely understood. Borrowing ideas from the orthogonal polynomial literature, we in…
The article presents a translation of some widespread financial terminology into the language of decision theory. For instance, financial leverage can be regarded as an object of choice or a decision. We show how the optics of decision theory allows perceiving the recently introduced metrics of see-through-leverage, wh…
Optimizes portfolios with utility theory, diversification, and leverage.
New image restoration method using localized patches and external databases.
In this paper we develop a statistical arbitrage trading strategy with two key elements in hi-frequency trading: stop-loss and leverage. We consider, as in Bertram (2009), a mean-reverting process for the security price with proportional transaction costs; we show how to introduce stop-loss and leverage in an optimal t…
There are some statistical anomalies in the Chinese stock market, i.e., positive return skewness, anti-leverage effect (positive returns induce higher volatility than negative returns); and reverse volatility asymmetry (contemporaneous return-volatility correlation is positive). In this paper, we first confirm the exis…
Graph learning improves FXRP and FXSA with significant statistical arbitrage gains.
We study the use of randomized value functions to guide deep exploration in reinforcement learning. This offers an elegant means for synthesizing statistically and computationally efficient exploration with common practical approaches to value function learning. We present several reinforcement learning algorithms that…
Adaptive data fusion boosts efficiency in multi-task optimization.
New method finds arbitrage opportunities in fluctuating asset bands.
Machine learning improves official statistics but needs rigorous validation.
Modern health data science applications leverage abundant molecular and electronic health data, providing opportunities for machine learning to build statistical models to support clinical practice. Time-to-event analysis, also called survival analysis, stands as one of the most representative examples of such statisti…
The paper uses statistics to improve the explainability of models.
Recent work has developed Bayesian methods for the automatic statistical analysis and description of single time series as well as of homogeneous sets of time series data. We extend prior work to create an interpretable kernel embedding for heterogeneous time series. Our method adds practically no computational cost co…
Proposes a framework for automated radiation therapy treatment planning with uncertainty quantification.
One approach to improving the running time of kernel-based machine learning methods is to build a small sketch of the input and use it in lieu of the full kernel matrix in the machine learning task of interest. Here, we describe a version of this approach that comes with running time guarantees as well as improved guar…
We propose a comprehensive treatment of the leverage effect, i.e. the relationship between returns and volatility of a specific asset, focusing on energy commodities futures, namely Brent and WTI crude oils, natural gas and heating oil. After estimating the volatility process without assuming any specific form of its b…
A new robust PCA method uses Innovation Search and Leverage Scores.
Efficiently learns private models using public data.
This article introduces a framework to estimate the value of evidence-based decision making.
In this paper, we propose a fast surrogate leverage weighted sampling strategy to generate refined random Fourier features for kernel approximation. Compared to the current state-of-the-art method that uses the leverage weighted scheme [Li-ICML2019], our new strategy is simpler and more effective. It uses kernel alignm…
aLTT selects hyperparameters efficiently with statistical guarantees.
RealStats detects fake images rigorously, combining multiple detectors for robustness.
New DP framework using data truncation for efficient estimation.
In this work we use Recurrent Neural Networks and Multilayer Perceptrons to predict NYSE, NASDAQ and AMEX stock prices from historical data. We experiment with different architectures and compare data normalization techniques. Then, we leverage those findings to question the efficient-market hypothesis through a formal…
NeurT-FDR controls FDR by incorporating feature hierarchy.
A new method for predicting insurance claims with statistical guarantees.
In this paper, we consider a statistical problem of learning a linear model from noisy samples. Existing work has focused on approximating the least squares solution by using leverage-based scores as an importance sampling distribution. However, no finite sample statistical guarantees and no computationally efficient o…
Automated rock fragmentation assessment using deep learning and spatial statistics.
The paper explores the tradeoffs between fairness measures in machine learning.
The paper tackles robust policy learning in MDPs using statistical methods.
Study improves interpretability in generative models by disentangling latent variables in scientific datasets.
Adaptive methods learn from multiple datasets, leveraging similarities and robust to outliers.
We present a novel family of deep neural architectures, named partially exchangeable networks (PENs) that leverage probabilistic symmetries. By design, PENs are invariant to block-switch transformations, which characterize the partial exchangeability properties of conditionally Markovian processes. Moreover, we show th…
Range penalization enhances statistical accuracy and resource efficiency in federated learning.