This report is an introduction to mathematical map colouring and the problems posed by Heawood in his paper of 1890. There will be a brief discussion of the Map Colour Theorem; then we will move towards investigating empire maps in the plane and the recent contributions by Wessel. Finally we will conclude with a discus…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Investigates model selection challenges in heterogeneous treatment effect estimation.
Paper explores using bootstrap methods to improve SGD's stability and robustness.
Study finds LLMs hallucinate in finance tasks, needing research.
Explaining how overparametrized neural networks simultaneously achieve low risk and zero empirical risk on benchmark datasets is an open problem. PAC-Bayes bounds optimized using variational inference (VI) have been recently proposed as a promising direction in obtaining non-vacuous bounds. We show empirically that thi…
We investigate the historical volatility of the 100 most capitalized stocks traded in US equity markets. An empirical probability density function (pdf) of volatility is obtained and compared with the theoretical predictions of a lognormal model and of the Hull and White model. The lognormal model well describes the pd…
The recently introduced dropout training criterion for neural networks has been the subject of much attention due to its simplicity and remarkable effectiveness as a regularizer, as well as its interpretation as a training procedure for an exponentially large ensemble of networks that share parameters. In this work we …
Investigates financial and economic systems using statistical mechanics and information theory.
We present a large-scale empirical study of catastrophic forgetting (CF) in modern Deep Neural Network (DNN) models that perform sequential (or: incremental) learning. A new experimental protocol is proposed that enforces typical constraints encountered in application scenarios. As the investigation is empirical, we ev…
Most positive and unlabeled data is subject to selection biases. The labeled examples can, for example, be selected from the positive set because they are easier to obtain or more obviously positive. This paper investigates how learning can be ena BHbled in this setting. We propose and theoretically analyze an empirica…
Bayesian neural networks can be partially stochastic without losing predictive power.
New research shows common ID estimators in neural representations are inaccurate.
Study investigates key design choices in on-policy RL algorithms.
We investigate the possible drawbacks of employing the standard Pearson estimator to measure correlation coefficients between financial stocks in the presence of non-stationary behavior, and we provide empirical evidence against the well-established common knowledge that using longer price time series provides better, …
Paper assesses error estimates of Random Forests classification.
Study investigates one-shot semi-supervised learning for image classification.
In many applications, multivariate samples may harbor previously unrecognized heterogeneity at the level of conditional independence or network structure. For example, in cancer biology, disease subtypes may differ with respect to subtype-specific interplay between molecular components. Then, both subtype discovery and…
An empirical analysis of interest rates in money and capital markets is performed. We investigate a set of 34 different weekly interest rate time series during a time period of 16 years between 1982 and 1997. Our study is focused on the collective behavior of the stochastic fluctuations of these time-series which is in…
In this study, the effects of eight representation regularization methods are investigated, including two newly developed rank regularizers (RR). The investigation shows that the statistical characteristics of representations such as correlation, sparsity, and rank can be manipulated as intended, during training. Furth…
Empirical median performs well in estimating location with varying scales.
We highlight a very simple statistical tool for the analysis of financial bubbles, which has already been studied in [1]. We provide extensive empirical tests of this statistical tool and investigate analytically its link with stocks correlation structure.
Empirical study finds variance swap rate is affine in spot variance for S&P500 data.
The paper has 2 main goals: 1. We propose a variant of the CAPM based on coherent risk. 2. In addition to the real-world measure and the risk-neutral measure, we propose the third one: the extreme measure. The introduction of this measure provides a powerful tool for investigating the relation between the first two mea…
We consider different levels of complexity which are observed in the empirical investigation of financial time series. We discuss recent empirical and theoretical work showing that statistical properties of financial time series are rather complex under several ways. Specifically, they are complex with respect to their…
We use a large census of hyperbolic 3-manifolds to experimentally investigate a conjecture of Neumann regarding the Bloch Group. We present an augmented census including, for feasible invariant trace fields, explicit manifolds (associated to that field) that appear to generate the Bloch group of that field. We also mak…
We empirically investigate distributions of individual consumption expenditure f or four commodity categories conditional on fixed income levels. The data stems from the Family Expenditure Survey carried out annually in the United Kingdom. W e use graphical techniques to test for normality and lognormality of these dis…
A new framework for dimension reduction using ensemble of random projections.
Investigates how multivariate Lévy models affect calibration and pricing.
Online social networks offer a new way to investigate financial markets' dynamics by enabling the large-scale analysis of investors' collective behavior. We provide empirical evidence that suggests social media and stock markets have a nonlinear causal relationship. We take advantage of an extensive data set composed o…
Neural Empirical Bayes estimates source distributions from noisy simulations.
This work investigates fundamental questions related to learning features in convolutional neural networks (CNN). Empirical findings across multiple architectures such as VGG, ResNet, Inception, DenseNet and MobileNet indicate that weights near the center of a filter are larger than weights on the outside. Current regu…
All too often measuring statistical dependencies between financial time series is reduced to a linear correlation coefficient. However this may not capture all facets of reality. We study empirical dependencies of daily stock returns by their pairwise copulas. Here we investigate particularly to which extent the non-st…
In this preliminary work, we study the generalization properties of infinite ensembles of infinitely-wide neural networks. Amazingly, this model family admits tractable calculations for many information-theoretic quantities. We report analytical and empirical investigations in the search for signals that correlate with…
ROI-driven data analytics guides investment in empirical data analysis.
Study investigates asymptotic risk of overparameterized models, including deep neural networks.
The study investigates the consistency of -means clustering under finite expectation assumptions.
We consider a financial market model which consists of a financial asset and a large number of interacting agents classified into many types. Different types of agents are heterogeneous in their price expectations. Each agent can change its type based on the current empirical distribution of the types and the equilibri…
In this preregistration submission, we propose an empirical study of how networks handle changes in complexity of the data. We investigate the effect of network capacity on generalization performance in the face of increasing data complexity. For this, we measure the generalization error for an image classification tas…
The paper investigates learning conditional distributions on multi-dimensional spaces using clustering and neural networks.
Deep learning is finding its way into the embedded world with applications such as autonomous driving, smart sensors and aug- mented reality. However, the computation of deep neural networks is demanding in energy, compute power and memory. Various approaches have been investigated to reduce the necessary resources, on…
We study some properties of eigenvalue spectra of financial correlation matrices. In particular, we investigate the nature of the large eigenvalue bulks which are observed empirically, and which have often been regarded as a consequence of the supposedly large amount of noise contained in financial data. We challenge t…
In this work we present the novel ASTRID method for investigating which attribute interactions classifiers exploit when making predictions. Attribute interactions in classification tasks mean that two or more attributes together provide stronger evidence for a particular class label. Knowledge of such interactions make…
The paper improves semi-supervised learning using -divergences and -Rényi divergences.
Multiplicative random cascade model naturally reproduces the intermittency or multifractality, which is frequently shown among hierarchical complex systems such as turbulence and financial markets. As described herein, we investigate the validity of a multiplicative hierarchical random cascade model through an empirica…
New framework explains how larger pre-trained models reduce downstream learning sample complexity.
We investigate the problem of semi-parametric maximum likelihood under constraints on summary statistics. Such a procedure results in a discrete probability distribution that maximises the likelihood among all such distributions under the specified constraints (called estimating equations), and is an approximation to t…
While defaults are rare events, losses can be substantial even for credit portfolios with a large number of contracts. Therefore, not only a good evaluation of the probability of default is crucial, but also the severity of losses needs to be estimated. The recovery rate is often modeled independently with regard to th…
Study predicts NFT bubbles using LPPL model.