In Bipartite Correlation Clustering (BCC) we are given a complete bipartite graph with `+' and `-' edges, and we seek a vertex clustering that maximizes the number of agreements: the number of all `+' edges within clusters plus all `-' edges cut across clusters. BCC is known to be NP-hard. We present a novel approx…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We apply random matrix theory to compare correlation matrix estimators C obtained from emerging market data. The correlation matrices are constructed from 10 years of daily data for stocks listed on the Johannesburg Stock Exchange (JSE) from January 1993 to December 2002. We test the spectral properties of C against ra…
New index improves anomaly detection in correlated time series data.
This work introduces significativity indices for agreement values between classifiers.
Unified framework for policy learning using weak supervision.
We conduct an empirical study using the quantile-based correlation function to uncover the temporal dependencies in financial time series. The study uses intraday data for the S\&P 500 stocks from the New York Stock Exchange. After establishing an empirical overview we compare the quantile-based correlation function to…
This paper uses rank correlation methods to construct MSTs from financial returns, finding them more stable and robust.
Signatures of universality are detected by comparing individual eigenvalue distributions and level spacings from financial covariance matrices to random matrix predictions. A chopping procedure is devised in order to produce a statistical ensemble of asset-price covariances from a single instance of financial data sets…
The study reveals how synaptic correlations promote dimension reduction in neural networks.
Financial time series exhibit two different type of non linear correlations: (i) volatility autocorrelations that have a very long range memory, on the order of years, and (ii) asymmetric return-volatility (or `leverage') correlations that are much shorter ranged. Different stochastic volatility models have been propos…
We investigate relaxation and correlations in a class of mean-reverting models for stochastic variances. We derive closed-form expressions for the correlation functions and leverage for a general form of the stochastic term. We also discuss correlation functions and leverage for three specific models -- multiplicative,…
Through simple analytical calculations and numerical simulations, we demonstrate the generic existence of a self-organized macroscopic state in any large multivariate system possessing non-vanishing average correlations between a finite fraction of all pairs of elements. The coexistence of an eigenvalue spectrum predic…
Complex systems are composed of mutually interacting components and the output values of these components are usually long-range cross-correlated. We propose a method to characterize the joint multifractal nature of such long-range cross correlations based on wavelet analysis, termed multifractal cross wavelet analysis…
Solves a 60-year-old question on agreement measures in statistics.
We consider random vectors drawn from a multivariate normal distribution and compute the sample statistics in the presence of non-stationary correlations. For this purpose, we construct an ensemble of random correlation matrices and average the normal distribution over this ensemble. The resulting distribution contains…
Multi-view spectral clustering, which aims at yielding an agreement or consensus data objects grouping across multi-views with their graph laplacian matrices, is a fundamental clustering problem. Among the existing methods, Low-Rank Representation (LRR) based method is quite superior in terms of its effectiveness, intu…
We investigate the general problem of how to model the kinematics of stock prices without considering the dynamical causes of motion. We propose a stochastic process with long-range correlated absolute returns. We find that the model is able to reproduce the experimentally observed clustering, power law memory, fat tai…
An explicit solution found for maximizing/minimizing agreement in a 2x2 table.
We define a random-matrix ensemble given by the infinite-time covariance matrices of Ornstein-Uhlenbeck processes at different temperatures coupled by a Gaussian symmetric matrix. The spectral properties of this ensemble are shown to be in qualitative agreement with some stylized facts of financial markets. Through the…
Study of correlated Wigner matrices with BBP transitions.
Among the proposed network models, the hidden variable (or good get richer) one is particularly interesting, even if an explicit empirical test of its hypotheses has not yet been performed on a real network. Here we provide the first empirical test of this mechanism on the world trade web, the network defined by the tr…
The purpose of this paper is introducing rigorous methods and formulas for bilateral counterparty risk credit valuation adjustments (CVA's) on interest-rate portfolios. In doing so, we summarize the general arbitrage-free valuation framework for counterparty risk adjustments in presence of bilateral default risk, as de…
We analyze the sequence of time intervals between consecutive stock trades of thirty companies representing eight sectors of the U. S. economy over a period of four years. For all companies we find that: (i) the probability density function of intertrade times may be fit by a Weibull distribution; (ii) when appropriate…
We study properties of the cross-sectional distribution of returns. A significant anti-correlation between dispersion and cross-sectional kurtosis is found such that dispersion is high but kurtosis is low in panic times, and the opposite in normal times. The co-movement of stock returns also increases in panic times. W…
The study of record statistics of correlated series is gaining momentum. In this work, we study the records statistics of the time series of select stock market data and the geometric random walk, primarily through simulations. We show that the distribution of the age of records is a power law with the exponent lyi…
Python tool detects economic crises from S&P500 correlation data.
PPM improves graph matching for correlated Gaussian Wigner models with high probability.
ExCIR provides efficient, consistent, and scalable explainability for complex models.
We show that the cost of market orders and the profit of infinitesimal market-making or -taking strategies can be expressed in terms of directly observable quantities, namely the spread and the lag-dependent impact function. Imposing that any market taking or liquidity providing strategies is at best marginally profita…
SNAP improves robust computation by emphasizing trustworthy items and downweighting outliers.
We consider a mean-reverting stochastic volatility model which satisfies some relevant stylized facts of financial markets. We introduce an algorithm for the detection of peaks in the volatility profile, that we apply to the time series of Dow Jones Industrial Average and Financial Times Stock Exchange 100 in the perio…
An automated metric to evaluate dialogue quality is vital for optimizing data driven dialogue management. The common approach of relying on explicit user feedback during a conversation is intrusive and sparse. Current models to estimate user satisfaction use limited feature sets and rely on annotation schemes with low …
We analyze cross-correlations between price fluctuations of different stocks using methods of random matrix theory (RMT). Using two large databases, we calculate cross-correlation matrices C of returns constructed from (i) 30-min returns of 1000 US stocks for the 2-yr period 1994--95 (ii) 30-min returns of 881 US stock…
Our work sheds new light on the role of oil prices in shaping the world economy by investigating flows of goods and services through global value chains between 1960 and 2011, by means of Markov Chain and network analysis. We show that over that time period the international division of labor and trade patterns are tig…
We study the volatility of the MIB30-stock-index high-frequency data from November 28, 1994 through September 15, 1995. Our aim is to empirically characterize the volatility random walk in the framework of continuous-time finance. To this end, we compute the index volatility by means of the log-return standard deviatio…
Individual's semantics have been used for guiding the learning process of Genetic Programming solving supervised learning problems. The semantics has been used to proposed novel genetic operators as well as different ways of performing parent selection. The latter is the focus of this contribution by proposing three he…
Discriminatory trade liberalization policies are becoming more popular among world economies. Countries are motivated to enter for regional trade agreements to capture faster economic growth for alleviating poverty. In developing economies like most of the member countries of the Association of South East Asian Nations…
We introduce a technique based on the singular vector canonical correlation analysis (SVCCA) for measuring the generality of neural network layers across a continuously-parametrized set of tasks. We illustrate this method by studying generality in neural networks trained to solve parametrized boundary value problems ba…
We study the dependence structure of market states by estimating empirical pairwise copulas of daily stock returns. We consider both original returns, which exhibit time-varying trends and volatilities, as well as locally normalized ones, where the non-stationarity has been removed. The empirical pairwise copula for ea…
LFD method improves text classification by making features clearer and less label-leaking.
New formalism solves kinematical constraints in curved backgrounds and non-trivial states.
Formula found for minimum ARI between clusterings of fixed sizes.
In complex financial systems, the sector structure and volatility clustering are respectively important features of the spatial and temporal correlations. However, the microscopic generation mechanism of the sector structure is not yet understood. Especially, how to produce these two features in one model remains chall…
In unsupervised machine learning, agreement between partitions is commonly assessed with so-called external validity indices. Researchers tend to use and report indices that quantify agreement between two partitions for all clusters simultaneously. Commonly used examples are the Rand index and the adjusted Rand index. …
In many machine learning problems, labeled training data is limited but unlabeled data is ample. Some of these problems have instances that can be factored into multiple views, each of which is nearly sufficent in determining the correct labels. In this paper we present a new algorithm for probabilistic multi-view lear…
The detrending moving average (DMA) algorithm is a widely used technique to quantify the long-term correlations of non-stationary time series and the long-range correlations of fractal surfaces, which contains a parameter determining the position of the detrending window. We develop multifractal detrending moving a…
We address the problem of long-range memory in the financial markets. There are two conceptually different ways to reproduce power-law decay of auto-correlation function: using fractional Brownian motion as well as non-linear stochastic differential equations. In this contribution we address this problem by analyzing e…
Model selection is a problem that has occupied machine learning researchers for a long time. Recently, its importance has become evident through applications in deep learning. We propose an agreement-based learning framework that prevents many of the pitfalls associated with model selection. It relies on coupling the t…