Solves a 60-year-old question on agreement measures in statistics.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
An explicit solution found for maximizing/minimizing agreement in a 2x2 table.
This work introduces significativity indices for agreement values between classifiers.
MAS scores cluster size consistency from points, robust to label changes.
Formula found for minimum ARI between clusterings of fixed sizes.
Study finds simple model-agreement scores perform well in various error estimation scenarios.
In unsupervised machine learning, agreement between partitions is commonly assessed with so-called external validity indices. Researchers tend to use and report indices that quantify agreement between two partitions for all clusters simultaneously. Commonly used examples are the Rand index and the adjusted Rand index. …
LFD method improves text classification by making features clearer and less label-leaking.
A new method for combining multiple data views in supervised learning.
Within the natural language processing (NLP) community, active learning has been widely investigated and applied in order to alleviate the annotation bottleneck faced by developers of new NLP systems and technologies. This paper presents the first theoretical analysis of stopping active learning based on stabilizing pr…
New index improves anomaly detection in correlated time series data.
Unified framework for comparing clusterings from information-theoretic and pair-counting perspectives.
Study confirms eurozone interbank market stability but finds higher collateral reuse.
Real life hedging in the Black-Scholes model must be imperfect and if the stock's drift is higher than the risk free rate, leads to a profit on average. Hence the option price is examined as a fair game agreement between the parties, based on expected payoffs and a simple measure of risk. The resulting prices result in…
We investigated distributions of short term price trends for high frequency stock market data. A number of trends as a function of their lengths was measured. We found that such a distribution does not fit to results following from an uncorrelated stochastic process. We proposed a simple model with a memory that gives …
We propose a projected gradient dynamical system as a model for a bargaining scheme for an asset for which the two interested agents have personal valuations which do not initially coincide. The personal valuations are formed using subjective beliefs concerning the future states of the world and the reservation prices …
SNAP improves robust computation by emphasizing trustworthy items and downweighting outliers.
We apply random matrix theory to compare correlation matrix estimators C obtained from emerging market data. The correlation matrices are constructed from 10 years of daily data for stocks listed on the Johannesburg Stock Exchange (JSE) from January 1993 to December 2002. We test the spectral properties of C against ra…
ValueBlindBench tests LLM-generated investment rationales for validity before returns are known.
The operator realizing a Dehn twist in quantum Teichmuller theory is diagonalized and continuous spectrum is obtained. This result is in agreement with the expected spectrum of conformal weights in quantum Liouville theory at c>1. The completeness condition of the eigenvectors includes the integration measure which app…
We show that the behaviour of Bitcoin has interesting similarities to stock and precious metal markets, such as gold and silver. We report that whilst Litecoin, the second largest cryptocurrency, closely follows Bitcoin's behaviour, it does not show all the reported properties of Bitcoin. Agreements between apparently …
Community moderation drifts towards majority, study finds.
Motivation: Untargeted metabolomics comprehensively characterizes small molecules and elucidates activities of biochemical pathways within a biological sample. Despite computational advances, interpreting collected measurements and determining their biological role remains a challenge. Results: To interpret measurement…
Discriminatory trade liberalization policies are becoming more popular among world economies. Countries are motivated to enter for regional trade agreements to capture faster economic growth for alleviating poverty. In developing economies like most of the member countries of the Association of South East Asian Nations…
In this article we study the dependence degree of the traded volume of the Dow Jones 30 constituent equities by using a nonextensive generalised form of the Kullback-Leibler information measure. Our results show a slow decay of the dependence degree as a function of the lag. This feature is compatible with the existenc…
The paper analyzes MENA region's energy consumption and policy needs for renewable energy.
The paper proposes a framework for information-theoretic predictive uncertainty measures.
Haircutting non-cash collateral has become a key element of the post-crisis reform of the shadow banking system and OTC derivatives markets. This article develops a parametric haircut model by expanding haircut definitions beyond the traditional value-at-risk measure and employing a double-exponential jump-diffusion mo…
We introduce the formalism of generalized Fourier transforms in the context of risk management. We develop a general framework to efficiently compute the most popular risk measures, Value-at-Risk and Expected Shortfall (also known as Conditional Value-at-Risk). The only ingredient required by our approach is the knowle…
Under Solvency II the computation of capital requirements is based on value at risk (V@R). V@R is a quantile-based risk measure and neglects extreme risks in the tail. V@R belongs to the family of distortion risk measures. A serious deficiency of V@R is that firms can hide their total downside risk in corporate network…
Novel method uses information theory to measure causal influences during transient neural events.
In many machine learning problems, labeled training data is limited but unlabeled data is ample. Some of these problems have instances that can be factored into multiple views, each of which is nearly sufficent in determining the correct labels. In this paper we present a new algorithm for probabilistic multi-view lear…
In Bipartite Correlation Clustering (BCC) we are given a complete bipartite graph with `+' and `-' edges, and we seek a vertex clustering that maximizes the number of agreements: the number of all `+' edges within clusters plus all `-' edges cut across clusters. BCC is known to be NP-hard. We present a novel approx…
We introduce a technique based on the singular vector canonical correlation analysis (SVCCA) for measuring the generality of neural network layers across a continuously-parametrized set of tasks. We illustrate this method by studying generality in neural networks trained to solve parametrized boundary value problems ba…
Model selection is a problem that has occupied machine learning researchers for a long time. Recently, its importance has become evident through applications in deep learning. We propose an agreement-based learning framework that prevents many of the pitfalls associated with model selection. It relies on coupling the t…
Complex systems are composed of mutually interacting components and the output values of these components are usually long-range cross-correlated. We propose a method to characterize the joint multifractal nature of such long-range cross correlations based on wavelet analysis, termed multifractal cross wavelet analysis…
Machine learning predicts greenhouse gas emissions for undisclosed companies.
New algorithms handle phase retrieval with rank d measurements, revealing phase transitions.
Computable contracts simplify financial transactions and reduce legal costs.
Individual's semantics have been used for guiding the learning process of Genetic Programming solving supervised learning problems. The semantics has been used to proposed novel genetic operators as well as different ways of performing parent selection. The latter is the focus of this contribution by proposing three he…
Transformer models improve financial sentiment measurement.
Generalization and reliability of multilingual translation often highly depend on the amount of available parallel data for each language pair of interest. In this paper, we focus on zero-shot generalization---a challenging setup that tests models on translation directions they have not been optimized for at training t…
We apply an asymmetric version of Kirman's herding model to volatile financial markets. In the relation between returns and agent concentration we use the square root law proposed by Zhang. This can be derived by extending the idea of a critical mean field theory suggested by Plerou et al. We show that this model is eq…
New method estimates robust multi-period portfolios using entropy.
We consider a simple stochastic model of a urban rental housing market, in which the interaction of tenants and landlords induces rent fluctuations. We simulate the model numerically and measure the equilibrium rent distribution, which is found to be close to a lognormal law. We also study the influence of the density …
Researchers develop multi-agent systems for quadcopters to collaborate in missions.
It is generally recognized that economical systems, and more in general complex systems, are characterized by power law distributions. Sometime, these distributions show a changing of the slope in the tail so that, more appropriately, they show a multi-power law behavior. We present a method to derive analytically a tw…
Ensemble learning is a powerful approach to construct a strong learner from multiple base learners. The most popular way to aggregate an ensemble of classifiers is majority voting, which assigns a sample to the class that most base classifiers vote for. However, improved performance can be obtained by assigning weights…