Study predicts synchronization state of financial time series using cross-recurrence plots.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New approach predicts stock price synchronization using RNNs and LSTMs.
Adversarial audio attacks can be considered as a small perturbation unperceptive to human ears that is intentionally added to the audio signal and causes a machine learning model to make mistakes. This poses a security concern about the safety of machine learning models since the adversarial attacks can fool such model…
Automate residual plot assessment with R package and Shiny application
This research optimizes Andrews plots for better visual clarity in high-dimensional data.
3D filament plots visualize curves in datasets, avoiding visual clutter.
In practical applications of machine learning, it is necessary to look beyond standard metrics such as test accuracy in order to validate various qualitative properties of a model. Partial dependence plots (PDP), including instance-specific PDPs (i.e., ICE plots), have been widely used as a visual tool to understand or…
Used to investigate the presence of distinctive recurrent behaviours in natural processes, the recurrence plots can be applied to the analysis of economic data, and, in particular, to the characterization of exchange rates of currencies too. In this paper, we will show that these plots are able to characterize the peri…
One aim of data mining is the identification of interesting structures in data. For better analytical results, the basic properties of an empirical distribution, such as skewness and eventual clipping, i.e. hard limits in value ranges, need to be assessed. Of particular interest is the question of whether the data orig…
Pareto distributions, and power laws in general, have demonstrated to be very useful models to describe very different phenomena, from physics to finance. In recent years, the econophysical literature has proposed a large amount of papers and models justifying the presence of power laws in economic data. Most of the ti…
Paper exposes vulnerabilities in interpreting machine learning models using adversarial attacks on PD plots.
Introduces PIT-plot for prioritizing projects based on their impact.
Charts are an excellent way to convey patterns and trends in data, but they do not facilitate further modeling of the data or close inspection of individual data points. We present a fully automated system for extracting the numerical values of data points from images of scatter plots. We use deep learning techniques t…
PLOT uses optimal transport to find neural site handles for causal abstraction.
Interpreting a nonparametric regression model with many predictors is known to be a challenging problem. There has been renewed interest in this topic due to the extensive use of machine learning algorithms and the difficulty in understanding and explaining their input-output relationships. This paper develops a unifie…
Plots show miscalibration directly as slopes of secant lines.
Computer vision model automates residual plot assessment for diagnosing model assumptions.
A new method detects interactions in machine learning models.
Deep Learning has managed to push boundaries in a wide variety of tasks. One area of interest is to tackle problems in reasoning and understanding, with an aim to emulate human intelligence. In this work, we describe a deep learning model that addresses the reasoning task of question-answering on categorical plots. We …
KM-GPT automates IPD reconstruction from KM plots with high accuracy and scalability.
The size distribution of land plots is a result of land allocation processes in the past. In the absence of regulation this is a Markov process leading an equilibrium described by a probabilistic equation used commonly in the insurance and financial mathematics. We support this claim by analyzing the distribution of tw…
The paper contains a short review of techniques examining regional wealth inequalities based on recently published research work but is also presenting unpublished features. The data pertains to Italy (IT), over the period 2007-2011: the number of cities in regions, the number of inhabitants in cities and in regions, a…
It is a truth universally acknowledged that an observed association without known mechanism must be in want of a causal estimate. However, causal estimation from observational data often relies on the (untestable) assumption of `no unobserved confounding'. Violations of this assumption can induce bias in effect estimat…
This text discusses several popular explanatory methods that go beyond the error measurements and plots traditionally used to assess machine learning models. Some of the explanatory methods are accepted tools of the trade while others are rigorously derived and backed by long-standing theory. The methods, decision tree…
This paper compares imputation and direct parameter estimation methods for missing data in correlation matrix visualization.
Visualizes classification accuracy and label bias in neural nets and trees.
Study compares Bitcoin and Ethereum tail behavior using Q-Q plots.
Recently, our group has published two papers that have received some attention in the finance community. One is about the profitability of trend following strategies over 200 years, the second is about the correlation between the profitability of "Risk Premia" and their skewness. In this short note, we present two addi…
We propose a procedure for supervised classification that is based on potential functions. The potential of a class is defined as a kernel density estimate multiplied by the class's prior probability. The method transforms the data to a potential-potential (pot-pot) plot, where each data point is mapped to a vector of …
According to the Loss Distribution Approach, the operational risk of a bank is determined as 99.9% quantile of the respective loss distribution, covering unexpected severe events. The 99.9% quantile can be considered a tail event. As supported by the Pickands-Balkema-de Haan Theorem, tail events exceeding some high thr…
Develops exact and invariant study-based decompositions for network meta-analysis.
A phase plot of the oil economy is built using the literature data of world oil production, price, and EROEI (Energy Returned on Energy Invested). An analogy between the oil economy and the Benard convection is proposed; some methods of interpretation and forecast of the system behavior are also shown based on "phase p…
Study tests rough fractional volatility model across different time scales, revealing new volatility patterns.
SHAP explains boosted trees with additively modeled features.
Modeling intraday electricity prices with a Hawkes process.
In fixed income sector, the yield curve is probably the most observed indicator by the market for trading and fifinancing purposes. A yield curve plots interest rates across different contract maturities from short end to as long as 30 years. For each currency, the corresponding curve shows the relation between the lev…
CDPs visualize causal dependencies in AI models.
Phase plotting is a useful way of visualising functions on complex space. We reinvent the method in the context of hyperbolic geometry, and we use it to plot functions on various representative surfaces for hyperbolic space, illustrating with direct motions in particular. The reinvention is nontrivial, and we discuss t…
Recent years have witnessed an increased focus on interpretability and the use of machine learning to inform policy analysis and decision making. This paper applies machine learning to examine travel behavior and, in particular, on modeling changes in travel modes when individuals are presented with a novel (on-demand)…
There are many statistical tests that verify the null hypothesis: the variable of interest has the same distribution among k-groups. But once the null hypothesis is rejected, how to present the structure of dissimilarity between groups? In this article, we introduce The Merging Path Plot - a methodology, and factorMerg…
Researchers formalize PD and PFI to relate them to data generating process.
An analytic process is iterative between two agents, an analyst and an analytic toolbox. Each iteration comprises three main steps: preparing a dataset, running an analytic tool, and evaluating the result, where dataset preparation and result evaluation, conducted by the analyst, are largely domain-knowledge driven. In…
We use PDPs with confidence bands to explain HPO results.
One of the most popular approaches to understanding feature effects of modern black box machine learning models are partial dependence plots (PDP). These plots are easy to understand but only able to visualize low order dependencies. The paper is about the question 'How much can we see?': A framework is developed to qu…
As an example of the transitions between some of the eight geometries of Thurston, investigated before, we study the geometries supported by the cone-manifolds obtained by surgery on the trefoil knot with singular set the core of the surgery. The geometric structures are explicitly constructed. The most interesting phe…
This project explores several Machine Learning methods to predict movie genres based on plot summaries. Naive Bayes, Word2Vec+XGBoost and Recurrent Neural Networks are used for text classification, while K-binary transformation, rank method and probabilistic classification with learned probability threshold are employe…
A new hypersurface of Tzitzeica type is obtained in all three forms: parametric, implicit and explicit. Its two-dimensional version, although well-known from a theoretical point of view, is plotted with Matlab.
Extends multidimensional scaling to analyze three-way asymmetric proximities.