A new method detects interactions in machine learning models.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The paper contains a short review of techniques examining regional wealth inequalities based on recently published research work but is also presenting unpublished features. The data pertains to Italy (IT), over the period 2007-2011: the number of cities in regions, the number of inhabitants in cities and in regions, a…
We use PDPs with confidence bands to explain HPO results.
This paper reviews and advocates against the use of permute-and-predict (PaP) methods for interpreting black box functions. Methods such as the variable importance measures proposed for random forests, partial dependence plots, and individual conditional expectation plots remain popular because they are both model-agno…
Interpreting a nonparametric regression model with many predictors is known to be a challenging problem. There has been renewed interest in this topic due to the extensive use of machine learning algorithms and the difficulty in understanding and explaining their input-output relationships. This paper develops a unifie…
RAMs improve GAMs' accuracy by fitting components to subregions of feature space.
Develops Austen plots for assessing bias from unobserved confounding in observational studies.
Hollow-tree Super resolves feature importance in large datasets.
PLOT uses optimal transport to find neural site handles for causal abstraction.
Study robustness of global feature effect explanations in machine learning models.
Computer vision model automates residual plot assessment for diagnosing model assumptions.
GADGET framework decomposes global feature effects using recursive partitioning.
Study shows cliff-learning in transfer learning from foundation models.
This paper compares imputation and direct parameter estimation methods for missing data in correlation matrix visualization.
Proposes a new method for interpreting feature importance and effects in dependent feature models.
Automate residual plot assessment with R package and Shiny application
This research optimizes Andrews plots for better visual clarity in high-dimensional data.
Develops exact and invariant study-based decompositions for network meta-analysis.
3D filament plots visualize curves in datasets, avoiding visual clutter.
New approach predicts stock price synchronization using RNNs and LSTMs.
One of the most popular approaches to understanding feature effects of modern black box machine learning models are partial dependence plots (PDP). These plots are easy to understand but only able to visualize low order dependencies. The paper is about the question 'How much can we see?': A framework is developed to qu…
Modeling intraday electricity prices with a Hawkes process.
In practical applications of machine learning, it is necessary to look beyond standard metrics such as test accuracy in order to validate various qualitative properties of a model. Partial dependence plots (PDP), including instance-specific PDPs (i.e., ICE plots), have been widely used as a visual tool to understand or…
Used to investigate the presence of distinctive recurrent behaviours in natural processes, the recurrence plots can be applied to the analysis of economic data, and, in particular, to the characterization of exchange rates of currencies too. In this paper, we will show that these plots are able to characterize the peri…
One aim of data mining is the identification of interesting structures in data. For better analytical results, the basic properties of an empirical distribution, such as skewness and eventual clipping, i.e. hard limits in value ranges, need to be assessed. Of particular interest is the question of whether the data orig…
Pareto distributions, and power laws in general, have demonstrated to be very useful models to describe very different phenomena, from physics to finance. In recent years, the econophysical literature has proposed a large amount of papers and models justifying the presence of power laws in economic data. Most of the ti…
Paper exposes vulnerabilities in interpreting machine learning models using adversarial attacks on PD plots.
Introduces PIT-plot for prioritizing projects based on their impact.
Charts are an excellent way to convey patterns and trends in data, but they do not facilitate further modeling of the data or close inspection of individual data points. We present a fully automated system for extracting the numerical values of data points from images of scatter plots. We use deep learning techniques t…
New methods for assessing and visualizing feature groups in machine learning models.
Plots show miscalibration directly as slopes of secant lines.
Recent years have witnessed an increased focus on interpretability and the use of machine learning to inform policy analysis and decision making. This paper applies machine learning to examine travel behavior and, in particular, on modeling changes in travel modes when individuals are presented with a novel (on-demand)…
Study predicts synchronization state of financial time series using cross-recurrence plots.
Deep Learning has managed to push boundaries in a wide variety of tasks. One area of interest is to tackle problems in reasoning and understanding, with an aim to emulate human intelligence. In this work, we describe a deep learning model that addresses the reasoning task of question-answering on categorical plots. We …
Herein, we applied statistical physics to study incomes of three (low-, medium- and high-income) society classes instead of the two (low- and medium-income)classes studied so far. In the frame of the threshold nonlinear Langevin dynamics and its threshold Fokker-Planck counterpart, we derived a unified formula for desc…
KM-GPT automates IPD reconstruction from KM plots with high accuracy and scalability.
The size distribution of land plots is a result of land allocation processes in the past. In the absence of regulation this is a Markov process leading an equilibrium described by a probabilistic equation used commonly in the insurance and financial mathematics. We support this claim by analyzing the distribution of tw…
Application of K-Means algorithm is restricted by the fact that the number of clusters should be known beforehand. Previously suggested methods to solve this problem are either ad hoc or require parametric assumptions and complicated calculations. The proposed method aims to solve this conundrum by considering cluster …
New approach interprets machine learning models through feature space transformations.
This text discusses several popular explanatory methods that go beyond the error measurements and plots traditionally used to assess machine learning models. Some of the explanatory methods are accepted tools of the trade while others are rigorously derived and backed by long-standing theory. The methods, decision tree…
Visualizes classification accuracy and label bias in neural nets and trees.
A new tree-based estimator, FastPD, efficiently estimates PD functions for machine learning models.
Study compares Bitcoin and Ethereum tail behavior using Q-Q plots.
Recently, our group has published two papers that have received some attention in the finance community. One is about the profitability of trend following strategies over 200 years, the second is about the correlation between the profitability of "Risk Premia" and their skewness. In this short note, we present two addi…
We propose a procedure for supervised classification that is based on potential functions. The potential of a class is defined as a kernel density estimate multiplied by the class's prior probability. The method transforms the data to a potential-potential (pot-pot) plot, where each data point is mapped to a vector of …
According to the Loss Distribution Approach, the operational risk of a bank is determined as 99.9% quantile of the respective loss distribution, covering unexpected severe events. The 99.9% quantile can be considered a tail event. As supported by the Pickands-Balkema-de Haan Theorem, tail events exceeding some high thr…
A phase plot of the oil economy is built using the literature data of world oil production, price, and EROEI (Energy Returned on Energy Invested). An analogy between the oil economy and the Benard convection is proposed; some methods of interpretation and forecast of the system behavior are also shown based on "phase p…
Study tests rough fractional volatility model across different time scales, revealing new volatility patterns.