FigureNet learns to answer questions about scientific plots.
problem Addressing reasoning tasks in question-answering on scientific plots.
method Introduces FigureNet, a deep learning model that identifies plot elements, quantifies values, and determines relative ordering.
result FigureNet outperforms state-of-the-art models by 7% on the FigureQA dataset.
cGAP visualizes high-dimensional categorical data with interpretable geometric structure.
problem Lack of visualization tools for high-dimensional categorical data.
method cGAP uses Homogeneity Analysis (HOMALS) to embed data in a 3D space and maps it to colors.
result cGAP reveals coherent clusters, outliers, and local-to-global structure in categorical data.
cGAP visualizes high-dimensional categorical data with interpretable geometric structure.
problem Lack of visualization tools for high-dimensional categorical data.
method cGAP uses Homogeneity Analysis (HOMALS) to embed data in a 3D space and maps it to colors for visualization.
result cGAP reveals coherent clusters, outliers, and local-to-global structure in categorical data.
Automate residual plot assessment with R package and Shiny application
problem Diagnosing linear models
method Computer vision model for residual plot assessment
result Predicts visual signal strength and supports model fit assessment
Earth observation embeddings can convert discrete biome maps into continuous representations that better capture ecological variation.
problem Biome maps impose categorical boundaries that compress continuous variation in biotic communities.
method Fit a linear classifier on Earth observation embeddings to predict biome labels.
result Continuous biome representation outperforms discrete biome labels for predicting species occurrence.
This research optimizes Andrews plots for better visual clarity in high-dimensional data.
problem Visualizing high-dimensional datasets with clarity and aesthetics.
method Developed a method to add spectral smoothing to Andrews plots to reduce visual clutter.
result Optimal spatial-spectral smoothing leads to more aesthetically pleasing and clutter-free visualizations.
Automates selection and visualization of model responses in various directions.
problem Manual selection and limited visualization of PDPs.
method Formalizes method for automating PDP selection and extends to arbitrary directions.
result Demonstrates usefulness across model selection, bias detection, and latent space exploration.
3D filament plots visualize curves in datasets, avoiding visual clutter.
problem Visualizing high-dimensional data relationships in scatterplots of points.
method Construct 3D filament plots using linear isometries and Frenet-Serret systems.
result 3D filament plots preserve Euclidean distances and avoid visual clutter.
A new visualization tool MD plot discovers interesting structures in continuous features.
problem Identifying interesting structures in data distributions, especially with skewed, clipped, or multimodal distributions.
method Proposes a new visualization tool called the mirrored density plot (MD plot) that does not require adjusting density estimation parameters.
result The MD plot outperforms conventional methods in identifying structures in complex distributions.
Used to investigate the presence of distinctive recurrent behaviours in natural processes, the recurrence plots can be applied to the analysis of economic data, and, in particular, to the characterization of exchange rates of currencies too. In this paper, we will show that these plots are able to characterize the peri…
Pareto distributions, and power laws in general, have demonstrated to be very useful models to describe very different phenomena, from physics to finance. In recent years, the econophysical literature has proposed a large amount of papers and models justifying the presence of power laws in economic data. Most of the ti…
Paper exposes vulnerabilities in interpreting machine learning models using adversarial attacks on PD plots.
problem Vulnerability of permutation-based interpretation methods, particularly PD plots, to adversarial attacks.
method Adversarial framework to manipulate black-box models and produce deceptive PD plots.
result It is possible to hide discriminatory behaviors in machine learning models through interpretation tools like PD plots.
Introduces PIT-plot for prioritizing projects based on their impact.
problem Optimizing R&D investments in project portfolios.
method Develops a new tool (PIT-plot) focusing on project impact rather than project properties.
result Identifies projects with the largest impact for risk mitigation or value-adding.
Unified framework for interpreting complex regression models with many predictors.
problem Interpreting nonparametric regression models with many predictors.
method Derivative-based approach for existing tools like partial-dependence plots.
result New technique called accumulated total derivative effects plot for complex models.
Charts are an excellent way to convey patterns and trends in data, but they do not facilitate further modeling of the data or close inspection of individual data points. We present a fully automated system for extracting the numerical values of data points from images of scatter plots. We use deep learning techniques t…
PLOT uses optimal transport to find neural site handles for causal abstraction.
problem Finding the relevant neural site for causal analysis is computationally challenging.
method PLOT employs optimal transport to localize causal variables from neural network outputs.
result PLOT efficiently finds intervention handles for causal abstraction in neural networks.
Develops Austen plots for assessing bias from unobserved confounding in observational studies.
problem Bias in causal estimates due to unobserved confounding.
method Formalizes confounding strength, uses Austen plots to visualize and quantify bias.
result Allows domain experts to assess the plausibility of strong confounders.
Plots show miscalibration directly as slopes of secant lines.
problem Detecting discrepancies between probabilistic predictions and actual outcomes.
method Cumulative differences between observed and expected values displayed as slopes of secant lines.
result Directly shows miscalibration without binning or kernel density estimation.
Automates identifying interesting plots in production yield data.
problem Identifying plots deemed interesting by analysts in production yield data.
method Uses Generative Adversarial Networks (GANs) to learn analyst's intent.
result Demonstrates improved identification of interesting plots in production yield data.
Computer vision model automates residual plot assessment for diagnosing model assumptions.
problem Automating residual plot assessment for model diagnostics.
method Trains a computer vision model to predict disparity between residual distributions and reference distributions using Kullback-Leibler divergence.
result Computer vision model is less sensitive to non-linearity but more sensitive than human judgment and conventional tests.
Study predicts synchronization state of financial time series using cross-recurrence plots.
problem Predicting the state of synchronization of financial time series.
method Cross-correlation analysis and deep learning framework for predicting synchronization state based on cross-recurrence plots.
result Satisfactory performance in predicting synchronization state for certain pairs of stocks.
A new method detects interactions in machine learning models.
problem Interpreting non-linear and interaction effects in machine learning models.
method Regional effect plots with implicit interaction detection.
result The method quantifies and interprets feature effects reliably, less confounded by interactions.
KM-GPT automates IPD reconstruction from KM plots with high accuracy and scalability.
problem Manual digitization of IPD from KM plots is error-prone and lacks scalability.
method KM-GPT integrates advanced image preprocessing, multi-modal reasoning, and iterative reconstruction algorithms.
result KM-GPT generates high-quality IPD without manual input or intervention, achieving superior accuracy.
The size distribution of land plots is a result of land allocation processes in the past. In the absence of regulation this is a Markov process leading an equilibrium described by a probabilistic equation used commonly in the insurance and financial mathematics. We support this claim by analyzing the distribution of tw…
The paper contains a short review of techniques examining regional wealth inequalities based on recently published research work but is also presenting unpublished features. The data pertains to Italy (IT), over the period 2007-2011: the number of cities in regions, the number of inhabitants in cities and in regions, a…
This paper compares imputation and direct parameter estimation methods for missing data in correlation matrix visualization.
problem Missing data challenges in estimating correlation coefficients for accurate visualization.
method Comparison of imputation and direct parameter estimation methods for handling missing data.
result Direct parameter estimation (DPER) outperforms imputation for accurate correlation matrix visualization.
Visualizes classification accuracy and label bias in neural nets and trees.
problem Identifying mislabeled cases in neural net and tree-based classifications.
method Silhouette plots and quasi residual plots of PAC (probability of alternative class).
result Comparison of different classifications using silhouette width and PAC plots.
Visualizes functions on hyperbolic geometry surfaces.
problem Visualizing functions on hyperbolic geometry surfaces.
method Reinvented phase plotting for hyperbolic geometry using conformal maps.
result Illustrates direct motions on geodesics.
Study compares Bitcoin and Ethereum tail behavior using Q-Q plots.
problem Examining tail risk in cryptocurrency returns.
method Used Q-Q plots and Generalized Tempered Stable (GTS) distribution.
result Ethereum shows more extreme values than Bitcoin, indicating greater tail risk.
We propose a procedure for supervised classification that is based on potential functions. The potential of a class is defined as a kernel density estimate multiplied by the class's prior probability. The method transforms the data to a potential-potential (pot-pot) plot, where each data point is mapped to a vector of …
The abstract discusses various machine learning explanation methods.
problem Improving the understanding and trust in machine learning models.
method Exploratory methods for assessing machine learning models.
result Various methods exist, each with its own strengths and applications.
Recently, our group has published two papers that have received some attention in the finance community. One is about the profitability of trend following strategies over 200 years, the second is about the correlation between the profitability of "Risk Premia" and their skewness. In this short note, we present two addi…
Paper quantifies how much machine learning models can be explained.
problem Quantifying the explainability of machine learning models.
method Developed a framework to quantify explainability of arbitrary machine learning models.
result Allows for judgment on sufficiency of model explanations.
According to the Loss Distribution Approach, the operational risk of a bank is determined as 99.9% quantile of the respective loss distribution, covering unexpected severe events. The 99.9% quantile can be considered a tail event. As supported by the Pickands-Balkema-de Haan Theorem, tail events exceeding some high thr…
Explains how machine learning models can be biased and presents interactive plots to visualize bias.
problem Bias in machine learning models can lead to unfair decisions.
method Develops interactive plots to visualize bias in machine learning models.
result Demonstrates that machine learning models can learn and propagate bias from the data.
Develops exact and invariant study-based decompositions for network meta-analysis.
problem Lack of exact contribution decompositions in network meta-analysis.
method Contrast-space projection formulation of NMA, study-based definition of direct and indirect evidence.
result Exact covariance-aware decompositions of NMA estimator into direct and indirect contributions.
A phase plot of the oil economy is built using the literature data of world oil production, price, and EROEI (Energy Returned on Energy Invested). An analogy between the oil economy and the Benard convection is proposed; some methods of interpretation and forecast of the system behavior are also shown based on "phase p…
Study tests rough fractional volatility model across different time scales, revealing new volatility patterns.
problem Testing robustness of rough fractional volatility model over various time scales.
method Used large dataset on FX rates, included smoothing and measurement errors, analyzed log-log plots of realized variance increments.
result Found new stylized facts in volatility patterns, including convexity and nonlinear behavior.
New approach predicts stock price synchronization using RNNs and LSTMs.
problem Forecasting synchronization of stock prices in the Indian market.
method Utilizing recurrence plots and CRQA for non-linear analysis, RNNs and LSTMs for prediction.
result Accuracy of 0.98 and F1 score of 0.83 in predicting stock price synchronization.
Modeling intraday electricity prices with a Hawkes process.
problem Capturing the dynamics of intraday electricity prices, especially microstructure noise.
method 2D marked Hawkes process with increasing baseline intensity, providing analytic moments and signature plot.
result The model fits German intraday electricity data well and converges to a Brownian motion with increasing volatility.
SHAP explains boosted trees with additively modeled features.
problem Explaining predictions of boosted trees models with additively modeled features.
method SHAP values for additively modeled features in boosted trees models.
result SHAP dependence plot matches partial dependence plot for additively modeled features.
CDPs visualize causal dependencies in AI models.
problem Understanding how AI models depend on data inputs causally.
method Developed Causal Dependence Plots (CDPs) to visualize causal dependencies.
result CDPs show causal changes in predictors and outcomes.
There are many statistical tests that verify the null hypothesis: the variable of interest has the same distribution among k-groups. But once the null hypothesis is rejected, how to present the structure of dissimilarity between groups? In this article, we introduce The Merging Path Plot - a methodology, and factorMerg…
Researchers formalize PD and PFI to relate them to data generating process.
problem Lack of theory linking PD and PFI to data generating process.
method Formalize PD and PFI as estimators of ground truth estimands, account for model variance with learner-PD and learner-PFI.
result PD and PFI estimates deviate from ground truth due to statistical biases, model variance, and Monte Carlo approximation errors.
StructureBoost improves gradient boosting for complex categorical variables efficiently.
problem Efficiently handling complex categorical variables with known structure.
method Two methods to overcome computational obstacles in SCDT enumeration for structured categorical variables.
result StructureBoost outperforms existing packages on complex categorical problems.
Categorical bundles provide a natural framework for gauge theories involving multiple gauge groups. Unlike the case of traditional bundles there are distinct notions of triviality, and hence also of local triviality, for categorical bundles. We study categorical principal bundles that are product bundles in the categor…
We use PDPs with confidence bands to explain HPO results.
problem Lack of explainable insights into HPO results.
method Use IML techniques, specifically PDPs, with estimated confidence bands.
result Increased quality of PDPs in relevant sub-regions.
The paper explains the concave shape of yield curves from trading perspectives.
problem Lack of explanation for the concavity of yield curves from economics theory.
method Explains the concavity of yield curves from trading perspectives.
result Offers an explanation for the concave shape of yield curves.