Model-agnostic interpretation methods can mislead if not used carefully.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study pitfalls of deep learning ensembles in uncertainty estimation.
Machine learning suffers from poor design, data, and evaluation practices.
This paper evaluates metrics for graph generative models, addressing common pitfalls.
Machine Learning (ML) and Deep Learning (DL) innovations are being introduced at such a rapid pace that model owners and evaluators are hard-pressed analyzing and studying them. This is exacerbated by the complicated procedures for evaluation. The lack of standard systems and efficient techniques for specifying and pro…
Study identifies pitfalls in assessing hierarchies for multi-class classification.
Benchmark assesses LLMs' causal inference skills, revealing significant limitations.
New categorization of community detection methods to avoid pitfalls.
Bayesian optimization improves molecule design by addressing three pitfalls.
Unified framework for ranking-and-selection with multiple correct answers and non-answerable estimates
This paper improves conditional sampling for VAEs by overcoming structural issues.
Recent progress in the field of artificial intelligence, machine learning and also in computer industry resulted in the ongoing boom of using these techniques as applied to solving complex tasks in both science and industry. Same is, of course, true for the financial industry and mathematical finance. In this paper we …
This paper evaluates test selection methods for deep neural networks, revealing their limitations.
Deep learning applied to SAR data is explored in this paper.
This paper highlights the overlooked role of preprocessing hyperparameters in machine learning model performance.
Cross-validation pitfalls in change-point regression are addressed with new approaches.
New findings show a balance between data fit and complexity in kernel hyperparameters.
LLA shows strong performance in Bayesian optimization but has unbounded search space issues.
Success in the quest for artificial intelligence has the potential to bring unprecedented benefits to humanity, and it is therefore worthwhile to investigate how to maximize these benefits while avoiding potential pitfalls. This article gives numerous examples (which should by no means be construed as an exhaustive lis…
New datasets reveal neural networks can rely on simple features, leading to poor generalization.
Machine learning methods have gained a great deal of popularity in recent years among public administration scholars and practitioners. These techniques open the door to the analysis of text, image and other types of data that allow us to test foundational theories of public administration and to develop new theories. …
Proposes a new undersampling method for imbalanced data classification.
Statistical analysis reveals pitfalls in climate network construction.
This document provides a tutorial description of the use of the MDL principle in complex graph analysis. We give a brief summary of the preliminary subjects, and describe the basic principle, using the example of analysing the size of the largest clique in a graph. We also provide a discussion of how to interpret the r…
TBAL reduces manual annotation but requires validated data.
Statistical tests for fairness in admissions data reveal hidden patterns.
Paper proposes an ensemble of attacks to evaluate adversarial robustness more reliably.
This document is an invited chapter covering the specificities of ABC model choice, intended for the incoming Handbook of ABC by Sisson, Fan, and Beaumont (2017). Beyond exposing the potential pitfalls of ABC based posterior probabilities, the review emphasizes mostly the solution proposed by Pudlo et al. (2016) on the…
This letter uses the Block Maxima Extreme Value approach to quantify catastrophic risk in international equity markets. Risk measures are generated from a set threshold of the distribution of returns that avoids the pitfall of using absolute returns for markets exhibiting diverging levels of risk. From an application t…
Framework adds human knowledge to AI decisions to improve outcomes.
We construct Zero-Coupon Bond markets driven by a cylindrical Brownian motion in which the notion of generalized portfolio has important flaws: There exist bounded smooth random variables with generalized hedging portfolios for which the price of their risky part is at each time. For these generalized portfol…
Abstract MDPs enable strategic exploration and fast reward transfer in complex environments.
The recent progress on capsule networks by Hinton et al. has generated considerable excitement in the machine learning community. The idea behind a capsule is inspired by a cortical minicolumn in the brain, whereby a vertically organised group of around 100 neurons receive common inputs, have common outputs, are interc…
In this paper we sketch some reflections on the pitfalls and inconsistencies of the research program - currently dominant among the profession - aimed at providing microfoundations to macroeconomics along a Walrasian perspective. We argue that such a methodological approach constitutes an unsatisfactory answer to a wel…
We highlight a pitfall when applying stochastic variational inference to general Bayesian networks. For global random variables approximated by an exponential family distribution, natural gradient steps, commonly starting from a unit length step size, are averaged to convergence. This useful insight into the scaling of…
When faced with high frequency streams of data, clustering raises theoretical and algorithmic pitfalls. We introduce a new and adaptive online clustering algorithm relying on a quasi-Bayesian approach, with a dynamic (i.e., time-dependent) estimation of the (unknown and changing) number of clusters. We prove that our a…
The purpose of these notes is to provide a systematic quantitative framework - in what is intended to be a "pedagogical" fashion - for discussing mean-reversion and optimization. We start with pair trading and add complexity by following the sequence "mean-reversion via demeaning -> regression -> weighted regression ->…
The success of modern Artificial Intelligence (AI) technologies depends critically on the ability to learn non-linear functional dependencies from large, high dimensional data sets. Despite recent high-profile successes, empirical evidence indicates that the high predictive performance is often paired with low robustne…
Machine learning aids excited-state molecular dynamics studies.
When confronted with massive data streams, summarizing data with dimension reduction methods such as PCA raises theoretical and algorithmic pitfalls. Principal curves act as a nonlinear generalization of PCA and the present paper proposes a novel algorithm to automatically and sequentially learn principal curves from d…
Correctly evaluating defenses against adversarial examples has proven to be extremely difficult. Despite the significant amount of recent work attempting to design defenses that withstand adaptive attacks, few have succeeded; most papers that propose defenses are quickly shown to be incorrect. We believe a large contri…
There is an increasing interest in estimating expectations outside of the classical inference framework, such as for models expressed as probabilistic programs. Many of these contexts call for some form of nested inference to be applied. In this paper, we analyse the behaviour of nested Monte Carlo (NMC) schemes, for w…
This survey aims to provide a guide to the literature on topological 4-manifolds. Foundational theorems on 4-manifolds are stated, especially in the topological category. Precise references are given, with indications of the strategies employed in the proofs. Where appropriate we give statements for manifolds of all di…
Paper explores limitations of generative models in finance, proposing a new method for portfolio generation.
A grand challenge in reinforcement learning is intelligent exploration, especially when rewards are sparse or deceptive. Two Atari games serve as benchmarks for such hard-exploration domains: Montezuma's Revenge and Pitfall. On both games, current RL algorithms perform poorly, even those with intrinsic motivation, whic…
GrowNet uses shallow neural networks for gradient boosting, outperforming existing methods.
In this paper, we briefly review the basic scheme of the pseudoinverse learning (PIL) algorithm and present some discussions on the PIL, as well as its variants. The PIL algorithm, first presented in 1995, is a non-gradient descent and non-iterative learning algorithm for multi-layer neural networks and has several adv…
In this paper we propose a novel approach for learning from data using rule based fuzzy inference systems where the model parameters are estimated using Bayesian inference and Markov Chain Monte Carlo (MCMC) techniques. We show the applicability of the method for regression and classification tasks using synthetic data…