New local MDI variable importances derived from global scores match Shapley values.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Methodology to measure lag relevance in time series models.
Study examines value relevance of oil and gas reserve disclosures in London Stock Exchange.
NP-PROV separates mean and variance spaces to improve function uncertainty.
A new method for efficiently estimating Shapley values in dataset valuation.
The Nonlinear autoregressive exogenous (NARX) model, which predicts the current value of a time series based upon its previous values as well as the current and past values of multiple driving (exogenous) series, has been studied for decades. Despite the fact that various NARX models have been developed, few of them ca…
We study the average shape of a fluctuation of a time series x(t), that is the average value <x(t)-x(0)>_T before x(t) first returns, at time T, to its initial value x(0). For large classes of stochastic processes we find that a scaling law of the form <x(t) - x(0)>_T = T^αf(t/T) is obeyed. The scaling function f(s) is…
System detects relevant financial news and predictions from unstructured text.
We analyze a notion of multiple valued sections of a vector bundle over an abstract smooth Riemannian manifold, which was suggested by W. Allard in the unpublished note "Some useful techniques for dealing with multiple valued functions" and generalizes Almgren's -valued functions. We study some relevant properties o…
PCA simplifies multivariate extreme data analysis.
In this note, we comment on the relevance of elicitability for backtesting risk measure estimates. In particular, we propose the use of Diebold-Mariano tests, and show how they can be implemented for Expected Shortfall (ES), based on the recent result of Fissler and Ziegel (2015) that ES is jointly elicitable with Valu…
In this work, we provide a framework linking microstructural properties of an asset to the tick value of the exchange. In particular, we bring to light a quantity, referred to as implicit spread, playing the role of spread for large tick assets, for which the effective spread is almost always equal to one tick. The rel…
Adequate evaluation of an information retrieval system to estimate future performance is a crucial task. Area under the ROC curve (AUC) is widely used to evaluate the generalization of a retrieval system. However, the objective function optimized in many retrieval systems is the error rate and not the AUC value. This p…
A new jump diffusion regime-switching model is introduced, which allows for linking jumps in asset prices with regime changes. We prove the existence and uniqueness of the solution to the risk-sensitive asset management criterion maximisation problem in this setting. We provide an ODE for the optimal value function, wh…
Recommender systems, medical diagnosis, network security, etc., require on-going learning and decision-making in real time. These -- and many others -- represent perfect examples of the opportunities and difficulties presented by Big Data: the available information often arrives from a variety of sources and has divers…
We describe two topologies on the space of unbounded Fredholm operators and we explain their K-theoretic relevance. In the process we also prove a very general result concerning the continuity of families of first order, elliptic boundary value problems.
For purposes of Value-at-Risk estimation, we consider several multivariate families of heavy-tailed distributions, which can be seen as multidimensional versions of Paretian stable and Student's t distributions allowing different marginals to have different tail thickness. After a discussion of relevant estimation and …
The paper calculates the value of information in high-dimensional decision making.
Measurements made by satellite remote sensing, Moderate Resolution Imaging Spectroradiometer (MODIS), and globally distributed Aerosol Robotic Network (AERONET) are compared. Comparison of the two datasets measurements for aerosol optical depth values show that there are biases between the two data products. In this pa…
Defines cost of MEV and shows its relevance in various settings.
Valid p-value for bounded random variables without distributional assumptions.
Blood lactate concentration is a strong indicator of mortality risk in critically ill patients. While frequent lactate measurements are necessary to assess patient's health state, the measurement is an invasive procedure that can increase risk of hospital-acquired infections. For this reason we formally define the prob…
Feature selection with high-dimensional data and a very small proportion of relevant features poses a severe challenge to standard statistical methods. We have developed a new approach (HARVEST) that is straightforward to apply, albeit somewhat computer-intensive. This algorithm can be used to pre-screen a large number…
Study improves self-normalized bounds for vector-valued processes beyond sub-Gaussianity.
New method uses extreme value theory to estimate neural network errors.
In many real-world machine learning problems, feature values are not readily available. To make predictions, some of the missing features have to be acquired, which can incur a cost in money, computational time, or human time, depending on the problem domain. This leads us to the problem of choosing which features to u…
OSIRIS reduces variance in off-policy evaluation by omitting irrelevant states.
The goal of feature selection is to identify important features that are relevant to explain an outcome variable. Most of the work in this domain has focused on identifying globally relevant features, which are features that are related to the outcome using evidence across the entire dataset. We study a more fine-grain…
The problem of explaining the behavior of deep neural networks has recently gained a lot of attention. While several attribution methods have been proposed, most come without strong theoretical foundations, which raises questions about their reliability. On the other hand, the literature on cooperative game theory sugg…
This work surveys algorithmic recourse, aiming to clarify definitions and solutions.
Factor analysis has proven to be a relevant tool for extracting tissue time-activity curves (TACs) in dynamic PET images, since it allows for an unsupervised analysis of the data. Reliable and interpretable results are possible only if considered with respect to suitable noise statistics. However, the noise in reconstr…
The paper tackles reward-relevance in offline RL with sparse decision dynamics.
Study proposes explainable analytics for manufacturing process planning.
The tick value is a crucial component of market design and is often considered the most suitable tool to mitigate the effects of high frequency trading. The goal of this paper is to demonstrate that the approach introduced in Dayri and Rosenbaum (2015) allows for an ex ante assessment of the consequences of a tick valu…
Providing accurate predictions is challenging for machine learning algorithms when the number of features is larger than the number of samples in the data. Prior knowledge can improve machine learning models by indicating relevant variables and parameter values. Yet, this prior knowledge is often tacit and only availab…
An unsupervised anomaly detection method for irregularly sampled time-series data.
Many real-life decision-making situations allow further relevant information to be acquired at a specific cost, for example, in assessing the health status of a patient we may decide to take additional measurements such as diagnostic tests or imaging scans before making a final assessment. Acquiring more relevant infor…
We discuss promising recent contributions on quantifying feature relevance using Shapley values, where we observed some confusion on which probability distribution is the right one for dropped features. We argue that the confusion is based on not carefully distinguishing between observational and interventional conditi…
We introduce a -valued cross ratio on Roller boundaries of cube complexes. We motivate its relevance by showing that every cross-ratio preserving bijection of Roller boundaries uniquely extends to a cubical isomorphism. Our results are strikingly general and even apply to infinite dimensional…
We define the Ricci curvature, as a measure, for certain singular torsion-free connections on the tangent bundle of a manifold. The definition uses an integral formula and vector-valued half-densities. We give relevant examples in which the Ricci measure can be computed. In the time dependent setting, we give a weak no…
Improved local feature attributions using neighbourhood reference distributions.
MuZero visualizes its internal representations to stabilize planning.
This paper considers a utility maximization and optimal asset allocation problem in the presence of a stochastic endowment that cannot be fully hedged through trading in the financial market. After studying continuity properties of the value function for general utility functions, we rely on the dynamic programming app…
Geometric structures on -manifolds, i.e.~non-negatively graded manifolds with an homological vector field, encode non-graded geometric data on Lie algebroids and their higher analogues. A particularly relevant class of structures consists of vector bundle valued differential forms. Symplectic forms, contac…
Correctly pricing products or services in an online marketplace presents a challenging problem and one of the critical factors for the success of the business. When users are looking to buy an item they typically search for it. Query relevance models are used at this stage to retrieve and rank the items on the search p…
Method selects features robust to concept shift using Shapley values.
Distributional (or distribution-valued) data are a new type of data arising from several sources and are considered as realizations of distributional variables. A new set of fuzzy c-means algorithms for data described by distributional variables is proposed. The algorithms use the Wasserstein distance between dist…
Based on forward curves modelled as Hilbert-space valued processes, we analyse the pricing of various options relevant in energy markets. In particular, we connect empirical evidence about energy forward prices known from the literature to propose stochastic models. Forward prices can be represented as linear functions…