Develops method to assess feature importance in black-box models for unconditional distribution.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
In this paper we study the setting where features are added or change interpretation over time, which has applications in multiple domains such as retail, manufacturing, finance. In particular, we propose an approach to provably determine the time instant from which the new/changed features start becoming relevant with…
Paper proposes a feature-wise change detection method for improving indoor positioning accuracy.
The paper uses AI to analyze ECG data, revealing age-related changes and identifying key features.
New methods interpret clustering outcomes without altering data structure.
Online change detection algorithm using random Fourier features.
Optimizes feature shifts for tree ensemble reclassification.
CDLEEDS detects local changes in evolving data streams for accurate feature attributions.
Paper introduces a method to explain concept drift using counterfactual explanations.
Proposes a method to update model weights by dynamically changing features during training.
When training a deep neural network for image classification, one can broadly distinguish between two types of latent features of images that will drive the classification. We can divide latent features into (i) "core" or "conditionally invariant" features whose distribution , cond…
This paper shows feature importance remains valid even in low-performing models.
Modular pipeline improves stock portfolio prediction robustness under regime changes.
This paper uses PCA and FA for feature selection in credit rating.
In Random Forests, proximity distances are a metric representation of data into decision space. By observing how changes in input map to the movement of instances in this space we are able to determine the independent contribution of each feature to the decision-making process. For binary feature vectors, this process …
Defense against adversarial attacks by manipulating feature thickness.
A new method detects change points in time series with conceptors.
Detects data drift and outliers affecting ML model performance over time.
VFDS selects dynamic features for efficient HAR tasks, optimizing performance-cost trade-offs.
Inverse classification is the process of perturbing an instance in a meaningful way such that it is more likely to conform to a specific class. Historical methods that address such a problem are often framed to leverage only a single classifier, or specific set of classifiers. These works are often accompanied by naive…
The performance of machine learning model can be further improved if contextual cues are provided as input along with base features that are directly related to an inference task. In offline learning, one can inspect historical training data to identify contextual clusters either through feature clustering, or hand-cra…
Inverse classification is the process of manipulating an instance such that it is more likely to conform to a specific class. Past methods that address such a problem have shortcomings. Greedy methods make changes that are overly radical, often relying on data that is strictly discrete. Other methods rely on certain da…
A crucial challenge in image-based modeling of biomedical data is to identify trends and features that separate normality and pathology. In many cases, the morphology of the imaged object exhibits continuous change as it deviates from normality, and thus a generative model can be trained to model this morphological con…
Scaling feature values is an important step in numerous machine learning tasks. Different features can have different value ranges and some form of a feature scaling is often required in order to learn an accurate classifier. However, feature scaling is conducted as a preprocessing task prior to learning. This is probl…
Regulatory compliance is an organization's adherence to laws, regulations, guidelines and specifications relevant to its business. Compliance officers responsible for maintaining adherence constantly struggle to keep up with the large amount of changes in regulatory requirements. Keeping up with the changes entail two …
Perturbation-based explanation methods often measure the contribution of an input feature to an image classifier's outputs by heuristically removing it via e.g. blurring, adding noise, or graying out, which often produce unrealistic, out-of-samples. Instead, we propose to integrate a generative inpainter into three rep…
New method helps interpret complex models by visualizing feature shifts.
Counterfactual approach explains AI decisions using causal data inputs.
An essential problem in domain adaptation is to understand and make use of distribution changes across domains. For this purpose, we first propose a flexible Generative Domain Adaptation Network (G-DAN) with specific latent variables to capture changes in the generating process of features across domains. By explicitly…
We present a scalable Gaussian process model for identifying and characterizing smooth multidimensional changepoints, and automatically learning changes in expressive covariance structure. We use Random Kitchen Sink features to flexibly define a change surface in combination with expressive spectral mixture kernels to …
Machine learning for healthcare often trains models on de-identified datasets with randomly-shifted calendar dates, ignoring the fact that data were generated under hospital operation practices that change over time. These changing practices induce definitive changes in observed data which confound evaluations which do…
Concept drift in learning and classification occurs when the statistical properties of either the data features or target change over time; evidence of drift has appeared in search data, medical research, malware, web data, and video. Drift adaptation has not yet been addressed in high dimensional, noisy, low-context d…
In this study, we attempted to determine how eigenvalues change, according to random matrix theory (RMT), in stock market data as the number of stocks comprising the correlation matrix changes. Specifically, we tested for changes in the eigenvalue properties as a function of the number and type of stocks in the correla…
Boosts change-point detection power with optimal sub-sampling.
We introduce a class of randomly time-changed fast mean-reverting stochastic volatility models and, using spectral theory and singular perturbation techniques, we derive an approximation for the prices of European options in this setting. Three examples of random time-changes are provided and the implied volatility sur…
The paper analyzes MAML's representation using RSA, revealing that feature reuse is not the primary reason for its success.
Distributed sensors compress and send features to a fusion center for linear regression.
While several feature scoring methods are proposed to explain the output of complex machine learning models, most of them lack formal mathematical definitions. In this study, we propose a novel definition of the feature score using the maximally invariant data perturbation, which is inspired from the idea of adversaria…
Study reveals how features influence deep network decision boundaries.
New method for interpreting non-linear models using forward marginal effects.
Method predicts disease outbreaks using search logs, overcoming instability.
Researchers prove a special case of Borde-Sorkin's conjecture about topology change in spacetimes.
PredDiff measures prediction changes while marginalizing features, offering new insights into interaction effects.
Robust ASR model removes fast-changing features to resist attacks.
Intrusion detection systems (IDSs) generate valuable knowledge about network security, but an abundance of false alarms and a lack of methods to capture the interdependence among alerts hampers their utility for network defense. Here, we explore a graph-based approach for fusing alerts generated by multiple IDSs (e.g.,…
Identifies features most relevant to concept drift in data.
DVE models dynamic changes in feature embeddings for better sequence-aware applications.
Although deep learning has shown great success in recent years, researchers have discovered a critical flaw where small, imperceptible changes in the input to the system can drastically change the output classification. These attacks are exploitable in nearly all of the existing deep learning classification frameworks.…