Bayesian method detects outliers and uncertain points in data.
problem Detecting outliers and uncertain points in data using Bayesian methods.
method Generative model of data curation for aleatoric uncertainty, combining with epistemic uncertainty and outlier exposure.
result Principled Bayesian approach outperforms methods using aleatoric or epistemic uncertainty alone.
It is important to detect anomalous inputs when deploying machine learning systems. The use of larger and more complex inputs in deep learning magnifies the difficulty of distinguishing between anomalous and in-distribution examples. At the same time, diverse image and text data are available in enormous quantities. We…
Few random images can improve anomaly detection performance.
problem Improving anomaly detection performance with limited labeled data.
method Utilizing large collections of random images to represent anomalousness.
result Standard classifiers and semi-supervised one-class methods can achieve strong performance with just a small collection of outlier exposure data.
Deep neural networks have achieved great success in classification tasks during the last years. However, one major problem to the path towards artificial intelligence is the inability of neural networks to accurately detect samples from novel class distributions and therefore, most of the existent classification algori…
DCASE 2021 ASD task tackles domain-shifted anomalous sound detection.
problem Detecting unknown anomalous sounds under domain-shifted conditions.
method Ensemble of outlier exposure and inlier modeling detectors, feature learning from machine identification.
result Two types of remarkable approaches were adopted by top teams.
This paper combines existing OOD detection methods to improve overall performance.
problem Improving robustness of neural networks in safety-critical applications.
method Integrates four strategies for combining multiple OOD detection scores.
result Enhanced OOD detection through multi-dimensional evaluation metrics.
We introduce negative binomial matrix factorization (NBMF), a matrix factorization technique specially designed for analyzing over-dispersed count data. It can be viewed as an extension of Poisson matrix factorization (PF) perturbed by a multiplicative term which models exposure. This term brings a degree of freedom fo…
Green stocks show less factor exposure heterogeneity compared to brown stocks.
problem Exploring differences in factor exposure between green and brown stocks.
method Examined S&P 500 firms grouped by greenhouse gas emissions, analyzing factor exposure over 2014-2020.
result Green stocks have less factor exposure heterogeneity than brown stocks, except for the value factor.
Deep learning approximates Bermudan option exposures and future values.
problem Computing accurate expected and future exposures for high-dimensional Bermudan options.
method Neural network-based approach combining Deep Optimal Stopping and regression.
result Neural network approximations of pathwise option values are more accurate.
Study shows short exposure and systematic risk exposure affect disposition effect asymmetries.
problem Understanding disposition effect in short vs long exposure positions and systematic risk.
method Generalized Odean measures, introduced Value metric, implemented dispositionEffect R package.
result Short positions exhibit weaker disposition effect than long positions under narrow framing, reversing in integrated framing.
Randomized neural networks improve exposure and CVA estimation for American options.
problem Estimation of exposure and CVA for American options
method Randomized neural networks
result Improves convergence and efficiency in high-dimensional problems
Valuation of Credit Valuation Adjustment (CVA) has become an important field as its calculation is required in Basel III, issued in 2010, in the wake of the credit crisis. Exposure, which is defined as the potential future loss of a default event without any recovery, is one of the key elementsfor pricing CVA. This pap…
Paper optimizes neural networks for Bermudan option pricing with faster convergence and risk management tools.
problem Efficiently pricing Bermudan options with static hedging and risk management.
method Monte-Carlo-based artificial neural network framework with novel optimisation algorithm.
result The proposed neural network accelerates convergence and provides improved risk management tools.
Exposure bias has been regarded as a central problem for auto-regressive language models (LM). It claims that teacher forcing would cause the test-time generation to be incrementally distorted due to the training-generation discrepancy. Although a lot of algorithms have been proposed to avoid teacher forcing and theref…
In epidemiology, identifying the effect of exposure variables in relation to a time-to-event outcome is a classical research area of practical importance. Incorporating propensity score in the Cox regression model, as a measure to control for confounding, has certain advantages when outcome is rare. However, in situati…
We study the impact of central clearing of over-the-counter (OTC) transactions on counterparty exposures in a market with OTC transactions across several asset classes with heterogeneous characteristics. The impact of introducing a central counterparty (CCP) on expected interdealer exposure is determined by the tradeof…
The paper proposes using function approximations to reduce the computational burden in measuring counterparty credit exposure.
problem The need for regular exposure calculations in finance, balancing between computational cost and risk simplification.
method Replacing derivative pricers with function approximations, proving error bounds, and using Chebyshev interpolation for convergence.
result Derives probabilistic and finite sample error bounds, showing significant run-time reductions and asymptotic efficiency gains.
Study prenatal PM2.5 exposure and 4th grade reading scores, identifying critical windows of susceptibility.
problem Understanding the impact of prenatal PM2.5 exposure on educational outcomes.
method Developed a locally adaptive Bayesian regression model with B-spline basis expansion and dynamic shrinkage priors.
result Prenatal PM2.5 exposure during early and late pregnancy is most adverse for 4th grade reading scores.
Mack's estimator improves chain ladder prediction for large exposure insurance models.
problem Uncertainty quantification in compound Poisson loss models.
method Large exposure asymptotics applied to Mack's estimator.
result Chain ladder prediction uncertainty can be quantified without model assumptions.
Study estimates personalized effects of maternal PM2.5 exposure on birth weight.
problem Identify critical windows and heterogeneity in maternal PM2.5 exposure effects on birth weight.
method Heterogeneous Distributed Lag Models and Bayesian Additive Regression Trees.
result Evidence of heterogeneity in PM2.5-birth weight relationship, with some dyads showing 3x larger decrease.
We introduce a new method to calculate the credit exposure of Bermudan, discretely monitored barrier and European options. Core of the approach is the application of the dynamic Chebyshev method of Glau et al. (2019). The dynamic Chebyshev method delivers a closed form approximation of the option prices along the paths…
New method debiases selection bias in PU classification with exposure data.
problem Binary classification from positive and unlabeled data with selection bias.
method Automatic Debiased PUE (ADPUE) learning method.
result ADPUE outperforms traditional PU learning methods on various datasets.
Collaborative filtering analyzes user preferences for items (e.g., books, movies, restaurants, academic papers) by exploiting the similarity patterns across users. In implicit feedback settings, all the items, including the ones that a user did not consume, are taken into consideration. But this assumption does not acc…
Modeling incentives for content creators on algorithm-curated platforms.
problem Maximizing exposure for content creators on algorithmic platforms.
method Formalized exposure game model, proving effects of algorithmic choices on equilibria, proposing tools for finding equilibria.
result Algorithmic choices significantly affect content exposure and creator behavior.
BN^2MF identifies unknown exposure patterns in environmental mixtures.
problem Identifying unknown exposure patterns in environmental mixtures.
method Bayesian non-parametric non-negative matrix factorization (BN^2MF) with non-negative continuous priors and a non-parametric sparse prior.
result Estimates patterns of chemical exposures without specifying the number of patterns.
We propose an inlier-based outlier detection method capable of both identifying the outliers and explaining why they are outliers, by identifying the outlier-specific features. Specifically, we employ an inlier-based outlier detection criterion, which uses the ratio of inlier and test probability densities as a measure…
A novel approach ODAR detects outliers for clustering.
problem Outliers interfere with clustering algorithms, leading to unreliable results.
method Feature transformation to separate outliers and normal objects into distinct clusters.
result ODAR improves clustering accuracy on 7 out of 10 datasets.
Clustering, or unsupervised classification, is a task often plagued by outliers. Yet there is a paucity of work on handling outliers in clustering. Outlier identification algorithms tend to fall into three broad categories: outlier inclusion, outlier trimming, and post hoc outlier identification methods, with the forme…
We introduce a new method to calculate the credit exposure of European and path-dependent options. The proposed method is able to calculate accurate expected exposure and potential future exposure profiles under the risk-neutral and the real-world measure. Key advantage of is that it delivers an accuracy comparable to …
Researchers develop a new method to assess variable importance in spatial machine learning models for air pollution exposure prediction.
problem Understanding the mechanism captured by machine learning models in air pollution studies, especially with spatial correlation.
method Leave-one-out approach for variable importance measure applicable to models with separable mean and covariance components.
result The new method highlights differences in model mechanisms even for similar prediction accuracies.
New framework detects outliers in non-IID categorical data.
problem Existing outlier detection methods fail in non-IID data.
method Value-value graph-based representation and outlierness propagation.
result Significant improvement in AUC on complex data sets.
Outlier detection aims to identify unusual data instances that deviate from expected patterns. The outlier detection is particularly challenging when outliers are context dependent and when they are defined by unusual combinations of multiple outcome variable values. In this paper, we develop and study a new conditiona…
CENNSurv models cumulative effects of time-dependent exposures on survival outcomes.
problem Challenges in modeling cumulative effects of time-dependent exposures on survival outcomes.
method CENNSurv, a novel deep learning approach that captures dynamic risk relationships from time-dependent data.
result CENNSurv reveals multi-year lagged and short-term behavioral shifts in survival outcomes.
Bayesian approach models nonignorable missing data using copulas and marginal quantiles.
problem Nonignorable missing data in lead exposure and test score analysis.
method Gaussian copula model with auxiliary marginal quantiles for missingness indicators and study variables.
result Efficient MCMC algorithm estimates copula correlation and marginal distributions consistently.
The study tests a functional-form restriction on risk exposure dynamics using margin debt data.
problem Understanding risk exposure dynamics under capital constraints and slack.
method Testing a regime-conditional functional-form restriction on aggregate risk-exposure dynamics implied by VaR-constrained intermediary models.
result The contraction and growth of exposures under capital constraints and slack are observed and tested.
A new method uses counterfactual learning to improve recommendation system evaluation.
problem Inconsistent results in recommender systems due to exposure mechanisms.
method Proposes a minimax empirical risk formulation with an adversarial game to account for exposure.
result Shows improved learning bounds and effectiveness over various recommendation settings.
Paper proposes methods to make OT robust to outliers.
problem Optimal transport is sensitive to outliers.
method Detect outliers using adversarial training, adjust transport cost based on classifier predictions.
result Outliers are detected and do not affect transport in experiments.
Outlier detection plays an essential role in many data-driven applications to identify isolated instances that are different from the majority. While many statistical learning and data mining techniques have been used for developing more effective outlier detection algorithms, the interpretation of detected outliers do…
Speeds up complex portfolio exposure calculations.
problem Calculating exposure of portfolios with exotic derivatives.
method Least Squares Monte Carlo (LSMC) technique.
result Significantly reduces computation time for nested Monte Carlo.
Advances in sensor technology have enabled the collection of large-scale datasets. Such datasets can be extremely noisy and often contain a significant amount of outliers that result from sensor malfunction or human operation faults. In order to utilize such data for real-world applications, it is critical to detect ou…
The paper clarifies long-horizon investment and DCA, showing no risk reduction but different exposure profiles.
problem Misleading claims about reducing risk with longer investment horizons and DCA.
method Unified probabilistic framework, defining risk and uncertainty, and introducing effective investment exposure.
result Different investment timing strategies can lead to distinct exposure profiles over time, affecting risk and uncertainty.
Survey compares methods for generating artificial outliers.
problem Difficulty in detecting genuine outliers.
method Generates artificial outliers to approximate genuine ones.
result Variability in quality of generation approaches.
Paper tackles outlier detection in signals modeled by generative models with theoretical guarantees.
problem Recovering signals from linear measurements with sparse outliers.
method Proposes an iterative ADMM algorithm and gradient descent algorithm for outlier detection using ℓ1 and squared ℓ1 norm minimization. result Establishes theoretical recovery guarantees for signal reconstruction under sparse outliers.
Transforms distance-based outlier scores into interpretable probabilistic estimates.
problem Difficult interpretation of distance-based outlier scores.
method Generic transformation of scores into probabilistic estimates using distance probability distributions.
result Probabilistic transformation improves interpretability without impacting detection performance.
TradeMech nets trades without changing counterparty relationships.
problem Netting trades without altering counterparty exposure in complex financial networks.
method Transforms contracts into chains and cycles, nets designated object multilaterally, and replaces contracts with new multiparty agreements.
result Maximal multilateral netting of a designated object while preserving each agent's profit and counterparty risk.
Generates synthetic data for benchmarking unsupervised outlier detection.
problem Difficulty in benchmarking unsupervised outlier detection due to rare and varied outliers in real data.
method Proposes a generic process to generate synthetic data with insightful characteristics.
result Demonstrates practicality of the generic process through a benchmark with state-of-the-art detection methods.
A novel unsupervised outlier detection method using Randomized PCA Forest.
problem Unsupervised outlier detection in datasets.
method Randomized Principal Component Analysis (RPCA) Forest for deriving an outlier score.
result Superior performance compared to classical and state-of-the-art methods.
OpFlow predicts robust OD flows by learning choice potentials conditioned on spatial exposures.
problem Deep models trained on raw counts are vulnerable to distribution shift.
method OpFlow learns row-centered choice potentials and reconstructs flows by combining them with a calibrated origin scale.
result OpFlow improves robustness under environment shifts, as shown by controlled synthetic shifts and a real-world experiment.