In recent years, a large amount of model-agnostic methods to improve the transparency, trustability and interpretability of machine learning models have been developed. We introduce local feature importance as a local version of a recent model-agnostic global feature importance method. Based on local feature importance…
Paper tackles estimating individual treatment effects from observational data.
problem Estimating the difference between outcomes with and without treatment from single observation.
method Formulated as inference from hidden variables, uses a model of four causal populations, proposes ECM algorithm.
result ECM algorithm provides better performance compared to baseline methods on synthetic and real-world data.
This paper discusses an alternative explanation for the empirical findings contradicting the positive relationship between risk (variance) and reward (expected return). We show that these contradicting results might be due to the false definition of risk-perception, which we correct by introducing Expected Downside Ris…
Proposes ICE-based metric for better understanding interactions in black-box models.
problem Misleading global sensitivity metrics in black-box models due to interaction effects.
method Individual Conditional Expectation (ICE) curves to compute feature importance and interactions.
result ICE-based metric provides richer insights into feature importance and interactions.
Recent exploration of optimal individualized decision rules (IDRs) for patients in precision medicine has attracted a lot of attention due to the heterogeneous responses of patients to different treatments. In the existing literature of precision medicine, an optimal IDR is defined as a decision function mapping from t…
Generative AI predicts economic activity from corporate transcripts.
problem Predicting economic activity using existing measures like surveys.
method Extracted managerial expectations from transcripts using generative AI.
result AI Economy Score predicts economic activity up to 10 quarters ahead.
Optimal insurance strategy for maximizing RDEU under various premium principles.
problem Maximizing a risk-averse individual's RDEU with insurance priced by a distortion-deviation principle.
method Proved necessary and sufficient conditions for the optimal solution, considered ambiguity orders, and analyzed specific examples.
result Conditions for no insurance or deductible insurance to be optimal.
This research simplifies computation of feature attribution methods under certain conditions.
problem Computational complexity of feature attribution methods, especially power indices.
method Identifying conditions for polynomial computation and introducing new indices.
result Conditions for efficient computation of feature attribution methods are identified.
Mean-field approximations simplify insurance liability calculations.
problem High-dimensional system of equations makes insurance liability calculation infeasible.
method Use mean-field model to replace high-dimensional system with a low-dimensional non-linear system.
result Insurance liability converges to mean-field approximation as cohort size increases.
New method reveals true causal functions in nonlinear time series, not just scores.
problem Causal discovery in nonlinear time series often uses scalar edge scores, which hide true function-valued causal influence.
method Formalized function-valued causal influence for additive, contribution-decomposable architectures. Introduced a practical framework based on ICE for estimating causal response functions directly from trained models.
result Edges with indistinguishable scalar scores can exhibit qualitatively different functional behaviors.
Recent years have witnessed an increased focus on interpretability and the use of machine learning to inform policy analysis and decision making. This paper applies machine learning to examine travel behavior and, in particular, on modeling changes in travel modes when individuals are presented with a novel (on-demand)…
We determine the optimal amount to invest in a Black-Scholes financial market for an individual who consumes at a rate equal to a constant proportion of her wealth and who wishes to minimize the expected time that her wealth spends in drawdown during her lifetime. Drawdown occurs when wealth is less than some fixed pro…
Unified asymptotic treatment for VaR- and expectile-based systemic risk measures.
problem Analyzing systemic risk measures under extreme system-wide disasters.
method Classified systemic risk measures into VaR- and expectile-based families, introduced new ICE and SICE measures, and provided second-order asymptotic results.
result Second-order asymptotics provide more accurate tail approximations for systemic risk measures.
Aims to create safe reinforcement learning policies by considering individual harm.
problem Optimal policies for a population may harm certain individuals.
method Formalizes individual harm, proposes a two-stage procedure, and establishes finite-sample properties.
result Learned policies maximize expected return while minimizing harm.
Proposes a new risk model using stable laws to manage company-wide losses.
problem Managing aggregate risks and pricing policies in the presence of systematic risk.
method Develops a modified risk model using multivariate stable distributions to account for various risk phenomena.
result Computes the Tail Conditional Expectation of aggregate risks and corresponding allocations.
This paper introduces individual fairness in clustering using f-divergence.
problem Ensuring fair clustering by treating similar individuals similarly.
method Uses f-divergence to measure statistical similarity and assigns individuals to probability distributions over cluster centers. result Provides an algorithm with provable approximation guarantee for clustering with individual fairness constraints.
Extends expected value framework for cost-sensitive causal decision-making.
problem Optimizing operational decision-making with cost-sensitive causal classification.
method Introduces a cost-sensitive decision boundary based on estimated individual treatment effects, positive outcome probability, and cost parameters.
result Effective in maximizing expected causal profit, outperforming cost-insensitive ranking approach.
Algorithm provides fair clustering guarantees for k-means and k-median.
problem Ensuring fair clustering guarantees for every point in a dataset.
method Local search algorithm for k-median and k-means, with individual fairness as a key metric. result Achieves fair clustering with a constant factor approximation of the optimal solution.
Kernel method optimizes personalized dose rules for patients.
problem Finding optimal individualized dose rules for patients.
method Kernel assisted learning method for estimating optimal dose rules.
result The method identifies the optimal individualized dose rule and produces favorable outcomes.
Study on estimating conditional risk in machine learning.
problem Estimating expected loss of prediction models given input features.
method Analyzed in classification and regression settings, showing equivalence to standard regression. Developed theoretical insights and empirical validation.
result Conditional risk calibration is distinct from existing uncertainty quantification problems.
In this paper we study the effect of network structure between agents and objects on measures for systemic risk. We model the influence of sharing large exogeneous losses to the financial or (re)insuance market by a bipartite graph. Using Pareto-tailed losses and multivariate regular variation we obtain asymptotic resu…
Network-based models predict user preferences for items like movies and research articles.
problem Filtering and delivering personalized advice for users with many available products.
method Network models based on group memberships, using Monte Carlo sampling and Expectation-Maximization methods.
result Network models outperform leading approaches for recommendation.
A new realized conditional autoregressive Value-at-Risk (VaR) framework is proposed, through incorporating a measurement equation into the original quantile regression model. The framework is further extended by employing various Expected Shortfall (ES) components, to jointly estimate and forecast VaR and ES. The measu…
Who {\em values} life annuities more? Is it the healthy retiree who expects to live long and might become a centenarian, or is the unhealthy retiree with a short life expectancy more likely to appreciate the pooling of longevity risk? What if the unhealthy retiree is pooled with someone who is much healthier and thus f…
The random cluster model is used to define an upper bound on a distance measure as a function of the number of data points to be classified and the expected value of the number of classes to form in a hybrid K-means and regression classification methodology, with the intent of detecting anomalies. Conditions are given …
The paper introduces a machine learning method to forecast market direction using efficient frontier coefficients.
problem Improving asset return estimation for portfolio optimization.
method Monthly directional market forecast using an online decision tree trained on efficient frontier coefficients.
result The method outperforms baseline portfolios and other feature sets.
The paper optimizes investment strategies with constraints for life-cycle models.
problem Maximizing consumption, death benefit, and wealth under trading constraints.
method Deep pricing kernel approach to solve constrained portfolio optimization.
result Individuals reduce consumption, insurance demand, and wealth due to constraints.
Proposes a new decision rule for continuous treatments.
problem Developing personalized treatment recommendations for continuous treatments.
method Jump interval-learning method to estimate conditional mean of outcomes.
result Optimal interval-valued decision rule (I2DR) for continuous treatments.
A new method models individual survival curves using conditional normalizing flows.
problem Precise per-individual predictions in survival analysis.
method Conditional normalizing flows for flexible and individualized survival distributions.
result Efficient estimation of individual survival curves without overfitting.
This text discusses several popular explanatory methods that go beyond the error measurements and plots traditionally used to assess machine learning models. Some of the explanatory methods are accepted tools of the trade while others are rigorously derived and backed by long-standing theory. The methods, decision tree…
New insights on Shapley value precision for tabular data predictions.
problem Precision of Shapley value explanations for individual observations.
method Conditional Shapley value estimation methods for tabular data.
result Shapley value explanations are less precise for outer observations.
Optimizes non-linear outcomes from summed contributions.
problem Maximizing a non-linear function of summed small contributions.
method Derives a scalable descent algorithm leveraging concentration properties.
result Directly optimizes for stated objective, e.g., A/B test success criterion.
We introduce a statistical model for operational losses based on heavy-tailed distributions and bipartite graphs, which captures the event type and business line structure of operational risk data. The model explicitly takes into account the Pareto tails of losses and the heterogeneous dependence structures between the…
We propose a new family of fairness definitions for classification problems that combine some of the best properties of both statistical and individual notions of fairness. We posit not only a distribution over individuals, but also a distribution over (or collection of) classification tasks. We then ask that standard …
For a risk vector V, whose components are shared among agents by some random mechanism, we obtain asymptotic lower and upper bounds for the individual agents' exposure risk and the aggregated risk in the market. Risk is measured by Value-at-Risk or Conditional Tail Expectation. We assume Pareto tails for the componen…
Proposes CLIQUE for improved local variable importance in multi-class classification.
problem Lack of methods to characterize local structure in model loss space.
method CLIQUE (Conditional Local Importance by Quantile Expectations)
result CLIQUE emphasizes locally dependent information and captures interaction behavior.
This study examines return and risk of Puerto Rico stock market IRA products.
problem Performance of Puerto Rico stock market IRA products not previously studied.
method Parametric modeling approach estimating conditional expected return and variance.
result PRIRAs underperform the stock market but carry substantial risk.
The Collective Graphical Model (CGM) models a population of independent and identically distributed individuals when only collective statistics (i.e., counts of individuals) are observed. Exact inference in CGMs is intractable, and previous work has explored Markov Chain Monte Carlo (MCMC) and MAP approximations for le…
This research improves demand forecasting by predicting complete probability density functions using machine learning.
problem Forecasting complete probability density functions for better operational decision making.
method Supervised machine learning method 'Cyclic Boosting' for explainable predictions.
result Predicted probability density functions are fully explainable and avoid 'black-box' models.
Social media reduces individual investors' disposition effect through negative information.
problem The disposition effect in individual investors selling profitable assets too early and holding onto losing assets for too long.
method Analysis of post data and trading data from Xueqiu.com.
result Social media information significantly reduces the disposition effect.
Proposes CCE to assess point-wise reliability of neural network predictions.
problem Overconfidence and misaligned predictive distributions in neural networks.
method Introduces Conditional Congruence (CCE) metric using conditional kernel mean embeddings.
result CCE exhibits correctness, monotonicity, reliability, and robustness in high-dimensional regression tasks.
We model the influence of sharing large exogeneous losses to the reinsurance market by a bipartite graph. Using Pareto-tailed claims and multivariate regular variation we obtain asymptotic results for the Value-at-Risk and the Conditional Tail Expectation. We show that the dependence on the network structure plays a fu…
Improves survival prediction model calibration for better individual decision-making.
problem Survival prediction's marginal and conditional calibration issues.
method Conformal prediction using individual survival probabilities.
result Effective marginal and conditional calibration without compromising discrimination.
Develops a prediction method based on sampling design.
problem Creating accurate individual predictions.
method Design-based approach using expected cross-validation results.
result Valid inference of unobserved prediction errors defined with respect to sampling design.
In the world of modern financial theory, portfolio construction has traditionally operated under at least one of two central assumptions: the constraints are derived from a utility function and/or the multivariate probability distribution of the underlying asset returns is fully known. In practice, both the performance…
The paper learns personalized treatment rules from observational data.
problem Developing effective treatment policies for individual patients.
method Contextual bandit approach to minimize expected risk of treatment policies.
result The proposed method outperforms physicians and baseline approaches in IV and VP administration.
Study optimal growth strategies in a continuous-time asset market.
problem Guaranteeing that individual agent strategies cannot outperform the market.
method Mean-field approximation of an infinite number of infinitesimal agents, focusing on optimal strategy distribution among assets.
result Optimal strategy for market agents is to invest proportionally to discounted expected relative dividend intensities.
We find the optimal investment strategy to minimize the expected time that an individual's wealth stays below zero, the so-called {\it occupation time}. The individual consumes at a constant rate and invests in a Black-Scholes financial market consisting of one riskless and one risky asset, with the risky asset's price…