This paper develops a new method for eliciting more flexible metrics, improving fairness and applicability.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Proposes a method to select fair performance metrics through metric elicitation.
Study creates web interface to elicit user-preferred metrics.
This thesis formalizes metric selection for machine learning applications.
Given a binary prediction problem, which performance metric should the classifier optimize? We address this question by formalizing the problem of Metric Elicitation. The goal of metric elicitation is to discover the performance metric of a practitioner, which reflects her innate rewards (costs) for correct (incorrect)…
New scoring rules compare probabilistic top lists in classification.
We revisit the notion of individual fairness proposed by Dwork et al. A central challenge in operationalizing their approach is the difficulty in eliciting a human specification of a similarity metric. In this paper, we propose an operationalization of individual fairness that does not rely on a human specification of …
Framework uses IRL and RL to elicit and optimize risk preferences robustly to noise.
Constructs new elicitable risk measures with multiplicative scoring functions.
A property, or statistical functional, is said to be elicitable if it minimizes expected loss for some loss function. The study of which properties are elicitable sheds light on the capabilities and limitations of point estimation and empirical risk minimization. While recent work asks which properties are elicitable, …
We discuss equivalent axiomatic characterizations of distortion risk measures, and give a novel and concise proof of the characterization of elicitable distortion risk measures. Elicitability has recently been discussed as a desirable criterion for risk measures, motivated by statistical considerations of forecasting. …
Robustifies elicitable functionals to handle small distribution misspecifications.
Study generalizes property elicitation to imprecise probabilities.
A statistical functional, such as the mean or the median, is called elicitable if there is a scoring function or loss function such that the correct forecast of the functional is the unique minimizer of the expected score. Such scoring functions are called strictly consistent for the functional. The elicitability of a …
It is important to collect credible training samples for building data-intensive learning systems (e.g., a deep learning system). Asking people to report complex distribution , though theoretically viable, is challenging in practice. This is primarily due to the cognitive loads required for human agents t…
A method for eliciting expert beliefs using preferential questions and normalizing flows.
Paper shows similarity learning can lead to strong binary classification performance.
The paper analyzes elicitability of return risk measures and their scoring functions.
Proposes method for eliciting non-parametric joint priors using normalizing flows.
The risk of a financial position is usually summarized by a risk measure. As this risk measure has to be estimated from historical data, it is important to be able to verify and compare competing estimation procedures. In statistical decision theory, risk measures for which such verification and comparison is possible,…
Method combines deep learning and elicitability for solving complex stochastic equations.
Develops a simulation-based method to translate expert knowledge into prior distributions for Bayesian models.
Paper establishes identifiability and elicitability of tail risk measures.
In this note, we comment on the relevance of elicitability for backtesting risk measure estimates. In particular, we propose the use of Diebold-Mariano tests, and show how they can be implemented for Expected Shortfall (ES), based on the recent result of Fissler and Ziegel (2015) that ES is jointly elicitable with Valu…
We consider settings in which the right notion of fairness is not captured by simple mathematical definitions (such as equality of error rates across groups), but might be more complex and nuanced and thus require elicitation from individual or collective stakeholders. We introduce a framework in which pairs of individ…
New algorithm achieves faster multicalibration in online settings.
New method allows backtesting of systemic risk forecasts.
We propose a cost-effective framework for preference elicitation and aggregation under the Plackett-Luce model with features. Given a budget, our framework iteratively computes the most cost-effective elicitation questions in order to help the agents make a better group decision. We illustrate the viability of the fram…
Formulates a Dueling Bandits problem for eliciting Kemeny rankings.
Paper proposes incentives for federated learning to ensure truthful contributions.
A framework for eliciting utility functions from investor preferences.
Learning predictive models from small high-dimensional data sets is a key problem in high-dimensional statistics. Expert knowledge elicitation can help, and a strong line of work focuses on directly eliciting informative prior distributions for parameters. This either requires considerable statistical expertise or is l…
Requirements elicitation can be very challenging in projects that require deep domain knowledge about the system at hand. As analysts have the full control over the elicitation process, their lack of knowledge about the system under study inhibits them from asking related questions and reduces the accuracy of requireme…
Providing accurate predictions is challenging for machine learning algorithms when the number of features is larger than the number of samples in the data. Prior knowledge can improve machine learning models by indicating relevant variables and parameter values. Yet, this prior knowledge is often tacit and only availab…
Study uses property elicitation to understand how fairness regularizers affect optimal decisions.
PPT optimizes transformer behavior by steering its latent posterior using prior samples.
In this paper we propose an approach to preference elicitation that is suitable to large configuration spaces beyond the reach of existing state-of-the-art approaches. Our setwise max-margin method can be viewed as a generalization of max-margin learning to sets, and can produce a set of "diverse" items that can be use…
Crowdsourced wisdom improves causal learning.
Generative Adversarial Regression (GAR) learns risk scenarios robustly across policies.
High-dimensional prediction is a challenging problem setting for traditional statistical models. Although regularization improves model performance in high dimensions, it does not sufficiently leverage knowledge on feature importances held by domain experts. As an alternative to standard regularization techniques, we p…
Platform uses queries to elicit investor preferences for portfolio trades, improving allocation efficiency.
Requirements elicitation requires extensive knowledge and deep understanding of the problem domain where the final system will be situated. However, in many software development projects, analysts are required to elicit the requirements from an unfamiliar domain, which often causes communication barriers between analys…
ContextBench benchmarks methods for generating linguistically fluent inputs that activate specific latent features in language models.
Proposes a new framework for risk-sensitive RL using deep nets.
Recently, financial industry and regulators have enhanced the debate on the good properties of a risk measure. A fundamental issue is the evaluation of the quality of a risk estimation. On the one hand, a backtesting procedure is desirable for assessing the accuracy of such an estimation and this can be naturally achie…
This paper explores how to choose scoring rules for estimating properties with parametric assumptions.
Topic models are typically evaluated with respect to the global topic distributions that they generate, using metrics such as coherence, but without regard to local (token-level) topic assignments. Token-level assignments are important for downstream tasks such as classification. Even recent models, which aim to improv…
Study shows LLM-advisors match human performance in eliciting preferences but struggle with conflicting needs and trust.