It is important to collect credible training samples for building data-intensive learning systems (e.g., a deep learning system). Asking people to report complex distribution , though theoretically viable, is challenging in practice. This is primarily due to the cognitive loads required for human agents t…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper proposes incentives for federated learning to ensure truthful contributions.
New model for recovering unverifiable signals from observers in decentralized networks.
Framework improves self-play for cooperative multi-agent learning.
Some online advertising offers pay only when an ad elicits a response. Randomness and uncertainty about response rates make showing those ads a risky investment for online publishers. Like financial investors, publishers can use portfolio allocation over multiple advertising offers to pursue revenue while controlling r…
We consider the problem of fitting a linear model to data held by individuals who are concerned about their privacy. Incentivizing most players to truthfully report their data to the analyst constrains our design to mechanisms that provide a privacy guarantee to the participants; we use differential privacy to model in…
This paper develops a new method for eliciting more flexible metrics, improving fairness and applicability.
Constructs new elicitable risk measures with multiplicative scoring functions.
A property, or statistical functional, is said to be elicitable if it minimizes expected loss for some loss function. The study of which properties are elicitable sheds light on the capabilities and limitations of point estimation and empirical risk minimization. While recent work asks which properties are elicitable, …
Proposes a method to select fair performance metrics through metric elicitation.
We discuss equivalent axiomatic characterizations of distortion risk measures, and give a novel and concise proof of the characterization of elicitable distortion risk measures. Elicitability has recently been discussed as a desirable criterion for risk measures, motivated by statistical considerations of forecasting. …
Robustifies elicitable functionals to handle small distribution misspecifications.
Study generalizes property elicitation to imprecise probabilities.
A statistical functional, such as the mean or the median, is called elicitable if there is a scoring function or loss function such that the correct forecast of the functional is the unique minimizer of the expected score. Such scoring functions are called strictly consistent for the functional. The elicitability of a …
Study creates web interface to elicit user-preferred metrics.
A method for eliciting expert beliefs using preferential questions and normalizing flows.
Given a binary prediction problem, which performance metric should the classifier optimize? We address this question by formalizing the problem of Metric Elicitation. The goal of metric elicitation is to discover the performance metric of a practitioner, which reflects her innate rewards (costs) for correct (incorrect)…
The paper analyzes elicitability of return risk measures and their scoring functions.
Proposes method for eliciting non-parametric joint priors using normalizing flows.
This thesis formalizes metric selection for machine learning applications.
The risk of a financial position is usually summarized by a risk measure. As this risk measure has to be estimated from historical data, it is important to be able to verify and compare competing estimation procedures. In statistical decision theory, risk measures for which such verification and comparison is possible,…
Method combines deep learning and elicitability for solving complex stochastic equations.
Develops a simulation-based method to translate expert knowledge into prior distributions for Bayesian models.
Eliciting labels from crowds is a potential way to obtain large labeled data. Despite a variety of methods developed for learning from crowds, a key challenge remains unsolved: \emph{learning from crowds without knowing the information structure among the crowds a priori, when some people of the crowds make highly corr…
Paper establishes identifiability and elicitability of tail risk measures.
In this note, we comment on the relevance of elicitability for backtesting risk measure estimates. In particular, we propose the use of Diebold-Mariano tests, and show how they can be implemented for Expected Shortfall (ES), based on the recent result of Fissler and Ziegel (2015) that ES is jointly elicitable with Valu…
New method evaluates language model forecasters by checking consistency of predictions.
We consider settings in which the right notion of fairness is not captured by simple mathematical definitions (such as equality of error rates across groups), but might be more complex and nuanced and thus require elicitation from individual or collective stakeholders. We introduce a framework in which pairs of individ…
New algorithm achieves faster multicalibration in online settings.
Framework uses IRL and RL to elicit and optimize risk preferences robustly to noise.
Duel-Evolve uses LLM self-preferences for test-time optimization of discrete outputs.
New method allows backtesting of systemic risk forecasts.
We propose a cost-effective framework for preference elicitation and aggregation under the Plackett-Luce model with features. Given a budget, our framework iteratively computes the most cost-effective elicitation questions in order to help the agents make a better group decision. We illustrate the viability of the fram…
Formulates a Dueling Bandits problem for eliciting Kemeny rankings.
New design detects confounders from treatment intent in ICU data.
In this work, we train fully convolutional networks to detect anger in speech. Since training these deep architectures requires large amounts of data and the size of emotion datasets is relatively small, we use transfer learning. However, unlike previous approaches that use speech or emotion-based tasks for the source …
A framework for eliciting utility functions from investor preferences.
Learning predictive models from small high-dimensional data sets is a key problem in high-dimensional statistics. Expert knowledge elicitation can help, and a strong line of work focuses on directly eliciting informative prior distributions for parameters. This either requires considerable statistical expertise or is l…
Requirements elicitation can be very challenging in projects that require deep domain knowledge about the system at hand. As analysts have the full control over the elicitation process, their lack of knowledge about the system under study inhibits them from asking related questions and reduces the accuracy of requireme…
Providing accurate predictions is challenging for machine learning algorithms when the number of features is larger than the number of samples in the data. Prior knowledge can improve machine learning models by indicating relevant variables and parameter values. Yet, this prior knowledge is often tacit and only availab…
Study uses property elicitation to understand how fairness regularizers affect optimal decisions.
PPT optimizes transformer behavior by steering its latent posterior using prior samples.
In this paper we propose an approach to preference elicitation that is suitable to large configuration spaces beyond the reach of existing state-of-the-art approaches. Our setwise max-margin method can be viewed as a generalization of max-margin learning to sets, and can produce a set of "diverse" items that can be use…
Crowdsourced wisdom improves causal learning.
Generative Adversarial Regression (GAR) learns risk scenarios robustly across policies.
Platform uses queries to elicit investor preferences for portfolio trades, improving allocation efficiency.
Requirements elicitation requires extensive knowledge and deep understanding of the problem domain where the final system will be situated. However, in many software development projects, analysts are required to elicit the requirements from an unfamiliar domain, which often causes communication barriers between analys…
ContextBench benchmarks methods for generating linguistically fluent inputs that activate specific latent features in language models.