Framework uses human judgment to distinguish algorithmically indistinguishable cases.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Approach for assessing supply chain cyber risks using expert judgment and forecasting.
Graphical models improve actuarial judgment in insurance claims analysis.
Paper proposes government indemnification for AI risks to solve judgment-proof problem.
Reply to Tetlock et al. on tail risk and probability gap.
Enhances BO with expert preferences about abstract properties.
A method for eliciting expert beliefs using preferential questions and normalizing flows.
This paper considers the challenge of evaluating a set of classifiers, as done in shared task evaluations like the KDD Cup or NIST TREC, without expert labels. While expert labels provide the traditional cornerstone for evaluating statistical learners, limited or expensive access to experts represents a practical bottl…
Inherent risk scoring is an important function in anti-money laundering, used for determining the riskiness of an individual during onboarding fraudulent transactions occur. It is, however, often fraught with two challenges: (1) inconsistent notions of what constitutes as high or low risk by experts a…
Potential violent criminals will often need to go through a sequence of preparatory steps before they can execute their plans. During this escalation process police have the opportunity to evaluate the threat posed by such people through what they know, observe and learn from intelligence reports about their activities…
Develops Austen plots for assessing bias from unobserved confounding in observational studies.
Proposes a new bankruptcy prediction model using Bayesian framework with expert knowledge.
Language-based methods improve human similarity approximations without requiring many human judgments.
Large language models predict human sensory judgments across multiple modalities.
No-knowledge alarms detect misaligned LLM judges without trusting them.
To study how mental object representations are related to behavior, we estimated sparse, non-negative representations of objects using human behavioral judgments on images representative of 1,854 object categories. These representations predicted a latent similarity structure between objects, which captured most of the…
This paper presents the development of a hybrid learning system based on Support Vector Machines (SVM), Adaptive Neuro-Fuzzy Inference System (ANFIS) and domain knowledge to solve prediction problem. The proposed two-stage Domain Knowledge based Fuzzy Information System (DKFIS) improves the prediction accuracy attained…
Automatically evaluates image quality based on human judgment.
In this paper we propose a method for a quantitative estimation of the decision maker's knowledge in the context of the Analytic Hierarchy Process (AHP) in cases, where the judgment matrix is inconsistent. We show that the matrix of deviation from the transitivity condition corresponds to the rate matrix for transactio…
Paper improves PBO using Skew Gaussian Processes for better optimization.
Unified framework to bridge human and LLM judgments.
DBOT uses AI to automate long-term stock valuation.
Topic models are typically evaluated with respect to the global topic distributions that they generate, using metrics such as coherence, but without regard to local (token-level) topic assignments. Token-level assignments are important for downstream tasks such as classification. Even recent models, which aim to improv…
It is inconceivable how chaotic the world would look to humans, faced with innumerable decisions a day to be made under uncertainty, had they been lacking the capacity to distinguish the relevant from the irrelevant---a capacity which computationally amounts to handling probabilistic independence relations. The highly …
Machine learning predicts ECHR judgments on human rights violations.
Tree-Query uses LLMs to discover causal relationships in a transparent, interpretable manner.
The goal of ordinal embedding is to represent items as points in a low-dimensional Euclidean space given a set of constraints in the form of distance comparisons like "item is closer to item than item ". Ordinal constraints like this often come from human judgments. To account for errors and variation in jud…
Bayesian taut splines estimate modes in probability densities.
Bayesian model averaging (BMA) is the state of the art approach for overcoming model uncertainty. Yet, especially on small data sets, the results yielded by BMA might be sensitive to the prior over the models. Credal Model Averaging (CMA) addresses this problem by substituting the single prior over the models by a set …
Comparative study of neural networks for short-term FOREX forecasting.
Accurate prediction of suicide risk in mental health patients remains an open problem. Existing methods including clinician judgments have acceptable sensitivity, but yield many false positives. Exploiting administrative data has a great potential, but the data has high dimensionality and redundancies in the recording …
Transformer models improve financial sentiment measurement.
Over the last few decades, psychologists have developed sophisticated formal models of human categorization using simple artificial stimuli. In this paper, we use modern machine learning methods to extend this work into the realm of naturalistic stimuli, enabling human categorization to be studied over the complex visu…
This paper proposes a simple technical approach for the analytical derivation of Point-in-Time PD (probability of default) forecasts, with minimal data requirements. The inputs required are the current and future Through-the-Cycle PDs of the obligors, their last known default rates, and a measurement of the systematic …
In this paper, we address the problem of measuring and analysing sensation, the subjective magnitude of one's experience. We do this in the context of the method of triads: the sensation of the stimulus is evaluated via relative judgments of the form: "Is stimulus S_i more similar to stimulus S_j or to stimulus S_k?". …
CJE calibrates cheap LLM judges against an oracle, achieving high accuracy at a fraction of the cost.
We revisit the notion of individual fairness proposed by Dwork et al. A central challenge in operationalizing their approach is the difficulty in eliciting a human specification of a similarity metric. In this paper, we propose an operationalization of individual fairness that does not rely on a human specification of …
LLM evaluation suffers from systematic biases and lacks reliable positive judgments.
The macroeconomic climate influences operations with regard to, e.g., raw material prices, financing, supply chain utilization and demand quotas. In order to adapt to the economic environment, decision-makers across the public and private sectors require accurate forecasts of the economic outlook. Existing predictive f…
E-Commerce (E-Com) search is an emerging important new application of information retrieval. Learning to Rank (LETOR) is a general effective strategy for optimizing search engines, and is thus also a key technology for E-Com search. While the use of LETOR for web search has been well studied, its use for E-Com search h…
In this paper, we present a new task that investigates how people interact with and make judgments about towers of blocks. In Experiment~1, participants in the lab solved a series of problems in which they had to re-configure three blocks from an initial to a final configuration. We recorded whether they used one hand …
Proposes RDASS for better Korean text summarization evaluation.
A test measures artificial agents' human-like behavior in video games.
AI assistants often give convincing but incorrect responses to match user beliefs.
Multi-expert L2D underfits more severely, requiring new methods.
New benchmark for causal reasoning from human video descriptions.
Ranking a set of objects involves establishing an order allowing for comparisons between any pair of objects in the set. Oftentimes, due to the unavailability of a ground truth of ranked orders, researchers resort to obtaining judgments from multiple annotators followed by inferring the ground truth based on the collec…
TENP prunes experts and neurons in Mixture-of-Experts models for efficient deployment.