AtC combines human judgments and model scores for better human-centered assessments.
problem Lack of verifiable ground truth in human-centered assessments.
method Two-stage framework: aggregate judgments, then calibrate model scores.
result AtC outperforms human-only or model-only assessments across datasets.
Transformer models improve financial sentiment measurement.
problem Capturing nuanced sentiment from financial news articles.
method Transformer-based language models for sentiment classification and aggregation.
result Transformer models outperform traditional dictionary-based methods in sentiment classification.
New method improves LLM judge accuracy by accounting for dependencies in aggregated binary labels.
problem Classical label aggregation methods fail to account for dependencies among LLM judges, leading to miscalibrated predictions.
method Dependence-aware models based on Ising graphical models and latent factors.
result The proposed method outperforms classical methods on real-world datasets, reducing excess risk.
We propose a probabilistic model to aggregate the answers of respondents answering multiple-choice questions. The model does not assume that everyone has access to the same information, and so does not assume that the consensus answer is correct. Instead, it infers the most probable world state, even if only a minority…
Ranking a set of objects involves establishing an order allowing for comparisons between any pair of objects in the set. Oftentimes, due to the unavailability of a ground truth of ranked orders, researchers resort to obtaining judgments from multiple annotators followed by inferring the ground truth based on the collec…
Framework uses human judgment to distinguish algorithmically indistinguishable cases.
problem Clarifying human-AI collaboration in prediction and decision tasks.
method Integrates human judgment to distinguish algorithmically indistinguishable cases.
result Improves performance of any feasible algorithmic predictor.
Language-based methods improve human similarity approximations without requiring many human judgments.
problem Approximating human similarity judgments using pre-trained deep neural networks (DNNs) is challenging and expensive.
method Developed language-based methods to approximate human similarity judgments, validated with adaptive tag collection pipeline.
result Language-based methods significantly improve performance over DNN-based methods with fewer human judgments.
Large language models predict human sensory judgments across multiple modalities.
problem Determining the extent of perceptual information in language.
method State-of-the-art large language models were used to predict sensory judgments across six psychophysical datasets.
result Large language models can predict human sensory judgments across multiple modalities with significant correlation to human data.
Graphical models improve actuarial judgment in insurance claims analysis.
problem Improving actuarial judgment in insurance claims analysis.
method Using graphical models to represent complex inter-dependencies and incorporate qualitative knowledge.
result Graphical models can be used to express and analyze non-life insurance claims data.
New method interprets object representations from human behavior.
problem Understanding how mental object representations relate to human behavior.
method Sparse, non-negative representations of objects estimated from behavioral judgments.
result Representations predict latent object similarity and are interpretable.
New metric correlates local topic quality with human judgments.
problem Evaluation of topic models focuses on global metrics, ignoring token-level assignments.
method Proposed a human evaluation task and automated metrics to assess local topic quality.
result Consistency metric correlates best with human judgments of local topic quality.
Automatically evaluates image quality based on human judgment.
problem Difficulty in rigorously evaluating generated image quality.
method Generative model embeddings, human labels regression, and statistical matching.
result 66% accuracy in predicting human scores of image realism.
FinRobot AI agent for equity research provides comprehensive insights.
problem Narrow focus and limited discretion in AI solutions for equity research.
method Multi-agent Chain of Thought system integrating quantitative and qualitative analyses.
result FinRobot delivers insights comparable to major brokerage firms.
Approach for assessing supply chain cyber risks using expert judgment and forecasting.
problem Supply chain managers face challenges in assessing cyber risks affecting business factors.
method Structured expert judgment and forecasting models to assess various attack techniques and impacts.
result Facilitates implementation of risk management activities and decision-making processes.
A new process-level model explains how humans detect dependencies quickly.
problem Understanding how humans handle probabilistic independence relations efficiently.
method Developed a rational, distributed, message-passing model called D*.
result D* shows a tendency to quickly detect dependencies, outperforming other algorithms.
In this paper we propose a method for a quantitative estimation of the decision maker's knowledge in the context of the Analytic Hierarchy Process (AHP) in cases, where the judgment matrix is inconsistent. We show that the matrix of deviation from the transitivity condition corresponds to the rate matrix for transactio…
Paper improves PBO using Skew Gaussian Processes for better optimization.
problem Optimizing with preference judgments, especially in A/B tests and recommender systems.
method Uses Skew Gaussian Processes to model preference function and exact posterior inference.
result Exact SkewGP posterior leads to better optimization results than Laplace approximation.
Unified framework to bridge human and LLM judgments.
problem Systematic discrepancies between human and LLM evaluations.
method Latent human preference score and linear transformations of covariates.
result Higher agreement with human ratings and exposure of systematic gaps.
Deep neural network features model human image categorization.
problem Modeling human categorization using natural images.
method Used convolutional neural network features to model human behavior.
result Representations from deep neural networks can model human natural image classifications.
Paper presents datasets from European Court of Human Rights judgments for classification studies.
problem Lack of accessible and reproducible datasets for legal judgments.
method Automated open-source scripts for data collection and feature transformation; experimental campaign on machine learning algorithms.
result Consistently good accuracy (75.86% - 98.32%) across binary datasets, with an average accuracy of 96.45%.
Ordinal embedding methods estimate perceptual scales from relative judgments.
problem Measuring subjective sensation using relative judgments.
method Ordinal embedding from machine learning applied to method of triads.
result Ordinal embedding allows estimating perceptual scales from few judgments, non-monotonous functions, and multi-dimensional scales.
New method approximates Individual Fairness using human judgments.
problem Enforcing fairness in classification tasks.
method Approximates a metric for Individual Fairness based on human queries.
result Constructs hypotheses for metric approximations that generalize.
A new model explains how people solve physical problems and judge hand usage.
problem Understanding how people solve physical problems and judge hand usage.
method Developed a model that plans over symbolic representation, uses geometric solver, and checks feasibility with physical constraints.
result Model explains participants' actions and judgments with high quantitative accuracy.
Machine learning predicts ECHR judgments on human rights violations.
problem Predicting the outcome of ECHR judgments on human rights violations.
method Auto-sklearn for model selection, N-grams, word embeddings, doc2vec, echr2vec for feature extraction, cross-validation for accuracy assessment.
result Features from echr2vec embedding provided the highest cross-validation accuracy for 5 Articles, overall test accuracy was 68.83%.
Paper proposes government indemnification for AI risks to solve judgment-proof problem.
problem Uninsurable risks from AI, especially existential risks, create a judgment-proof problem.
method A government-provided, mandatory indemnification program using risk-priced fees and Bayesian Truth Serum.
result The approach better leverages private information and signals risk mitigation efforts.
The goal of ordinal embedding is to represent items as points in a low-dimensional Euclidean space given a set of constraints in the form of distance comparisons like "item i is closer to item j than item k". Ordinal constraints like this often come from human judgments. To account for errors and variation in jud…
This paper explores LETOR for E-Com search, addressing practical challenges and reporting key findings.
problem Applying LETOR to E-Com search presents unique challenges.
method Investigates practical challenges in LETOR for E-Com search, including feature representation, relevance judgments, and feedback signal exploitation.
result LETOR methods can effectively optimize combinations of popularity-based and relevance-based features, and order rate is the most robust training objective.
Accurate prediction of suicide risk in mental health patients remains an open problem. Existing methods including clinician judgments have acceptable sensitivity, but yield many false positives. Exploiting administrative data has a great potential, but the data has high dimensionality and redundancies in the recording …
LLM evaluation suffers from systematic biases and lacks reliable positive judgments.
problem LLM evaluation suffers from systematic biases and lacks reliable positive judgments.
method Formulate LLM evaluation as a positive-unlabelled learning problem and propose a geometric auditing framework based on Partial Optimal Transport.
result Improved alignment with human preferences, increased robustness to presentation biases, and interpretable confidence estimates.
New risk measure and quadrangle improve financial decision-making.
problem Heterogeneous risk assessments among analysts.
method Established analytical characterizations of WGRM and incorporated FRQ into WRQ.
result WGRM and WRQ framework improves risk-adjusted performance and downside resilience.
Paper operationalizes individual fairness using side-information and a unified representation.
problem Difficulty in eliciting a human specification of a similarity metric for individual fairness.
method Proposes a Pairwise Fair Representation (PFR) model that learns from fairness graph and side-information.
result Unified PFR model effectively operationalizes individual fairness without human specification.
Develops Austen plots for assessing bias from unobserved confounding in observational studies.
problem Bias in causal estimates due to unobserved confounding.
method Formalizes confounding strength, uses Austen plots to visualize and quantify bias.
result Allows domain experts to assess the plausibility of strong confounders.
Proposes RDASS for better Korean text summarization evaluation.
problem ROUGE scores fail to capture semantic meaning in Korean text summarization.
method Introduces RDASS metrics and a method to improve their correlation with human judgment.
result RDASS metrics correlate better with human judgment than ROUGE scores.
Method quantifies relation similarity using entity pair distributions.
problem Measuring similarity between relations in knowledge bases.
method Simple neural network parameterizes conditional probability distributions over entity pairs. Sampling-based approximation for similarity computation.
result Approximation correlates with human judgments and detects redundant relations.
A test measures artificial agents' human-like behavior in video games.
problem Measuring the believability of artificial agents' human-like behavior.
method Developed a non-parametric two-sample hypothesis test.
result The p-value correlates with human judgment of human-like behavior. AI assistants often give convincing but incorrect responses to match user beliefs.
problem Sycophancy in AI assistants that use human feedback.
method Examined five AI assistants across four tasks, analyzed human preference data, and compared model outputs against preference models.
result Sycophancy is a general behavior of AI assistants, driven in part by human preference judgments.
FinReflectKG - EvalBench benchmarks financial KG extraction from SEC 10-K filings.
problem Lack of universal benchmark and evaluation framework for financial KG construction.
method Agentic and holistic evaluation principles, deterministic commit-then-justify judging protocol, binary and ordinal evaluations.
result Reflection-based extraction outperforms single-pass extraction in comprehensiveness, precision, and relevance.
New benchmark for causal reasoning from human video descriptions.
problem Lack of diversity in event types and natural language descriptions, and differences from human judgments.
method Iterative event cloze task and data augmentation techniques.
result Improved data collection efficiency and diverse causal judgments.
Reply to Tetlock et al. on tail risk and probability gap.
problem Expert judgment fails to account for tail risk.
method Comparison of forecasting tournaments and extreme value theory.
result Greater gap between tail expectation and probability properties.
The Frame Problem (FP) is a puzzle in philosophy of mind and epistemology, articulated by the Stanford Encyclopedia of Philosophy as follows: "How do we account for our apparent ability to make decisions on the basis only of what is relevant to an ongoing situation without having explicitly to consider all that is not …
The artistic style of a painting is a subtle aesthetic judgment used by art historians for grouping and classifying artwork. The recently introduced `neural-style' algorithm substantially succeeds in merging the perceived artistic style of one image or set of images with the perceived content of another. In light of th…
This paper considers the challenge of evaluating a set of classifiers, as done in shared task evaluations like the KDD Cup or NIST TREC, without expert labels. While expert labels provide the traditional cornerstone for evaluating statistical learners, limited or expensive access to experts represents a practical bottl…
Novel approach trains LLMs for inductive reasoning using probabilistic programs.
problem Training LLMs for inductive reasoning with sparse, ambiguous data.
method Program-based Posterior Training (PPT) using probabilistic inference.
result Significant improvement in estimation accuracy and alignment with human judgments.
Paper tackles inherent risk scoring with choice-based data labeling and synthetic data collection.
problem Inconsistent expert judgments and lack of labeled data in inherent risk scoring.
method Choice-based data labeling and synthetic data collection.
result System achieves 89% accuracy on a test set of 52 examples.
ValueBlindBench tests LLM-generated investment rationales for validity before returns are known.
problem Delayed-ground-truth evaluation of LLM-generated investment rationales.
method Agreement-gated stress testing protocol to validate LLM-judged rationales.
result ValueBlindBench prevents overclaims and identifies flawed financial constructs.
When dealing with subjective, noisy, or otherwise nebulous features, the "wisdom of crowds" suggests that one may benefit from multiple judgments of the same feature on the same object. We give theoretically-motivated `feature multi-selection' algorithms that choose, among a large set of candidate features, not only wh…
Proposes a new optimization-based method for aggregating sets in neural networks.
problem Limited representational power of existing aggregation methods.
method Equilibrium Aggregation: an optimization-based approach.
result Equilibrium Aggregation outperforms existing methods in various tasks.
Study examines how people perceive fairness in criminal risk prediction algorithms.
problem Concerns about fairness in algorithmic decision making, especially in criminal risk prediction.
method Survey of 576 people to understand perceptions of fairness in algorithmic decision making.
result People's fairness judgments are influenced by eight latent properties of features in algorithms.