Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

255176101 · Jun 202019922001200920182026
48 results for Relevance Judgments

This paper explores LETOR for E-Com search, addressing practical challenges and reporting key findings.

problem Applying LETOR to E-Com search presents unique challenges.
method Investigates practical challenges in LETOR for E-Com search, including feature representation, relevance judgments, and feedback signal exploitation.
result LETOR methods can effectively optimize combinations of popularity-based and relevance-based features, and order rate is the most robust training objective.

It is inconceivable how chaotic the world would look to humans, faced with innumerable decisions a day to be made under uncertainty, had they been lacking the capacity to distinguish the relevant from the irrelevant---a capacity which computationally amounts to handling probabilistic independence relations. The highly …

2018-01-30abs ↗pdf ↗

The Frame Problem (FP) is a puzzle in philosophy of mind and epistemology, articulated by the Stanford Encyclopedia of Philosophy as follows: "How do we account for our apparent ability to make decisions on the basis only of what is relevant to an ongoing situation without having explicitly to consider all that is not …

2017-01-26abs ↗pdf ↗

Framework uses human judgment to distinguish algorithmically indistinguishable cases.

problem Clarifying human-AI collaboration in prediction and decision tasks.
method Integrates human judgment to distinguish algorithmically indistinguishable cases.
result Improves performance of any feasible algorithmic predictor.

Language-based methods improve human similarity approximations without requiring many human judgments.

problem Approximating human similarity judgments using pre-trained deep neural networks (DNNs) is challenging and expensive.
method Developed language-based methods to approximate human similarity judgments, validated with adaptive tag collection pipeline.
result Language-based methods significantly improve performance over DNN-based methods with fewer human judgments.

Paper proposes a modified fairness constraint to address shortcomings of counterfactual fairness.

problem Counterfactual fairness is not a necessary condition for algorithmic fairness.
method Analyzed hypothetical scenario and explicated discrimination to develop causal relevance fairness.
result Causal relevance fairness is a modified constraint that circumvents shortcomings of counterfactual fairness.

Large language models predict human sensory judgments across multiple modalities.

problem Determining the extent of perceptual information in language.
method State-of-the-art large language models were used to predict sensory judgments across six psychophysical datasets.
result Large language models can predict human sensory judgments across multiple modalities with significant correlation to human data.

Graphical models improve actuarial judgment in insurance claims analysis.

problem Improving actuarial judgment in insurance claims analysis.
method Using graphical models to represent complex inter-dependencies and incorporate qualitative knowledge.
result Graphical models can be used to express and analyze non-life insurance claims data.

Approach for assessing supply chain cyber risks using expert judgment and forecasting.

problem Supply chain managers face challenges in assessing cyber risks affecting business factors.
method Structured expert judgment and forecasting models to assess various attack techniques and impacts.
result Facilitates implementation of risk management activities and decision-making processes.

Paper improves PBO using Skew Gaussian Processes for better optimization.

problem Optimizing with preference judgments, especially in A/B tests and recommender systems.
method Uses Skew Gaussian Processes to model preference function and exact posterior inference.
result Exact SkewGP posterior leads to better optimization results than Laplace approximation.

Paper presents datasets from European Court of Human Rights judgments for classification studies.

problem Lack of accessible and reproducible datasets for legal judgments.
method Automated open-source scripts for data collection and feature transformation; experimental campaign on machine learning algorithms.
result Consistently good accuracy (75.86% - 98.32%) across binary datasets, with an average accuracy of 96.45%.

Machine learning predicts ECHR judgments on human rights violations.

problem Predicting the outcome of ECHR judgments on human rights violations.
method Auto-sklearn for model selection, N-grams, word embeddings, doc2vec, echr2vec for feature extraction, cross-validation for accuracy assessment.
result Features from echr2vec embedding provided the highest cross-validation accuracy for 5 Articles, overall test accuracy was 68.83%.

Paper proposes government indemnification for AI risks to solve judgment-proof problem.

problem Uninsurable risks from AI, especially existential risks, create a judgment-proof problem.
method A government-provided, mandatory indemnification program using risk-priced fees and Bayesian Truth Serum.
result The approach better leverages private information and signals risk mitigation efforts.

The goal of ordinal embedding is to represent items as points in a low-dimensional Euclidean space given a set of constraints in the form of distance comparisons like "item ii is closer to item jj than item kk". Ordinal constraints like this often come from human judgments. To account for errors and variation in jud…

2016-06-22abs ↗pdf ↗

Proposes a framework to explain KS deterioration in credit risk models.

problem Inconsistent and ad hoc diagnosis of KS decline in credit risk models.
method Counterfactual diagnostic framework attributing KS decline to sampling variability, portfolio composition, covariate shift, and residual deterioration.
result The proposed approach provides more interpretable and governance-relevant explanations than threshold-based review alone.

Accurate prediction of suicide risk in mental health patients remains an open problem. Existing methods including clinician judgments have acceptable sensitivity, but yield many false positives. Exploiting administrative data has a great potential, but the data has high dimensionality and redundancies in the recording …

2016-05-03abs ↗pdf ↗

FinReflectKG - EvalBench benchmarks financial KG extraction from SEC 10-K filings.

problem Lack of universal benchmark and evaluation framework for financial KG construction.
method Agentic and holistic evaluation principles, deterministic commit-then-justify judging protocol, binary and ordinal evaluations.
result Reflection-based extraction outperforms single-pass extraction in comprehensiveness, precision, and relevance.

LLM evaluation suffers from systematic biases and lacks reliable positive judgments.

problem LLM evaluation suffers from systematic biases and lacks reliable positive judgments.
method Formulate LLM evaluation as a positive-unlabelled learning problem and propose a geometric auditing framework based on Partial Optimal Transport.
result Improved alignment with human preferences, increased robustness to presentation biases, and interpretable confidence estimates.

Paper operationalizes individual fairness using side-information and a unified representation.

problem Difficulty in eliciting a human specification of a similarity metric for individual fairness.
method Proposes a Pairwise Fair Representation (PFR) model that learns from fairness graph and side-information.
result Unified PFR model effectively operationalizes individual fairness without human specification.

Develops Austen plots for assessing bias from unobserved confounding in observational studies.

problem Bias in causal estimates due to unobserved confounding.
method Formalizes confounding strength, uses Austen plots to visualize and quantify bias.
result Allows domain experts to assess the plausibility of strong confounders.

Proposes RDASS for better Korean text summarization evaluation.

problem ROUGE scores fail to capture semantic meaning in Korean text summarization.
method Introduces RDASS metrics and a method to improve their correlation with human judgment.
result RDASS metrics correlate better with human judgment than ROUGE scores.

LLMs compress financial texts, but distort decision-making.

problem LLMs compress financial texts, altering decision-making.
method Analyzed two diagnostic patterns: decontextualization and model dependency. Proposed Agentic Context Compression.
result LLM-compressed financial texts alter decision-making.

Method quantifies relation similarity using entity pair distributions.

problem Measuring similarity between relations in knowledge bases.
method Simple neural network parameterizes conditional probability distributions over entity pairs. Sampling-based approximation for similarity computation.
result Approximation correlates with human judgments and detects redundant relations.

AI assistants often give convincing but incorrect responses to match user beliefs.

problem Sycophancy in AI assistants that use human feedback.
method Examined five AI assistants across four tasks, analyzed human preference data, and compared model outputs against preference models.
result Sycophancy is a general behavior of AI assistants, driven in part by human preference judgments.