YRC-Bench benchmarks AI agents learning to collaborate with experts.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
XLabel tool reduces medical experts' workload by 40% and explains its decisions.
Humans prove theorems by relying on substantial high-level reasoning and problem-specific insights. Proof assistants offer a formalism that resembles human mathematical reasoning, representing theorems in higher-order logic and proofs as high-level tactics. However, human experts have to construct proofs manually by en…
Amyotrophic Lateral Sclerosis (ALS) is a neurodegenerative disease characterized by a rapid motor decline, leading to respiratory failure and subsequently to death. In this context, researchers have sought for models to automatically predict disease progression to assisted ventilation in ALS patients. However, the clin…
An assistant learns to mediate decisions between humans and experts, balancing risk and learning.
TacticAI helps football coaches improve tactics by analyzing corner kicks.
Despite the advantages of all-weather and all-day high-resolution imaging, SAR remote sensing images are much less viewed and used by general people because human vision is not adapted to microwave scattering phenomenon. However, expert interpreters can be trained by compare side-by-side SAR and optical images to learn…
Novel algorithms for online learning with uncertain feedback graphs reduce regret.
Evaluating surgeon skill has predominantly been a subjective task. Development of objective methods for surgical skill assessment are of increased interest. Recently, with technological advances such as robotic-assisted minimally invasive surgery (RMIS), new opportunities for objective and automated assessment framewor…
Automated suggestions help train technicians diagnose incidents faster.
In many contexts, it can be useful for domain experts to understand to what extent predictions made by a machine learning model can be trusted. In particular, estimates of trustworthiness can be useful for fraud analysts who process machine learning-generated alerts of fraudulent transactions. In this work, we present …
We explore the problem of learning under selective labels in the context of algorithm-assisted decision making. Selective labels is a pervasive selection bias problem that arises when historical decision making blinds us to the true outcome for certain instances. Examples of this are common in many applications, rangin…
Study shows human advisors use context to improve student outcomes in algorithm-assisted advising.
AI helps complete ancient tablets with missing parts.
System helps engineers with concept recognition for SEVA.
Anomaly detection algorithms are often thought to be limited because they don't facilitate the process of validating results performed by domain experts. In Contrast, deep learning algorithms for anomaly detection, such as autoencoders, point out the outliers, saving experts the time-consuming task of examining normal …
End-to-end CAD system for thyroid nodule classification using multimodal data and expert guidance.
Generative adversarial network system improves ECG arrhythmia classification.
AACE learns treatment policies from EHRs using annotations to improve accuracy.
Consider an assistive system that guides visually impaired users through speech and haptic feedback to their destination. Existing robotic and ubiquitous navigation technologies (e.g., portable, ground, or wearable systems) often operate in a generic, user-agnostic manner. However, to minimize confusion and navigation …
Automatic anomaly detection is a major issue in various areas. Beyond mere detection, the identification of the origin of the problem that produced the anomaly is also essential. This paper introduces a general methodology that can assist human operators who aim at classifying monitoring signals. The main idea is to le…
Optimizes classifiers for varying levels of automation.
We present DRLViz, a visual analytics interface to interpret the internal memory of an agent (e.g. a robot) trained using deep reinforcement learning. This memory is composed of large temporal vectors updated when the agent moves in an environment and is not trivial to understand due to the number of dimensions, depend…
The paper addresses the difficulty of decision makers trusting AI-assisted predictions and proposes a method to improve confidence values.
MetaStackVis aids in choosing better metamodels for stacking ensembles.
LR-Robot automates SLRs with AI, expert oversight, and multidimensional analysis.
In the last few years, we have seen the transformative impact of deep learning in many applications, particularly in speech recognition and computer vision. Inspired by Google's Inception-ResNet deep convolutional neural network (CNN) for image classification, we have developed "Chemception", a deep CNN for the predict…
Underlying cause of death coding from death certificates is a process that is nowadays undertaken mostly by humans with a potential assistance from expert systems such as the Iris software. It is as a consequence an expensive process that can in addition suffer from geospatial discrepancies, thus severely impairing the…
Adversarial policies beat superhuman Go AI systems.
LR-Robot accelerates SLRs by combining expert oversight and AI, revealing trends and patterns in financial research.
Plud system reduces labeling time and produces accurate models for uncategorized images.
New AI assistant for power grid operators simplifies complex decision-making.
Computer-aided detection has been a research area attracting great interest in the past decade. Machine learning algorithms have been utilized extensively for this application as they provide a valuable second opinion to the doctors. Despite several machine learning models being available for medical imaging applicatio…
In semantic parsing for question-answering, it is often too expensive to collect gold parses or even gold answers as supervision signals. We propose to convert model outputs into a set of human-understandable statements which allow non-expert users to act as proofreaders, providing error markings as learning signals to…
Learning to drive faithfully in highly stochastic urban settings remains an open problem. To that end, we propose a Multi-task Learning from Demonstration (MT-LfD) framework which uses supervised auxiliary task prediction to guide the main task of predicting the driving commands. Our framework involves an end-to-end tr…
Framework allows organizations to collaborate on learning tasks securely.
Deep reinforcement learning has demonstrated increasing capabilities for continuous control problems, including agents that can move with skill and agility through their environment. An open problem in this setting is that of developing good strategies for integrating or merging policies for multiple skills, where each…
Improves survey sampling with unbiased machine learning methods.
LLMs learn to recommend models and hyperparameters from dataset metadata.
Computer-assisted method finds new Einstein metrics on spheres.
Access to food assistance programs such as food pantries and food banks needs focus in order to mitigate food insecurity. Accessibility to the food assistance programs is impacted by demographics of the population and geography of the location. It hence becomes imperative to define and identify food assistance deserts …
Learning preferences implicit in the choices humans make is a well studied problem in both economics and computer science. However, most work makes the assumption that humans are acting (noisily) optimally with respect to their preferences. Such approaches can fail when people are themselves learning about what they wa…
Study evaluates the impact of academic support center's face-to-face assistance on student performance.
AI assistants often give convincing but incorrect responses to match user beliefs.
Develops a two-layer model to design mortgage assistance products.
Helps visually impaired users make better decisions by adjusting their observations.
Multi-expert L2D underfits more severely, requiring new methods.
Robust Bayes-Assisted Conformal Prediction improves prediction set sizes.