Proposes PA-DSL for correcting noisy human labels in automated data labeling.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Human input has enabled autonomous systems to improve their capabilities and achieve complex behaviors that are otherwise challenging to generate automatically. Recent work focuses on how robots can use such input - like demonstrations or corrections - to learn intended objectives. These techniques assume that the huma…
Learning robot objective functions from human input has become increasingly important, but state-of-the-art techniques assume that the human's desired objective lies within the robot's hypothesis space. When this is not true, even methods that keep track of uncertainty over the objective fail because they reason about …
Deep Reinforcement Learning (DRL) has become a powerful strategy to solve complex decision making problems based on Deep Neural Networks (DNNs). However, it is highly data demanding, so unfeasible in physical systems for most applications. In this work, we approach an alternative Interactive Machine Learning (IML) stra…
Experiment shows cognitive biases impact human-AI collaboration, highlighting the need for diverse evaluator samples.
Theoretical analysis shows LLMs can self-correct responses through in-context learning.
Corrects bias in LLM-as-a-judge evaluations using adaptive calibration.
Paper proposes an efficient method for bounding box annotation in object detection.
Optimal allocation of human effort to correct AI assessments in decision-making.
This work advances collaborative decision making by combining human and AI strengths in uncertainty quantification.
Explainable AI improves human decision accuracy but does not enhance it significantly.
Owing to the advancement of deep learning, artificial systems are now rival to humans in several pattern recognition tasks, such as visual recognition of object categories. However, this is only the case with the tasks for which correct answers exist independent of human perception. There is another type of tasks for w…
Develops methods to correct bias in AI feedback for more accurate alignment.
Paper stabilizes generative model training with synthetic data.
Develops a new tensor model for clustering with degree correction.
New method uses explicit human demonstrations to teach missing features in reward learning.
Learning from human feedback is a viable alternative to control design that does not require modelling or control expertise. Particularly, learning from corrective advice garners advantages over evaluative feedback as it is a more intuitive and scalable format. The current state-of-the-art in this field, COACH, has pro…
We present an approach to interactive-predictive neural machine translation that attempts to reduce human effort from three directions: Firstly, instead of requiring humans to select, correct, or delete segments, we employ the idea of learning from human reinforcements in form of judgments on the quality of partial tra…
A new L2D system produces calibrated probabilities of expert correctness without sacrificing accuracy.
New approach shows AI can adapt like toddlers by correcting old knowledge.
AI assistants often give convincing but incorrect responses to match user beliefs.
Deep Reinforcement Learning has enabled the control of increasingly complex and high-dimensional problems. However, the need of vast amounts of data before reasonable performance is attained prevents its widespread application. We employ binary corrective feedback as a general and intuitive manner to incorporate human …
The concept of progress has characterized human society from millennia. However, this concept is elusive and too often given for certain. The goal of this paper is to suggest a general definition of human progress that satisfies, whenever possible the conditions of independence, generality, epistemological applicabilit…
LLM evaluation suffers from systematic biases and lacks reliable positive judgments.
This paper improves deep learning model consistency through ensemble methods.
Recent years have seen a boom in interest in machine learning systems that can provide a human-understandable rationale for their predictions or decisions. However, exactly what kinds of explanation are truly human-interpretable remains poorly understood. This work advances our understanding of what makes explanations …
We consider the problem of learning object arrangements in a 3D scene. The key idea here is to learn how objects relate to human poses based on their affordances, ease of use and reachability. In contrast to modeling object-object relationships, modeling human-object relationships scales linearly in the number of objec…
We conduct a large-scale, systematic study to evaluate the existing evaluation methods for natural language generation in the context of generating online product reviews. We compare human-based evaluators with a variety of automated evaluation procedures, including discriminative evaluators that measure how well machi…
While linear mixed model (LMM) has shown a competitive performance in correcting spurious associations raised by population stratification, family structures, and cryptic relatedness, more challenges are still to be addressed regarding the complex structure of genotypic and phenotypic data. For example, geneticists hav…
We propose Disentanglement based Active Learning (DAL), a new active learning technique based on self-supervision which leverages the concept of disentanglement. Instead of requesting labels from human oracle, our method automatically labels the majority of the datapoints, thus drastically reducing the human labeling b…
IAL uses interactive learning to improve model performance with minimal human feedback.
A new watermarking method corrects bias in language models using maximal coupling.
TIMELY improves consistency in labeling blood cell images.
STAS selects optimal spatio-temporal scales for bias correction in precipitation forecasts.
In this study, we propose a novel deep neural network and its supervised learning method that uses a feedforward supervisory signal. The method is inspired by the human visual system and performs human-like association-based learning without any backward error propagation. The feedforward supervisory signal that produc…
Machines, not humans, are the world's dominant knowledge accumulators but humans remain the dominant decision makers. Interpreting and disseminating the knowledge accumulated by machines requires expertise, time, and is prone to failure. The problem of how best to convey accumulated knowledge from computers to humans i…
Smartphones have been the most popular and widely used devices among means of communication. Nowadays, human activity recognition is possible on mobile devices by embedded sensors, which can be exploited to manage user behavior on mobile devices by predicting user activity. To reach this aim, storing activity character…
SRPO improves AI alignment with human preferences through self-improvement and task-independent optimization.
This research explores inductive biases for deep learning to improve AI's higher-level cognition.
In this paper, we propose a game theoretical adversarial intervention detection mechanism for reliable smart road signs. A future trend in intelligent transportation systems is ``smart road signs" that incorporate smart codes (e.g., visible at infrared) on their surface to provide more detailed information to smart veh…
k-Rater reliability corrects under-reporting of aggregated data reliability.
In autonomous vehicle (AV) control, allowing mistakes can be quite dangerous and costly in the real world. For this reason we investigate methods of training an AV without allowing the agent to explore and instead having a human explorer collect the data. Supervised learning has been explored for AV control, but it enc…
GICDM corrects hubness in embedding spaces for better generative model evaluation.
We proposed a probabilistic approach to joint modeling of participants' reliability and humans' regularity in crowdsourced affective studies. Reliability measures how likely a subject will respond to a question seriously; and regularity measures how often a human will agree with other seriously-entered responses coming…
XAI methods fail to explain ML models reliably.
Method curates cost-effective, high-quality datasets using AI models.
The paper proposes a method to evaluate superhuman models by checking for logical inconsistencies.
Rapid intensification (RI) of tropical cyclones often causes major destruction to human civilization due to short response time. It is an important yet challenging task to accurately predict this kind of extreme weather event in advance. Traditionally, meteorologists tackle the task with human-driven feature extraction…