The abstract discusses how humans use visualizations in machine learning.
problem The reliance on human involvement in AI systems and analytics.
method Review of seven steps in the ML process and different visualization techniques.
result Different visualizations are used at various stages of the ML process.
Framework for responsible LLM deployment with human involvement and decentralized technologies.
problem Challenges in deploying LLMs for high-stakes decisions, including data security and accountability.
method Interactive human involvement through multiple iterations, decentralized technologies, and automated auditing.
result Enhanced security and accountability in LLM deployment for financial decisions.
A new approach to fine-tuning LLMs with human feedback.
problem Inability of current reward models to fully represent human preferences.
method Introducing NLHF, a new pipeline for LLM fine-tuning using pairwise human feedback.
result NLHF produces a sequence of policies converging to the regularized Nash equilibrium.
Parenting algorithm improves AI safety by learning from human input.
problem Safety concerns in reinforcement learning, especially reward hacking and unsafe exploration.
method Inspired by parenting, a precise framework for learning from human input.
result Parenting algorithm solves safety problems in AI Safety gridworlds.
Optimizes model interpretability by involving humans in the loop.
problem Optimizing for interpretability while maintaining model accuracy.
method Develops an algorithm that minimizes user studies to find interpretable models.
result Human subjects results show varying preferences for interpretability proxies across different datasets.
A general theory of innovation and progress in human society is outlined, based on the combat between two opposite forces (conservatism/inertia and speculative herding "bubble" behavior). We contend that human affairs are characterized by ubiquitous ``bubbles'', which involve huge risks which would not otherwise be tak…
Unified framework for human-like decision making in various sequential tasks.
problem Real-life decision-making involves diverse strategies leading to similar outcomes.
method Two-stream reward processing mechanism for flexible and unified models.
result Framework unified MAB, CB, and RL with comparable performance.
Study shows how machine learning can improve human performance in deception detection.
problem Improving human performance in critical tasks involving ethical and legal concerns.
method Investigated how machine learning models and their predictions can assist humans in deception detection tasks.
result Explanations and predicted labels from machine learning models can improve human performance in deception detection.
Deep RL mimics human driving for collision avoidance in self-driving cars.
problem Developing human-like driving policies for autonomous vehicles in mixed traffic environments.
method Model-free, deep reinforcement learning approach using a combination of rule-based and expert-driven data.
result Demonstrated human-like driving policies through Gaussian process modeling of track position and speed distributions.
Study finds visual explanations do not significantly improve human accuracy or trust in model predictions.
problem Measuring the impact of visual explanations on human accuracy and trust in model predictions.
method Randomized controlled trial with image-based age prediction task, varying levels of explanation quality.
result Visual explanations do not significantly alter human accuracy or trust in the model.
Enhances AI models with human feedback for noisy data.
problem Improving AI model alignment with human feedback in noisy environments.
method Two-stage SL+LHF framework connecting machine learning with human feedback.
result The LNCA ratio identifies conditions for SL+LHF superiority over pure SL.
Unified LP framework for offline reward learning from human demonstrations and feedback.
problem Reward learning from human demonstrations and feedback with robustness and sample efficiency.
method A novel linear programming framework for offline reward learning.
result Unified LP framework achieves better performance compared to MLE.
IDS improves RLHF by smoothing reward data, enhancing model performance.
problem Reward model performance degrades and overoptimization hinders true objective.
method Iterative Data Smoothing (IDS) updates model and data labels during each epoch.
result IDS outperforms traditional methods in RLHF.
Partner-aware algorithms improve AI collaboration in multi-agent settings.
problem Improving AI cooperation in teams with shared rewards.
method Proposed Partner-Aware strategy extending Upper Confidence Bound for decentralized MAB.
result Achieves logarithmic regret in collaborative decision-making.
Paper tackles RLHF with DCPPO method, proving near-optimal suboptimality.
problem Challenges in offline RLHF with limited human feedback and bounded rationality.
method DCPPO method involving three stages: MLE, reward function recovery, and pessimistic value iteration.
result DCPPO's suboptimality almost matches classical pessimistic offline RL in terms of distribution shift and dimension.
An approach for learning ancestral causal relationships in high dimensions, validated on human genome-wide data.
problem Learning ancestral causal relationships in high-dimensional biological data.
method Supervised learning approach with discrete indicators treated as labels, scalable to large problems.
result The approach is highly effective and scalable to the human genome-wide setting, robust to perturbations of input information.
AI systems learn complex goals via debate with human judges.
problem Complex human goals and preferences for AI systems.
method Training agents via self-play on a debate game, where humans judge which agent gives more true, useful information.
result Boosted classifier accuracy from 48.2% to 85.2% given 4 pixels, and from 59.4% to 88.9% given 6 pixels.
New clustering method for POIs from spatio-temporal data with temporal constraints.
problem Lack of temporal constraints in existing clustering methods for POIs.
method POI clustering with temporal constraints (PC-TC).
result PC-TC outperforms existing methods in next place prediction.
TRIBE model uses LLMs to simulate human trading behavior in bond markets.
problem Complexities in decentralized bond market transactions.
method Agent-based model augmented with LLMs to simulate human-like decision-making.
result Slight trade aversion in LLMs can lead to complete market collapse.
Study extends cognitive modeling to natural images, revealing the importance of image representation.
problem Extending cognitive modeling to natural images and understanding human categorization.
method Conducted a large-scale study with over 500,000 human judgments. Used deep and shallow machine learning methods to represent images. Applied psychological models of categorization to natural images.
result Simple models with abstract prototypes outperform complex exemplar accounts when using expressive, data-driven image representations.
In this paper the theory of semi-bounded rationality is proposed as an extension of the theory of bounded rationality. In particular, it is proposed that a decision making process involves two components and these are the correlation machine, which estimates missing values, and the causal machine, which relates the cau…
How do we assign value to economic transactions? To answer this question, we must consider whether the value of objects is inherent, is a product of social interaction, or involves other mechanisms. Economic theory predicts that there is an optimal price for any market transaction, and can be observed during auctions o…
Improves prediction accuracy in document classification by measuring uncertainty.
problem Ensuring limited human resources focus on uncertain predictions in text classification.
method Proposes a neural-network-based model using dropout-entropy for uncertainty measurement and metric learning on feature representations.
result Significant improvement in overall prediction accuracy, from 0.78 to 0.92, when 30% of most uncertain predictions are handed over to human experts.
Deep RL agent outperforms humans in constructing towers using relational reasoning.
problem Current deep learning systems struggle with constructing and modifying complex systems.
method Introduced a deep reinforcement learning agent with object- and relation-centric scene and policy representations.
result Structured representations allow the agent to outperform humans and naive approaches.
A grand challenge in machine learning is the development of computational algorithms that match or outperform humans in perceptual inference tasks that are complicated by nuisance variation. For instance, visual object recognition involves the unknown object position, orientation, and scale in object recognition while …
UCFE benchmarks LLMs in financial tasks with human feedback.
problem Evaluating LLMs' financial task performance and user satisfaction.
method Hybrid approach combining human expert evaluations and dynamic interactions.
result Significant alignment between benchmark scores and human preferences (Pearson correlation coefficient of 0.78).
Study human-machine interaction with private info using offline RL.
problem Confounding bias and distributional mismatch in offline RL for human-guided interaction.
method Developed a novel identification result and OPE method to address confounding bias, and used pessimism to tackle distributional mismatch.
result Policy pair converges to optimal one at satisfactory rate under mild assumptions.
The paper predicts human-like driving behavior of other vehicles for safer AVs.
problem Safe and efficient interaction of AVs with other vehicles.
method Hierarchical inverse reinforcement learning considering both discrete and continuous decisions.
result The proposed approach accurately predicts both discrete and continuous driving behaviors.
Automates feature engineering for predictive models using reinforcement learning.
problem Lack of a well-defined basis for effective feature engineering.
method Performance-driven exploration of a transformation graph using reinforcement learning.
result Automated feature engineering reduces human intervention and costs.
Machine learning uses crowdworkers; determining their status as human subjects is tricky.
problem Determining the appropriate status of ML crowdworkers as human subjects.
method Investigation of natural language processing studies to expose challenges and propose solutions.
result Potential loophole in the U.S. Common Rule for ML research oversight.
New modeling approach for self-organizing complex systems.
problem Identifying general principles of self-organizing complex systems.
method Self-deploying system structure and activities for goal-driven agents.
result Self-organization emerges from rational activity algorithm based on goals dependency network.
This paper improves sample efficiency for off-policy evaluation with preference data.
problem Improving sample efficiency for off-policy evaluation with preference data.
method Using a deep neural network to learn the value function and leveraging manifold structure.
result Established a provably efficient guarantee for off-policy evaluation with RLHF.
This paper formalizes reward training in RLHF and provides theoretical guarantees.
problem Costly human feedback collection in RLHF.
method Linear contextual dueling bandit method for regret minimization.
result Derives bounds on simple regret for offline reward training.
Visual observations of dynamic phenomena, such as human actions, are often represented as sequences of smoothly-varying features . In cases where the feature spaces can be structured as Riemannian manifolds, the corresponding representations become trajectories on manifolds. Analysis of these trajectories is challengin…
Paper proposes Unification Networks to learn invariants from examples.
problem Learning to recognize common underlying principles across examples.
method End-to-end differentiable neural network approach with soft unification.
result Learning invariants improves performance on various datasets.
AI benchmarks evaluate football team performance using generative models.
problem Evaluating human performance in complex interactive tasks is error-prone and unreliable.
method Trained Conditional VRNN Model on player and ball tracking data to imitate and predict team interactions.
result Trained model as a useful benchmark for evaluating team performance in football.
Human computation or crowdsourcing involves joint inference of the ground-truth-answers and the worker-abilities by optimizing an objective function, for instance, by maximizing the data likelihood based on an assumed underlying model. A variety of methods have been proposed in the literature to address this inference …
Teaching machines to accomplish tasks by conversing naturally with humans is challenging. Currently, developing task-oriented dialogue systems requires creating multiple components and typically this involves either a large amount of handcrafting, or acquiring costly labelled datasets to solve a statistical learning pr…
Traditional disease surveillance can be augmented with a wide variety of real-time sources such as, news and social media. However, these sources are in general unstructured and, construction of surveillance tools such as taxonomical correlations and trace mapping involves considerable human supervision. In this paper,…
Interview study reveals considerations for designing semi-automated bias detection tools.
problem Detecting and mitigating machine learning biases.
method Interview study with 11 machine learning practitioners.
result Four considerations identified for tool design.
Enhances BO with expert preferences about abstract properties.
problem Lack of expert knowledge in BO for black-box experimental design.
method Human-AI collaboration to incorporate expert preferences into surrogate modeling.
result Superior performance compared to baselines in synthetic and real-world datasets.
Study uses machine learning to predict rain in Australia.
problem Challenging task of predicting rainfall with uncertain outcomes.
method Machine learning techniques, including modeling inputs, methods, and pre-processing.
result Comparison of various machine learning techniques' reliability in predicting rainfall.
Paper proposes using NLPD for cGANs to improve image realism and segmentation.
problem Lack of perceptual evaluation in cGANs.
method Used NLPD to measure perceptual similarity, comparing with L1 distance.
result NLPD improves realism and segmentation accuracy in generated images.
Collective learning leverages human collaboration for semi-supervised learning.
problem Distributed semi-supervised learning with heterogeneous systems.
method Two phases: self-training and collective training with proxy-labels.
result Collective learning improves performance over individual learning.
Iterated Amplification uses subproblem solutions to build training signals for complex tasks.
problem Learning complex tasks when humans can't directly evaluate performance.
method Progressively builds training signal by combining solutions to easier subproblems.
result Efficiently learns complex behaviors in algorithmic environments.
Mitigates biases in reward models using variational inference.
problem Spurious correlations in reward models that align large language models with human preferences.
method Formulates data-generating process, identifies non-spurious latent variables, and uses variational inference to recover them.
result Effective mitigation of spurious correlation issues, yielding more robust reward models.
Study shows trust and trustworthiness emerge through reinforcement learning.
problem Trust and trustworthiness are universal but not predicted by traditional economic models.
method Used Q-learning algorithm to simulate trust and trustworthiness dynamics in a trust game.
result High levels of trust and trustworthiness emerge when individuals consider both past and future experiences.
Cognitive model discovery framework for ill-structured domains.
problem Discover accurate cognitive models without student data or human knowledge.
method Cognitive Representation Learner (CogRL) framework for ill-structured domains.
result CogRL representations can accurately discover cognitive models in ill-structured domains.