Hotel booking chatbot handles daily tens of thousands of searches.
problem Improving hotel booking through conversational AI.
method Frame-based dialogue management system with ML models for intent, entity recognition, and info retrieval.
result Chatbot deployed on a commercial scale, handling hotel searches daily.
Chatbot uses BERT to handle financial investment questions, improving accuracy and decision-making.
problem Improving accuracy and decision-making in financial investment customer service.
method Deep Bidirectional Transformer (BERT) model, uncertainty measure comparison, mixed-integer programming, automatic spelling correction.
result Chatbot can recognize 381 intents and decide when to escalate questions.
Paper improves emotion expression in AI chatbots.
problem AI chatbots generate responses that lack emotion.
method Developed neural models to express specific emotions in generated responses.
result An encoder-decoder model with multiple attention layers performs best in expressing required emotion.
Study uses chatbot to understand users' needs for ML model explanations.
problem Lack of understanding of user needs for model explanations.
method Developed a conversational system (dr_ant) to collect user questions about a machine learning model trained on Titanic data.
result Collected a corpus of 1000+ dialogues to identify common user questions.
AI helps simplify complex ship finance processes.
problem Complexity in ship finance due to data and regulatory requirements.
method Integrates large language models for document comprehension, information extraction, and workflow automation.
result AI-assisted systems can support maritime finance professionals in managing complex information and reporting requirements.
Self-feeding chatbots learn from user feedback to improve performance.
problem Lack of training data after deployment of dialogue agents.
method Extracts new training examples from user responses and uses feedback to improve dialogue abilities.
result Self-feeding chatbots significantly improve performance on PersonaChat dataset.
Topic-aware chatbot learns from NMF topic vectors.
problem Improving chatbot relevance based on user topics.
method Combines RNN with NMF for topic learning and attention.
result Chatbot provides more relevant answers based on topic.
Meena is a chatbot trained on social media data, achieving human-like conversation quality.
problem Creating a chatbot that can have human-like conversations in an open-domain setting.
method End-to-end training of a 2.6B parameter neural network on social media data, using perplexity and a human evaluation metric (SSA) to assess quality.
result Meena achieves a high human-like conversation quality (79% SSA) when compared to existing chatbots.
Statistical framework improves LLM chatbot ranking.
problem Improving evaluation of LLM-based chatbots through pairwise comparisons.
method Factored tie model, covariance modeling, and parameter constraints.
result Substantial improvements in modeling pairwise comparison data.
Self-attentional models improve chatbot efficiency and performance.
problem Training efficient task-oriented dialogue generation systems.
method Applied self-attentional models to three datasets for chatbot training.
result Self-attentional models outperform recurrence-based models in efficiency and performance.
Information extraction and user intention identification are central topics in modern query understanding and recommendation systems. In this paper, we propose DeepProbe, a generic information-directed interaction framework which is built around an attention-based sequence to sequence (seq2seq) recurrent neural network…
We study reinforcement learning of chatbots with recurrent neural network architectures when the rewards are noisy and expensive to obtain. For instance, a chatbot used in automated customer service support can be scored by quality assurance agents, but this process can be expensive, time consuming and noisy. Previous …
Dropping a tiny fraction of preferences can significantly alter the rankings of top LLMs.
problem Robustness of LLM ranking systems to small changes in preference data.
method A computational method based on the Bradley-Terry model to evaluate robustness.
result Top LLM rankings can be highly sensitive to the removal of a small fraction of preferences.
We present MILABOT: a deep reinforcement learning chatbot developed by the Montreal Institute for Learning Algorithms (MILA) for the Amazon Alexa Prize competition. MILABOT is capable of conversing with humans on popular small talk topics through both speech and text. The system consists of an ensemble of natural langu…
We present MILABOT: a deep reinforcement learning chatbot developed by the Montreal Institute for Learning Algorithms (MILA) for the Amazon Alexa Prize competition. MILABOT is capable of conversing with humans on popular small talk topics through both speech and text. The system consists of an ensemble of natural langu…
This paper tackles interpretability of LLMs in finance.
problem Complexity and lack of transparency of LLMs in finance.
method Mechanistic interpretability to understand LLM behavior.
result Demonstrates practical relevance of mechanistic interpretability in financial use cases.
Unified framework to bridge human and LLM judgments.
problem Systematic discrepancies between human and LLM evaluations.
method Latent human preference score and linear transformations of covariates.
result Higher agreement with human ratings and exposure of systematic gaps.
This study evaluates the performances of an LSTM network for detecting and extracting the intent and content of com- mands for a financial chatbot. It presents two techniques, sequence to sequence learning and Multi-Task Learning, which might improve on the previous task.
Quriosity analyzes curiosity-driven questions from diverse sources.
problem Understanding and analyzing human curiosity-driven questions.
method Collection of 13.5K questions from various sources, development of an iterative prompt improvement framework.
result 42% of questions are causal, revealing unique linguistic properties.
New DL approach reveals feature construction in dense samples.
problem Understanding deep learning's effectiveness across diverse applications.
method High-density sample task with 5 unique tokens, 500 exemplars per token.
result Emergence of category structure and feature detectors observed.
In sequence generation task, many works use policy gradient for model optimization to tackle the intractable backpropagation issue when maximizing the non-differentiable evaluation metrics or fooling the discriminator in adversarial learning. In this paper, we replace policy gradient with proximal policy optimization (…
Low-rank framework for task-specific LLM ranking from sparse comparisons.
problem Challenges in reliable task-specific ranking of LLMs under sparse, imbalanced comparisons.
method Low-rank modeling of task-by-model ability matrix, max-norm accurate estimator, task-wise top-K recovery guarantees, uncertainty quantification framework.
result Improves sample efficiency and produces tighter, better-calibrated ranking certificates.
AI stocks hedge against AI singularity's economic impact.
problem AI singularity's displacement of consumption.
method Developed an asset pricing model with incomplete markets.
result AI stocks command a premium due to market incompleteness.
Study analyzes AI's impact on firms, markets, and workers using large language model data.
problem Understanding AI's effect on firms, markets, and workers.
method Used 380 trillion tokens from 400+ large language models to analyze AI's impact.
result Firms with higher AI exposure earn higher returns, creating an AI premium.
New AI stock indices classify firms' AI engagement using 10-K filings.
problem Opaque AI selection criteria in existing ETFs.
method NLP analysis of 10-K filings to classify AI stocks.
result Companies with higher AI engagement have greater positive returns.
AI updates can disrupt human-AI teams; new objective improves compatibility.
problem AI updates can disrupt human-AI team performance.
method Introduce compatibility concept, propose re-training objective.
result Current machine learning algorithms do not produce compatible updates.
The paper analyzes risk spillovers between AI ETFs, AI tokens, and green markets.
problem Risk spillovers among AI ETFs, AI tokens, and green markets.
method R2 decomposition method
result AI ETFs and clean energy act as risk transmitters, while AI tokens and green assets act as receivers.
This review covers AI in finance, challenges, techniques, and opportunities.
problem Challenges and opportunities in AI applications in finance.
method Comprehensive categorization and overview of AI research in finance over decades.
result A dense roadmap of AI challenges, techniques, and opportunities in finance.
Explainable AI improves human decision accuracy but does not enhance it significantly.
problem Improving human decision-making through explainable AI.
method Comparing human decision accuracy with and without AI predictions, including or excluding explanations.
result Providing AI predictions improves human decision accuracy, but explanations do not significantly enhance it.
The paper explores AI in finance, focusing on XAI's role in enhancing interpretability and trust.
problem The need for AI in finance and the importance of XAI for better decision-making.
method Tracing AI's evolution in finance, highlighting XAI's role, and demonstrating through simulations.
result XAI enhances trust in AI systems, leading to more responsible decision-making.
Paper defines AI-specific loss reconstruction problem and introduces CER framework.
problem Reconstructing AI-generated losses, especially in agentic systems.
method CER framework: C (control boundary), E (evidence reconstruction), R (insurance response).
result Defines AI-specific reconstruction problem and operationalizes it.
Experiment shows cognitive biases impact human-AI collaboration, highlighting the need for diverse evaluator samples.
problem Cognitive biases affect human-AI collaboration, leading to suboptimal outcomes.
method Randomized experiment with 2,784 participants, manipulating AI suggestion quality, task burden, and financial incentives.
result Individual attitudes toward AI are the strongest predictor of performance, influencing accuracy and overcorrection.
AI pipeline simplifies AI deployment on embedded devices.
problem Training and deployment of custom AI solutions on embedded devices requires expertise and integration barriers.
method Modular AI pipeline integrating data, algorithms, and deployment tools.
result LPDNN consistently outperforms other deployment frameworks on embedded platforms.
AI enhances financial forecasting with challenges in regulation and privacy.
problem Challenges in integrating AI with financial services and regulations.
method Integration of AI technologies like deep learning and reinforcement learning.
result AI improves financial forecasting but faces regulatory and privacy issues.
AI+MPS workshop aims to strengthen AI's role in science.
problem AI's potential to enhance scientific discovery and education.
method Proposes activities and strategic priorities to strengthen AI-MPS link.
result AI and science are becoming increasingly intertwined.
Improved AI patent classifier measures U.S. and China's AI patenting.
problem Measuring AI patents with high precision and generalization.
method Fine-tuning PatentSBERTa on manually labeled data from USPTO's AI Patent Dataset.
result Rapid growth in AI patenting in both countries, but different organizational patterns.
This paper evaluates AI ethics guidelines and their implementation.
problem Lack of comprehensive and effective AI ethics guidelines.
method Comprehensive evaluation and comparison of released ethics guidelines.
result Identifies overlaps and omissions in AI ethics guidelines.
FST.ai 2.0 improves Taekwondo decision-making with AI, reducing review time and increasing trust.
problem Fair, transparent, and explainable decision-making in Taekwondo.
method Pose-based action recognition, epistemic uncertainty modeling, interactive dashboards.
result 85% reduction in decision review time, 93% referee trust in AI-assisted decisions.
Article evaluates AI security threats and proposes multiple measures.
problem Threats to AI integrity and security.
method Literature review, analysis of AI supply chain, discussion of mitigations.
result Multiple protective measures are necessary for AI security.
Paper tackles AI risks by customizing metrics and models.
problem AI risks are multidimensional and immaturely managed.
method Decomposes AI risks into data protection, fairness, etc., and develops metrics and models.
result Customized metrics and models reduce AI risk uncertainty.
The DoD needs a robust process to evaluate AI/ML model performance and robustness.
problem AI/ML models are brittle and nonrobust, posing risks in national security.
method Reviews AI/ML development process and best practices for evaluation.
result Recommendations for DoD evaluators to ensure robust AI/ML capabilities.
Adversarial policies beat superhuman Go AI systems.
problem Vulnerability of superhuman AI systems to adversarial attacks.
method Training adversarial policies to trick KataGo into making blunders.
result Adversarial policies achieve >97% win rate against KataGo at superhuman settings.
AI threatens financial stability through misuse and stealth adoption.
problem Misuse and stealth adoption of AI in financial regulations.
method Analysis of AI's potential risks and criteria for AI suitability.
result AI will likely become widely used by stealth, affecting high-level financial functions.
Qlib aims to integrate AI into quantitative investment.
problem Challenges in applying AI to quantitative investment.
method Design and develop Qlib to accommodate AI-driven workflow.
result Qlib realizes the potential of AI technologies in quantitative investment.
Paper withdrawn; AI game behavior needs diversity.
problem Creating varied human-like playing styles in games.
method Evolutionary multi-objective deep reinforcement learning.
result Generated diverse AI behaviors for games.
Open AI models affect bond yields differently than closed ones.
problem Understanding how market reactions to AI releases impact bond yields.
method Analyzed US bond yields before and after the release of open and closed AI models.
result Open AI models shift bond yields in the opposite direction of closed models.
Do-AIQ framework evaluates AI algorithms' quality using DOE.
problem Quality evaluation of AI mislabel detection algorithms.
method Design-of-experiment approach with high-dimensional constraint space design and surrogate modeling.
result Established framework for evaluating AI algorithm quality robustly.
AI enhances ESG practices in finance, but requires careful consideration.
problem Regulatory pressures and stakeholder awareness drive ESG adoption.
method Industrial survey categorizing AI applications in ESG.
result AI improves analytical capabilities, risk assessment, and customer engagement.