Neural story generation gains common sense through targeted training.
problem Lack of common sense reasoning in neural-generated stories.
method Multi-task learning with auxiliary datasets for common sense grounding.
result Improved common sense reasoning and state-of-the-art perplexity.
Study evaluates progress in common-sense reasoning tasks.
problem Assessing genuine progress in common-sense reasoning systems.
method Case studies of WSC and SWAG, protocol design to clarify results.
result Previous experimental designs had flaws, need for new protocols.
The problem of replicating the flexibility of human common-sense reasoning has captured the imagination of computer scientists since the early days of Alan Turing's foundational work on computation and the philosophy of artificial intelligence. In the intervening years, the idea of cognition as computation has emerged …
A new method learns object hierarchies from images to reason about physical interactions.
problem Learning about the interactions of complex objects and their dynamics.
method Unsupervised learning of object hierarchies from raw visual images.
result Improves over a strong baseline at modeling synthetic and real-world videos.
ZSL-KG learns class representations from common sense knowledge graphs.
problem Predicting classes without labeled examples using semantic class representations.
method TrGCN, a novel transformer graph convolutional network, embeds nodes from common sense knowledge graphs in a vector space.
result ZSL-KG improves over existing methods on five out of six zero-shot benchmark datasets.
A novel fact-checking method using debate dynamics on knowledge graphs.
problem Fact-checking on knowledge graphs with user comprehension and interactive reasoning.
method Reinforcement learning agents debate on paths in the graph to classify facts as true or false.
result Interactive reasoning and user understanding of AI decisions on knowledge graphs.
Deep Reinforcement Learning (deep RL) has made several breakthroughs in recent years in applications ranging from complex control tasks in unmanned vehicles to game playing. Despite their success, deep RL still lacks several important capacities of human intelligence, such as transfer learning, abstraction and interpre…
Spatial understanding is a fundamental problem with wide-reaching real-world applications. The representation of spatial knowledge is often modeled with spatial templates, i.e., regions of acceptability of two objects under an explicit spatial relationship (e.g., "on", "below", etc.). In contrast with prior work that r…
We will compare three types of prices, namely, rational (hedging) prices, geometric (growth rate) prices, and martingale (measure) prices. We will show that rational prices in the complete market theory are sometimes contrary to common sense. In the continuous-time case, we insist that the market model should differ be…
PredictaBoard benchmarks LLM score predictors to assess their ability to anticipate errors.
problem Inconsistent performance of LLMs in common sense reasoning tasks.
method Collaborative benchmarking framework evaluating pairs of LLMs and assessors using rejection rate at different tolerance errors.
result Highlights the need to evaluate predictability alongside performance for safer AI systems.
It is a common belief that the behavior of shareholders depends upon the direction of price fluctuations: if prices increase they buy, if prices decrease they sell. That belief, however, is more based on ``common sense'' than on facts. In this paper we present evidence for a specific class of shareholders which shows t…
New method calibrates LLMs for safety-critical tasks with scalable Bayesian inference.
problem Overconfidence in LLMs after fine-tuning for specific tasks.
method Orthogonalized Low-Rank Adapters (PoLAR) with variational Bayesian inference.
result Scalable and well-calibrated uncertainty estimation for LLMs.
Online news recommender systems aim to address the information explosion of news and make personalized recommendation for users. In general, news language is highly condensed, full of knowledge entities and common sense. However, existing methods are unaware of such external knowledge and cannot fully discover latent k…
Induction of common sense knowledge about prototypical sequences of events has recently received much attention. Instead of inducing this knowledge in the form of graphs, as in much of the previous work, in our method, distributed representations of event realizations are computed based on distributed representations o…
Recent research in psycholinguistics has provided increasing evidence that humans predict upcoming content. Prediction also affects perception and might be a key to robustness in human language processing. In this paper, we investigate the factors that affect human prediction by building a computational model that can …
BIG-bench benchmarks language models, revealing their strengths and weaknesses.
problem Understanding and quantifying the capabilities of large language models.
method Developed the BIG-bench benchmark with 204 tasks from various domains, evaluated across multiple model sizes and types.
result Model performance improves with scale but remains poor, with breakthrough behaviors often involving multiple steps.
Our goal here is to discuss the pricing problem of European and American options in discrete time using elementary calculus so as to be an easy reference for first year undergraduate students. Using the binomial model we compute the fair price of European and American options. We explain the notion of Arbitrage and the…
Geometric deep learning detects fake news on social media.
problem Detecting fake news on social media due to lack of context understanding.
method Propagation-based geometric deep learning model.
result Highly accurate fake news detection (92.7% ROC AUC).
The paper addresses optimal control in modern tontines with bequest preferences, showing a linear investment strategy.
problem Optimal controls and decreasing allocation in modern tontines with bequest preferences.
method Dual approach to solve optimal control problems with power utilities, modeling bequest preferences.
result Investment strategy almost linearly adjusts from 0% to 100% over time.
Unified approach for learning with weak labels across various tasks.
problem Learning with noisy or incomplete labels in diverse machine learning settings.
method Implicit posterior models for joint label inference.
result Unified training objective for various machine learning tasks.
We propose a method to clean covariance matrices of nonstationary systems by using time-independent eigenvalues.
problem Noise in covariance matrices of nonstationary systems with time-independent eigenvalues.
method Data-driven approach to use independent eigenvalues encoding long-term influence of future on present.
result Our method outperforms optimal stationary methods for filtering covariance matrix and its inverse.
Automatically finds effective security strategies through reinforcement learning and self-play.
problem Finding effective security strategies for intrusion prevention.
method Modeling interaction as a Markov game, evolving attack and defense strategies through reinforcement learning and self-play.
result Effective security strategies emerge from self-play, reflecting common-sense knowledge.
The paper compares theoretical and empirical performance of imputation methods for missing data.
problem Missing data in real-world datasets.
method Contrast of theoretical and empirical imputation methods for prediction.
result Mean-imputation is asymptotically optimal for prediction, while mode-imputation is sub-optimal.
LaTRO optimizes latent reasoning in LLMs without external reward.
problem Training LLMs to perform complex reasoning tasks.
method Formulates reasoning as latent distribution sampling and optimizes via variational approaches.
result LLMs improve reasoning and evaluation quality through self-improvement.
Auto-CEI improves LLM reasoning by balancing assertiveness and conservativeness.
problem Hallucinations and laziness in LLM reasoning tasks.
method Expert Iteration explores reasoning trajectories, guiding incorrect paths back on track and promoting appropriate 'I don't know' responses.
result Auto-CEI achieves superior alignment in logical reasoning, mathematics, and planning tasks.
A framework isolates VQA reasoning from perception for better model evaluation.
problem Improper separation of visual perception and reasoning in VQA models.
method Introducing a framework and a top-down calibration technique to decouple reasoning from perception.
result Improved evaluation of VQA models by separating reasoning from perception.
Framework uses human annotations to make models robust to spurious correlations.
problem Machine learning models fail when unmeasured variables change test distributions.
method Human annotations to augment training examples, UV-DRO objective for robustness.
result Improvements of 5-10% on digit recognition task and 1.5-5% on NYPD Police Stops analysis.
New corpus improves coreference resolution by removing gender and number cues.
problem Challenges in resolving ambiguous pronoun references.
method Developed a new annotated corpus, introduced antecedent switching technique.
result Models perform poorly on ambiguous pronoun references, but antecedent switching improves performance.
A new method for math reasoning that allows for iterative correction.
problem Standard reasoning models commit to each token and cannot recover from early errors.
method Generative framework with latent thought vectors for iterative self-correction.
result 30 rethinking iterations surpass baselines with 15 times more parameters.
Cognitive KR model learns new facts from few examples.
problem Inferring new facts from small KG datasets.
method CogKR combines summary and reasoning modules with cognitive science principles.
result Significantly outperforms previous models on one-shot KG reasoning.
Collaborative filtering analyzes user preferences for items (e.g., books, movies, restaurants, academic papers) by exploiting the similarity patterns across users. In implicit feedback settings, all the items, including the ones that a user did not consume, are taken into consideration. But this assumption does not acc…
Transformers learn multi-step reasoning through gradient descent.
problem Understanding how transformers solve symbolic multi-step reasoning tasks.
method Theoretical analysis of gradient descent dynamics and multi-phase training.
result Trained one-layer transformers can solve both backward and forward reasoning tasks with generalization guarantees.
Transformers with CoT don't enhance reasoning power across all tasks.
problem Does CoT enhance the reasoning power of transformers?
method Examined the memorization capabilities of fixed-precision transformers with and without CoT.
result Transformers with CoT cannot memorize all reasoning tasks, leading to a negative answer.
Early stopping methods reduce unnecessary reasoning steps in LLMs by monitoring uncertainty signals.
problem LLMs sometimes generate unnecessary reasoning steps, especially under uncertainty.
method Statistically principled early stopping methods that monitor uncertainty signals during generation.
result Uncertainty-aware early stopping improves efficiency and reliability in LLM reasoning, especially in math reasoning.
Achieving artificial visual reasoning - the ability to answer image-related questions which require a multi-step, high-level process - is an important step towards artificial general intelligence. This multi-modal task requires learning a question-dependent, structured reasoning process over images from language. Stand…
DAFT models attention as a dynamical system to make neural networks more interpretable.
problem Uninterpretable features learned by neural networks without human priors.
method DAFT models attention as a continuous dynamical system using neural ODEs.
result DAFT reduces the number of reasoning steps while maintaining similar performance.
FinTradeBench benchmarks LLMs for financial reasoning combining company fundamentals and market signals.
problem Challenges in evaluating financial reasoning models for LLMs.
method Developed a benchmark integrating company fundamentals and trading signals, using a calibration-then-scaling framework.
result Clear performance gap between LLMs, retrieval improves reasoning over textual fundamentals but not trading signals.
MultiVerse uses importance sampling for efficient causal reasoning in probabilistic programming.
problem Efficient causal reasoning in probabilistic models, especially counterfactual inference.
method Native implementation of importance sampling in probabilistic programming, optimizing inference through query structure.
result Significant optimisation of inference process through careful design choices and query structure consideration.
What return should you expect when you take on a given amount of risk? How should that return depend upon other people's behavior? What principles can you use to answer these questions? In this paper, we approach these topics by exploring the consequences of two simple hypotheses about risk. The first is a common-sense…
Hybrid framework injects TSLM insights into GRLM for robust time-series reasoning.
problem Lack of domain-specific knowledge in large language models for time-series reasoning.
method Hybrid knowledge-injection framework combining RLVR for efficient knowledge transfer.
result Consistently outperforms existing models by 7.9%-26.1% on multivariate time-series benchmarks.
Neural Logic Reasoning integrates deep learning and symbolic logic for better prediction tasks.
problem Lack of cognitive reasoning in deep neural networks limits their ability to solve complex prediction tasks.
method Proposes Logic-Integrated Neural Network (LINN) that learns logical operations and conducts propositional logical reasoning.
result LINN significantly outperforms state-of-the-art recommendation models in Top-K recommendation.
Forward-prediction models enhance physical reasoning, but only for specific tasks.
problem Improving physical reasoning in complex tasks involving many objects.
method Incorporated forward-prediction models into simple physical-reasoning agents and evaluated their performance on the PHYRE benchmark.
result Forward-prediction models improve physical-reasoning performance, especially on complex tasks, but generalization to new task templates is challenging.
The paper tackles efficient exploration in MDPs to learn accurate models.
problem Efficient exploration in MDPs to learn accurate models.
method Formalizes the problem, introduces an algorithm for ε-accurate model estimation, and proposes a heuristic-based algorithm. result Heuristic-based algorithm outperforms original algorithm in small sample regime.
Meta-reinforcement learning enables causal reasoning in complex environments.
problem Discovering causal structure in complex environments.
method Training a recurrent network with model-free reinforcement learning to solve problems with causal structure.
result The trained agent can perform causal reasoning in novel situations, select informative interventions, draw causal inferences, and make counterfactual predictions.
RACER optimizes LLM-as-judge accuracy with dynamic reasoning selection.
problem Balancing reasoning accuracy with computational cost in LLM-as-judge settings.
method Formulates routing as a constrained distributionally robust optimization problem, accounting for distribution shift via KL-divergence uncertainty set.
result RACER achieves superior accuracy-cost trade-offs under distribution shift.
VTA combines verbal and latent reasoning for accurate stock time-series forecasts.
problem Challenges in combining textual analysis with time-series data for financial forecasting.
method Converts stock price data into textual annotations, optimizes reasoning trace using inverse MSE, conditions time-series model outputs on reasoning attributes.
result VTA achieves state-of-the-art forecasting accuracy and interpretable reasoning traces.
Neural networks struggle with reasoning tasks that require specialized structures.
problem Understanding why and when neural network structures generalize better.
method Developing a framework to characterize algorithmic alignment with reasoning tasks.
result Neural networks align better with dynamic programming (DP) for certain reasoning tasks.
Fractured Sampling improves LLM reasoning efficiency by truncating CoT trajectories.
problem Efficiently scaling reasoning in large language models with limited tokens.
method Integrating truncated Chain-of-Thought (CoT) with Fractured Sampling across multiple dimensions.
result Fractured Sampling achieves superior accuracy-cost trade-offs compared to full CoT.