Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

22446587 · May 202619922001200920172026
48 results for causal thinking

CausalGame benchmarks LLM agents' causal thinking in games.

problem Evaluating causal thinking in AI Scientists with LLMs.
method Interactive games with 14 scenarios incorporating selection bias, measurement error, and hidden confounders.
result None of the 30 LLM agents demonstrated reliable causal thinking, with the best model achieving only 68.0% survival.

New method uses causal thinking to make AI fairer decisions.

problem Designing fair machine learning models that treat equal individuals equally and unequals unequally.
method Rank-preserving interventional distributions and warping method.
result Warping method effectively identifies discriminated individuals and mitigates unfairness.

Recent progress in artificial intelligence (AI) has renewed interest in building systems that learn and think like people. Many advances have come from using deep neural networks trained end-to-end in tasks such as object recognition, video games, and board games, achieving performance that equals or even beats humans …

2016-04-01abs ↗pdf ↗

Human inertial thinking schemes can be formed through learning, which are then applied to quickly solve similar problems later. However, when problems are significantly different, inertial thinking generally presents the solutions that are definitely imperfect. In such cases, people will apply creative thinking, such a…

2018-03-01abs ↗pdf ↗

The paper tackles spurious correlations in machine learning models and introduces counterfactual invariance.

problem Spurious correlations in machine learning models that depend on irrelevant parts of input data.
method The paper uses causal inference to stress test models and introduces counterfactual invariance as a formal requirement.
result Counterfactual invariance is a requirement for models to be robust to irrelevant perturbations in input data.

Machine learning has made major advances in categorizing objects in images, yet the best algorithms miss important aspects of how people learn and think about categories. People can learn richer concepts from fewer examples, including causal models that explain how members of a category are formed. Here, we explore the…

2019-04-17abs ↗pdf ↗

Researchers use human-in-the-loop to create counterfactually augmented data, improving model performance.

problem Creating ML models less reliant on spurious patterns in NLP datasets.
method A human-in-the-loop process to curate counterfactually augmented data (CAD), prohibiting unnecessary edits.
result Models trained on CAD appear to rely less on semantically irrelevant words and generalize better out of domain.

Popular culture has contemplated societies of thinking machines for generations, envisioning futures from utopian to dystopian. These futures are, arguably, here now-we find ourselves at the doorstep of technology that can at least simulate the appearance of thinking, acting, and feeling. The real question is: now what…

2019-08-30abs ↗pdf ↗

A new reinforcement learning method for robots thinking and moving simultaneously.

problem Concurrent control in robotic systems where actions must be decided while the system is still evolving.
method Continuous-time Bellman equations, discretization aware of system delays, and architectural extension to deep reinforcement learning.
result The method successfully handles tasks requiring simultaneous decision-making and action execution.

Leo Breiman's Rashomon Effect and Occam Dilemma are re-evaluated in the context of modern machine learning.

problem The tradeoff between model complexity and accuracy in machine learning.
method Modern perspective on Breiman's arguments using current computational capabilities.
result Algorithmic models can be accurate without being complex, nullifying the Occam Dilemma.

Model learns brevity by exposing to easy problems, improving efficiency without explicit length penalties.

problem Excessive verbosity in step-by-step reasoning models trained with RLVR.
method Retaining and up-weighting moderately easy problems as implicit length regularizers.
result Model generates solutions that are, on average, nearly twice as short without explicit length penalties.

Thinking LLMs struggle with stock prediction, especially as data complexity increases.

problem Evaluating the performance of 'thinking' LLMs in stock prediction, especially under varying levels of cross-sectional complexity.
method Rolling 48m/1m walk-forward evaluation, comparing direct LLMs, TLLMs, and classical learners on cross-sectional ranking loss, MSE, and backtests with transaction costs.
result TLLMs' ranking quality deteriorates as cross-sectional complexity grows, while direct LLMs remain stable.

DGP learns speech recognition by modeling complex relationships between utterances.

problem Modeling complex relationships in speech recognition without relational data.
method Bayesian nonparametric deep learning method (DGP) that generates infinite probabilistic graphs.
result DGP successfully infers relationships among utterances without relational data during training.

Simple model outperforms neural networks on language understanding tasks.

problem Neural networks struggle with creating novel expressions from familiar ones.
method Attention-inspired modification of a baseline model, focusing on sequential thinking and acting.
result Simple model achieves good performance on gSCAN tasks, validating the benchmark.

The success of deep neural networks has inspired many to wonder whether other learners could benefit from deep, layered architectures. We present a general framework called forward thinking for deep learning that generalizes the architectural flexibility and sophistication of deep neural networks while also allowing fo…

2017-05-20abs ↗pdf ↗

Breiman's paper sparked debate on the future of statistics and machine learning.

problem The tension between traditional statistical modeling and model-free machine learning approaches.
method Discussion of the implications of machine learning's success and the need for new inferential approaches.
result The importance of understanding 'why' and 'if' questions in machine learning is now recognized.

Research suggests using deep learning for better recommendation systems.

problem Recommender systems rely on proxies for A/B testing, leading to random success.
method Advocates for using deep learning to improve recommendation performance.
result Deep learning can potentially optimize reward in recommendation systems.

In this paper we study n-composition series of affine manifolds. One composition series are classified using gerbe theory. It is natural to think that n-composition series must be classified using n-gerbe theory. In the last section of this, we propose a notion of abelian n-gerbe theory

2001-05-24abs ↗pdf ↗

We introduce Deep Reasoning Networks (DRNets), an end-to-end framework that combines deep learning with reasoning for solving complex tasks, typically in an unsupervised or weakly-supervised setting. DRNets exploit problem structure and prior knowledge by tightly combining logic and constraint reasoning with stochastic…

2019-06-03abs ↗pdf ↗

The paper converts metric bounds to distance function Hölder bounds and proves compactness theorems.

problem Proving geometric stability results with scalar curvature bounds.
method Transforming LpL^p bounds to Hölder bounds for distance functions.
result Compactness theorems and convergence guarantees for Riemannian manifolds.

We argue that the present crisis and stalling economy continuing since 2007 are rooted in the delusionary belief in policies based on a "perpetual money machine" type of thinking. We document strong evidence that, since the early 1980s, consumption has been increasingly funded by smaller savings, booming financial prof…

2012-12-12abs ↗pdf ↗

We prove new results on existence of solutions for the prescribed gaussian curvature problem on the euclidean sphere S^2. Those results are achieved by relating this problem with the holomorphic triples theory on Riemann surfaces. We think this approach might be applied to study some other semi-linear elliptic equation…

2015-03-19abs ↗pdf ↗

Doctors often rely on their past experience in order to diagnose patients. For a doctor with enough experience, almost every patient would have similarities to key cases seen in the past, and each new patient could be viewed as a mixture of these key past cases. Because doctors often tend to reason this way, an efficie…

2018-09-10abs ↗pdf ↗

Complex behaviors are often driven by an internal model, which integrates sensory information over time and facilitates long-term planning. Inferring an agent's internal model is a crucial ingredient in social interactions (theory of mind), for imitation learning, and for interpreting neural activities of behaving agen…

2018-05-24abs ↗pdf ↗

We prove some general results about quasi-actions on trees and define Property (QFA), which is analogous to Serre's Property (FA), but in the coarse setting. This property is shown to hold for a class of groups, including SL(n,Z)SL(n,\Z) for n3n\geq 3. We also give a way of thinking about Property (QFA) by breaking it down …

2003-10-06abs ↗pdf ↗

In this short paper, we re-derive the Bochner formula for the Laplacian by considering local variations of volume. The derivation is rooted in the fact that the Laplacian of a function measures the volume variation along the flow of the gradient vector of the function. Possible extensions of this approach/technique are…

2013-06-17abs ↗pdf ↗

This is a survey on the geometry of warped products, without, or essentially with only soft, calculation. Somewhere in the paper, the goal was to give a synthetic account since existing approaches are rather analytic. Somewhere else, we have interpreted statements, especially by means of a physical terminology. This is…

2011-07-02abs ↗pdf ↗

New framework learns disentangled causal representations from observed labels.

problem Learning meaningful disentangled causal representations from observed data.
method ICM-VAE framework using flow-based diffeomorphic functions and causal disentanglement prior.
result Induces highly disentangled causal factors and improves robustness.

Reinterprets Granger causality with causal Bayesian networks and Reichenbach's principles.

problem Lack of a rigorous causal foundation in Granger causality.
method Reinterpreting Granger causality through Reichenbach's principles and causal Bayesian networks, implementing as c-GC.
result c-GC provides a more principled framework for causal discovery in observational datasets.