Solves a model for sudden problem-solving ability in deep learning.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Deep learning models generate languages that lack abstract reasoning.
Noise-driven neural networks emerge modular structures, improving robustness and generalization.
MPNNs generalize poorly, study reviews current research.
Infants are experts at playing, with an amazing ability to generate novel structured behaviors in unstructured environments that lack clear extrinsic reward signals. We seek to replicate some of these abilities with a neural network that implements curiosity-driven intrinsic motivation. Using a simple but ecologically …
A simple model explains phase transition in large language models.
The VAE's reconstruction ability is studied using PAC-Bayes theory.
Emergenet predicts animal influenza strain emergence, outperforming current methods.
How do we know if communication is emerging in a multi-agent system? The vast majority of recent papers on emergent communication show that adding a communication channel leads to an increase in reward or task success. This is a useful indicator, but provides only a coarse measure of the agent's learned communication a…
Neural networks learn discrete tasks on continuous data via emergent geometry.
Neural networks learn task-specific features, influenced by nonlinearity.
Many recent works have discussed the propensity, or lack thereof, for emergent languages to exhibit properties of natural languages. A favorite in the literature is learning compositionality. We note that most of those works have focused on communicative bandwidth as being of primary importance. While important, it is …
New PAC-Bayes bounds use Wasserstein distances to improve generalization.
New approach predicts tokens in context, explaining how ICL emerges.
Computational swarm intelligence consists of multiple artificial simple agents exchanging information while exploring a search space. Despite a rich literature in the field, with works improving old approaches and proposing new ones, the mechanism by which complex behavior emerges in these systems is still not well und…
LaRT models LLMs' response accuracy and CoT length to evaluate reasoning ability and speed.
Heavy-tailed distributions emerge in SGD's parameter evolution.
Survey examines distillation methods for large language models.
Nonnegative matrix factorization (NMF) is a powerful tool for data mining. However, the emergence of `big data' has severely challenged our ability to compute this fundamental decomposition using deterministic algorithms. This paper presents a randomized hierarchical alternating least squares (HALS) algorithm to comput…
Extending spatio-temporal scale limitations of models for complex atomistic systems considered in biochemistry and materials science necessitates the development of enhanced sampling methods. The potential acceleration in exploring the configurational space by enhanced sampling methods depends on the choice of collecti…
Recently the GAN generated face images are more and more realistic with high-quality, even hard for human eyes to detect. On the other hand, the forensics community keeps on developing methods to detect these generated fake images and try to guarantee the credibility of visual contents. Although researchers have develo…
This work tackles representation learning by introducing stochastic competition-based activations.
Transformers can cluster data from Gaussian mixtures without supervision.
Study investigates how simple speech sounds can form abstract categories.
The challenge in controlling stochastic systems in which low-probability events can set the system on catastrophic trajectories is to develop a robust ability to respond to such events without significantly compromising the optimality of the baseline control policy. This paper presents CelluDose, a stochastic simulatio…
A model for groundwater trading among stakeholders.
In this work we afford the statistical characterization of a linear Stochastic Volatility Model featuring Inverse Gamma stationary distribution for the instantaneous volatility. We detail the derivation of the moments of the return distribution, revealing the role of the Inverse Gamma law in the emergence of fat tails,…
Deep Reinforcement Learning (DRL) has emerged as a powerful control technique in robotic science. In contrast to control theory, DRL is more robust in the thorough exploration of the environment. This capability of DRL generates more human-like behaviour and intelligence when applied to the robots. To explore this capa…
This work shows how transformers use multi-concept word semantics for efficient in-context learning.
A distinctive property of human and animal intelligence is the ability to form abstractions by neglecting irrelevant information which allows to separate structure from noise. From an information theoretic point of view abstractions are desirable because they allow for very efficient information processing. In artifici…
There is emerging interest in performing regression between distributions. In contrast to prediction on single instances, these machine learning methods can be useful for population-based studies or on problems that are inherently statistical in nature. The recently proposed distribution regression network (DRN) has sh…
Low-rank training improves neural network training on edge devices with non-volatile memory.
Transformers can approximate posterior predictive distributions through in-context learning.
We develop a model to study the role of rationality in economics and biology. The model's agents differ continuously in their ability to make rational choices. The agents' objective is to ensure their individual survival over time or, equivalently, to maximize profits. In equilibrium, however, rational agents who maxim…
A recent trend in machine learning has been to enrich learned models with the ability to explain their own predictions. The emerging field of Explainable AI (XAI) has so far mainly focused on supervised learning, in particular, deep neural network classifiers. In many practical problems however, label information is no…
Most artificial intelligence models have limiting ability to solve new tasks faster, without forgetting previously acquired knowledge. The recently emerging paradigm of continual learning aims to solve this issue, in which the model learns various tasks in a sequential fashion. In this work, a novel approach for contin…
New IRT method identifies useful datasets for ML classifier evaluation.
We demystify attention patterns in multi-head softmax models for linear data.
Study shows neural ODEs generalize well on synthetic graphs but struggle with degree heterogeneity and clustering.
The global financial crisis in 2007-2009 demonstrated that systemic risk can spread all over the world through a complex web of financial linkages, yet we still lack fundamental knowledge about the evolution of the financial web. In particular, interbank credit networks shape the core of the financial system, in which …
Panda predicts chaotic systems without retraining, showing emergent properties.
Paper develops an attention mechanism for long-term scientific impact prediction.
Improved noise estimation in latent neural SDEs enhances model accuracy.
The paper examines efficient algorithms for linear regression over resource-limited networks.
Tailoring the presentation of information to the needs of individual students leads to massive gains in student outcomes~\cite{bloom19842}. This finding is likely due to the fact that different students learn differently, perhaps as a result of variation in ability, interest or other factors~\cite{schiefele1992interest…
Modern society heavily relies on strongly connected, socio-technical systems. As a result, distinct risks threatening the operation of individual systems can no longer be treated in isolation. Consequently, risk experts are actively seeking for ways to relax the risk independence assumption that undermines typical risk…
The equity risk premium puzzle is that the return on equities has far exceeded the average return on short-term risk-free debt and cannot be explained by conventional representative-agent consumption based equilibrium models. We review a few attempts done over the years to explain this anomaly: 1. Inclusion of highly u…
Bayesian hybrid models fuse physics-based insights with machine learning constructs to correct for systematic bias. In this paper, we compare Bayesian hybrid models against physics-based glass-box and Gaussian process black-box surrogate models. We consider ballistic firing as an illustrative case study for a Bayesian …