A common assumption about neural networks is that they can learn an appropriate internal representations on their own, see e.g. end-to-end learning. In this work we challenge this assumption. We consider two simple tasks and show that the state-of-the-art training algorithm fails, although the model itself is able to r…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
DORA analyzes deep neural networks' internal representations to detect spurious correlations.
Current machine learning techniques proposed to automatically discover a robot kinematics usually rely on a priori information about the robot's structure, sensors properties or end-effector position. This paper proposes a method to estimate a certain aspect of the forward kinematics model with no such information. An …
Neural networks learn distance-based representations, not just intensity.
Method determines latent dimensionality in international trade flows.
Enhances RL agents with predictive internal representations.
Recurrent neural networks (RNNs) are a vital modeling technique that rely on internal states learned indirectly by optimization of a supervised, unsupervised, or reinforcement training loss. RNNs are used to model dynamic processes that are characterized by underlying latent states whose form is often unknown, precludi…
New algorithms for hierarchical classification using conformal prediction.
MuZero visualizes its internal representations to stabilize planning.
Graphs are general and powerful data representations which can model complex real-world phenomena, ranging from chemical compounds to social networks; however, effective feature extraction from graphs is not a trivial task, and much work has been done in the field of machine learning and data mining. The recent advance…
This paper evaluates and validates cluster results using external and internal evaluation methods.
Autoencoder learns group representations from actions, improving future prediction accuracy.
Probably the most important problem in machine learning is the preliminary biasing of a learner's hypothesis space so that it is small enough to ensure good generalisation from reasonable training sets, yet large enough that it contains a good solution to the problem being learnt. In this paper a mechanism for {\em aut…
New approach extracts AI model representations for steering and monitoring.
A new framework detects anomalous inputs to DNNs.
This study investigates abrupt learning dynamics in Transformers, revealing plateau formation and internal representation collapse.
LLMs prefer Bitcoin under crisis frames, affecting financial decisions.
Study finds neural dialog models struggle with conversational tasks.
Goal-directed manipulation of representations is a key element of human flexible behaviour, while consciousness is often related to several aspects of higher-order cognition and human flexibility. Currently these two phenomena are only partially integrated (e.g., see Neurorepresentationalism) and this (a) limits our un…
Recurrent neural networks (RNN) are at the core of modern automatic speech recognition (ASR) systems. In particular, long-short term memory (LSTM) recurrent neural networks have achieved state-of-the-art results in many speech recognition tasks, due to their efficient representation of long and short term dependencies …
We introduce the BriarPatch, a pixel-space intervention that obscures sensitive attributes from representations encoded in pre-trained classifiers. The patches encourage internal model representations not to encode sensitive information, which has the effect of pushing downstream predictors towards exhibiting demograph…
New algorithm constrains SOMs to create supervised low-dimensional mappings.
The worldwide trade network has been widely studied through different data sets and network representations with a view to better understanding interactions among countries and products. Here we investigate international trade through the lenses of the single-layer, multiplex, and multi-layer networks. We discuss diffe…
Vision Transformers show different internal representations compared to CNNs.
Moment Pooling reduces latent space dimensions in machine learning models.
The utility of learning a dynamics/world model of the environment in reinforcement learning has been shown in a many ways. When using neural networks, however, these models suffer catastrophic forgetting when learned in a lifelong or continual fashion. Current solutions to the continual learning problem require experie…
New model explains how concepts grow based on experience.
Reinforcement learning (RL) algorithms allow artificial agents to improve their selection of actions to increase rewarding experiences in their environments. Temporal Difference (TD) Learning -- a model-free RL method -- is a leading account of the midbrain dopamine system and the basal ganglia in reinforcement learnin…
In this paper, we present a two-stage stochastic international portfolio optimisation model to find an optimal allocation for the combination of both assets and currency hedging positions. Our optimisation model allows a "currency overlay", or a deviation of currency exposure from asset exposure, to provide flexibility…
Study benchmarks 26 clustering validity measures.
TabPFN's internal geometry topology correlates with dataset reliability.
Using the work of Bonahon-Dreyer and Fock-Goncharov, one can construct a real-analytic parameterization for the PSL(n,R) Hitchin component of a surface S, that is explicitly analogous to the Fenchel-Nielsen coordinates on the Teichmuller space of S. Given a Hitchin representation, we give a lower bound on the "length" …
We propose a novel learning method for multilayered neural networks which uses feedforward supervisory signal and associates classification of a new input with that of pre-trained input. The proposed method effectively uses rich input information in the earlier layer for robust leaning and revising internal representat…
New approach uses prior knowledge to improve neural network representations.
Transformers infer tasks from context via two modes, geometrically shaped task vectors explain their behavior.
The discovery of novel materials and functional molecules can help to solve some of society's most urgent challenges, ranging from efficient energy harvesting and storage to uncovering novel pharmaceutical drug candidates. Traditionally matter engineering -- generally denoted as inverse design -- was based massively on…
Sparse Transformers degrade semantic information first, with early layers encoding more.
Study examines flaws in probing LLMs' knowledge and introduces a new method.
Recurrent neural networks (RNNs) are powerful architectures to model sequential data, due to their capability to learn short and long-term dependencies between the basic elements of a sequence. Nonetheless, popular tasks such as speech or images recognition, involve multi-dimensional input features that are characteriz…
An internal model of the own body can be assumed a fundamental and evolutionary-early representation as it is present throughout the animal kingdom. Such functional models are, on the one hand, required in motor control, for example solving the inverse kinematic or dynamic task in goal-directed movements or a forward t…
This work characterizes how data augmentation shapes neural representations.
Develops non-standard analysis for coherent risk estimation.
Internal Lagrangians derived from variational principles.
Model predicts internal fraud in retail banking is cyclical and influenced by corruption.
Deep networks are well-known to be fragile to adversarial attacks. We conduct an empirical analysis of deep representations under the state-of-the-art attack method called PGD, and find that the attack causes the internal representation to shift closer to the "false" class. Motivated by this observation, we propose to …
A 12D spinor encodes fermions in a 4D Kaluza-Klein model.
Transformer architectures show significant promise for natural language processing. Given that a single pretrained model can be fine-tuned to perform well on many different tasks, these networks appear to extract generally useful linguistic features. A natural question is how such networks represent this information in…
New measure shows how LSTM models compose hierarchical representations.