EMIX minimizes surprise in multi-agent reinforcement learning.
problem Surprise and approximation bias in multi-agent reinforcement learning.
method Energy-based MIXer (EMIX) for minimizing surprise across multiple agents.
result EMIX demonstrates consistent stable performance in challenging StarCraft II scenarios.
Surprise describes a range of phenomena from unexpected events to behavioral responses. We propose a measure of surprise and use it for surprise-driven learning. Our surprise measure takes into account data likelihood as well as the degree of commitment to a belief via the entropy of the belief distribution. We find th…
TradeR uses RL to execute trades in real markets, minimizing surprise and catastrophe.
problem Minimizing surprise and catastrophe in high-frequency trading.
method Hierarchical RL with energy-based surprise value function.
result TradeR outperforms in abrupt price changes and maintains profitability.
Study shows surprising cobordism distances between certain torus knots.
problem Determining cobordism distances between thin and thick torus knots.
method Analyzes locally flat cobordisms between torus knots with small and large braid indices.
result Surprising fact about torus knots as cross-sections of almost minimal cobordisms.
Every living organism struggles against disruptive environmental forces to carve out and maintain an orderly niche. We propose that such a struggle to achieve and preserve order might offer a principle for the emergence of useful behaviors in artificial agents. We formalize this idea into an unsupervised reinforcement …
Derives time-averaged active inference from control principles.
problem Finite-horizon or discounted-surprise problems in active inference.
method Derives infinite-horizon, average-surprise active inference from optimal control principles.
result Unified objective functional for sensorimotor control.
The study explores how agents learn and adapt preferences in dynamic environments.
problem Adaptive behavior and preference learning in reinforcement learning tasks.
method The approach involves self-supervised learning of preferences, distinguishing between environmental and intrinsic observations, and evaluating with model-free and model-based reinforcement learning.
result The methodology successfully minimizes surprisal and expected free energy in dynamic environments.
Model financial markets using information theory with a single parameter.
problem Capture the complexity of financial markets with a simple model.
method Derive an idealized model based on four information-theoretic assumptions, minimizing surprisal and divergence.
result The model uses squared radial Ornstein-Uhlenbeck processes for state variables and their sums.
Auto-Surprise automates recommender system selection and optimization.
problem Finding the best algorithm and hyperparameters for recommender systems.
method Extends Surprise library with TPE optimization for algorithm selection and hyperparameter tuning.
result Significantly faster in finding optimal hyperparameters compared to grid search.
Unifies 18 definitions of surprise, classifies them into four categories.
problem Lack of consensus on surprise definition.
method Technical classification into three groups based on agent's belief; conceptual categorization into four types.
result Taxonomy of surprise definitions provides foundation for brain studies.
New limits of minimal surface systems have surprising large interior parts.
problem Minimal surface system limits with large interior vertical and non-minimal portions.
method Construction of limits with smallest possible dimension and codimension.
result Limits of minimal surface systems can have surprising large interior parts.
Surprise-based learning allows agents to rapidly adapt to non-stationary stochastic environments characterized by sudden changes. We show that exact Bayesian inference in a hierarchical model gives rise to a surprise-modulated trade-off between forgetting old observations and integrating them with the new ones. The mod…
DG separates successes and failures by gating updates with advantage and surprisal.
problem Negative learning from surprising data in distributed reinforcement learning.
method DG gates each update with the product of advantage and surprisal, suppressing failures and preserving successes.
result DG outperforms other methods in various challenging reinforcement learning tasks.
The paper uses Bayesian Surprise to identify unexpected structures in indoor environments.
problem Identifying unexpected structures in indoor environments.
method Bayesian Surprise applied to Isovist Analysis of 2D floor plans.
result Surprise regions in indoor environments can be used to focus on important areas in LBS.
New mathematical surfaces without boundaries found.
problem Existence of nonlocal free boundary minimal surfaces.
method Fractional perimeter critical points with invariant boundary.
result Existence of nonlocal free boundary minimal surfaces without boundaries.
Exploration in environments with continuous control and sparse rewards remains a key challenge in reinforcement learning (RL). Recently, surprise has been used as an intrinsic reward that encourages systematic and efficient exploration. We introduce a new definition of surprise and its RL implementation named Variation…
Stochastic gradient descent outperforms traditional force-directed methods.
problem Improving graph layout quality and efficiency.
method Applying stochastic gradient descent for stress minimization.
result Stochastic gradient descent is simpler and more robust than traditional methods.
New proof shows nonholonomic motions are geodesics, minimizing distance.
problem Nonholonomic motion equations are not variational.
method Proved geodesic property of nonholonomic trajectories using Riemannian metrics.
result Nonholonomic motions minimize distance in their manifold.
The Surprise index assesses autonomous systems' competency in uncertain environments.
problem Evaluating competency of autonomous systems in dynamic, uncertain environments.
method Surprise index, a measure that quantifies system performance based on available data.
result The Surprise index can be computed for dynamic systems with Gaussian marginal distributions.
DE is a new exploration method that limits resource usage based on expected improvement and surprise.
problem Limited exploration in large action spaces when resources are scarce.
method Delight-gated exploration (DE) that limits exploration actions based on a gate price set by the product of expected improvement and surprise.
result DE outperforms ε-greedy and Thompson Sampling in terms of regret across various bandit and MDP settings. Paper uses surprisal to dynamically allocate computation between fast and slow models.
problem Dynamic allocation of computation in neural networks.
method Surprisal-based dynamic model selection.
result Model can match baseline performance with 15% fewer FLOPs.
We show that reinforcement learning agents that learn by surprise (surprisal) get stuck at abrupt environmental transition boundaries because these transitions are difficult to learn. We propose a counter-intuitive solution that we call Mutual Information Minimising Exploration (MIME) where an agent learns a latent rep…
We consider the problem of rank loss minimization in the setting of multilabel classification, which is usually tackled by means of convex surrogate losses defined on pairs of labels. Very recently, this approach was put into question by a negative result showing that commonly used pairwise surrogate losses, such as ex…
Curiosity-driven exploration using Bayesian surprise in latent space.
problem Enhance exploration capabilities in reinforcement learning.
method Apply Bayesian surprise in a latent space to favor exploration.
result Our method is computationally cheap and performs well on various tasks.
Model compresses event-like contexts using gated surprise signals.
problem Perceiving a dynamic world as organized events.
method Hierarchical, surprise-gated recurrent neural network architecture.
result Achieves best performance on multiple event processing tasks.
In this paper, we describe a new surprising example of a fibration of the Clifford torus S3 x S3 in the 7-sphere by great 3-spheres, which is fiberwise homogeneous but whose fibers are not parallel to one another. In particular it is not part of a Hopf fibration. A fibration is fiberwise homogeneous when for any two fi…
The purpose of the present paper is to introduce and explore two surprises that arise when we apply a standard procedure to study the number of finite type invariants of 3-manifolds introduced independently by M. Goussarov and K. Habiro based on surgery on claspers, Y-graphs or clovers, \cite{Gu,Ha,GGP}. One surprise i…
Surprising circles found in Coxeter group boundaries.
problem Embedded circles in Morse boundaries of Coxeter groups.
method Analysis of Morse boundaries and defining graphs.
result Circles not arising from visible Fuchsian subgroups.
SAE-FiRE extracts key financial info from long documents, improving earnings surprise predictions.
problem Predicting earnings surprises from long, redundant financial documents.
method Sparse Autoencoder feature selection to filter out noise and identify key dimensions.
result SAE-FiRE significantly outperforms baseline approaches in financial datasets.
SRFE clarifies KL divergences without unifying learning frameworks.
problem Inductive biases of KL divergences and their limitations.
method Introducing SRFE, a log-moment-based functional of the likelihood ratio.
result SRFE recovers KL divergences as limits and reveals a mean-variance tradeoff.
A model simulates how different types of traders react to macroeconomic news.
problem Understanding how various market participants respond to macroeconomic surprises.
method Developed a calibrated data generation process (DGP) with four trader archetypes and a Monte Carlo simulation.
result Higher information and lower risk-averse traders take larger positions and achieve higher average wealth.
New shapes enclose less volume than the sphere, surprising in 3D.
problem Finding the minimal volume enclosed by smooth spheres with bounded curvatures.
method Produced a family of bodies parameterized by ε, each bounded by a smooth topological sphere with principal curvatures in [-1, 1].
result The unit sphere does not enclose the minimal volume among all smooth spheres in R^3 with principal curvatures in [-1, 1].
Plotting a learner's average performance against the number of training samples results in a learning curve. Studying such curves on one or more data sets is a way to get to a better understanding of the generalization properties of this learner. The behavior of learning curves is, however, not very well understood and…
We establish several new stylised facts concerning the intra-day seasonalities of stock dynamics. Beyond the well known U-shaped pattern of the volatility, we find that the average correlation between stocks increases throughout the day, leading to a smaller relative dispersion between stocks. Somewhat paradoxically, t…
LemonadeBench evaluates LLMs' economic intuition through a simulated lemonade stand.
problem Evaluating LLMs' economic understanding and decision-making in simple markets.
method Simulated lemonade stand business to test LLMs' long-term planning and profit maximization.
result Models achieve profitability but exhibit local rather than global optimization.
In 1974, Gehring posed the problem of minimizing the length of two linked curves separated by unit distance. This constraint can be viewed as a measure of thickness for links, and the ratio of length over thickness as the ropelength. In this paper we refine Gehring's problem to deal with links in a fixed link-homotopy …
Bitcoin reacts negatively to inflation surprises, contrary to belief.
problem Bitcoin's ability to hedge inflation is questioned.
method Examined cryptocurrency responses to macroeconomic news announcements.
result Bitcoin's price decreases by 24 bps in response to inflationary surprises.
New framework detects near vs. far out-of-distribution samples for AI safety.
problem Binary OOD detection fails to distinguish between semantically close and distant unknown risks.
method Ternary classification based on Low-Entropy Semantic Manifolds and Semantic Surprise Vector.
result Framework achieves state-of-the-art performance on ternary OOD detection task.
Firms disclosing positive earnings surprises are more likely to disclose ESG information.
problem Transparency vs. performance in financial markets.
method Empirical analysis of earnings surprises and ESG disclosures.
result Positive earnings firms disclose more ESG information than negative earnings firms.
Study on double descent behavior in two-layer neural networks for binary classification.
problem Understanding the double descent phenomenon in model test error.
method Two-layer neural network with ReLU activation for binary classification. Quantified model size by sample-to-dimension ratio. Empirical risk minimization using Convex Gaussian Min Max Theorem.
result Observed and investigated the double descent behavior of model test error.
We establish when the two problems of minimizing a function of lifetime minimum wealth and of maximizing utility of lifetime consumption result in the same optimal investment strategy on a given open interval O in wealth space. To answer this question, we equate the two investment strategies and show that if the indi…
Chance-constrained ActInf allows for small violations of constraints to drive goal-directed behavior.
problem Goal-directed behavior constrained by prior beliefs.
method Introducing chance constraints to ActInf, allowing for small violations of constraints.
result Chance-constrained ActInf allows for a trade-off between robust control and chance constraint violation.
New proof shows D-SGD and SAM are equivalent, revealing advantages of decentralization.
problem The generalization benefits of decentralized learning.
method Proved D-SGD implicitly minimizes SAM's loss function.
result Decentralized SGD and Average-direction SAM are asymptotically equivalent.
One of the most surprising and exciting discoveries in supervised learning was the benefit of overparameterization (i.e. training a very large model) to improving the optimization landscape of a problem, with minimal effect on statistical performance (i.e. generalization). In contrast, unsupervised settings have been u…
We discuss eight new(?) configuration theorems of classical projective geometry in the spirit of the Pappus and Pascal theorems.
A new k-NN algorithm using surprisal for robust and interpretable nonparametric learning.
problem Complex patterns and relationships in data without strong distribution assumptions.
method Surprisal-driven k-NN framework for classification, regression, density estimation, and anomaly detection. result State-of-the-art results in classification and anomaly detection, competitive regression results.
We study randomized sketching methods for approximately solving least-squares problem with a general convex constraint. The quality of a least-squares approximation can be assessed in different ways: either in terms of the value of the quadratic objective function (cost approximation), or in terms of some distance meas…
As traditional neural network consumes a significant amount of computing resources during back propagation, \citet{Sun2017mePropSB} propose a simple yet effective technique to alleviate this problem. In this technique, only a small subset of the full gradients are computed to update the model parameters. In this paper …