Unified model explains international trade patterns using reinforced urns.
problem Understanding the complex patterns of international trade networks.
method A unified modelling framework using reinforced urns and the Reinforced Urn Process.
result The model predicts power law behavior and accounts for various network properties.
New urn problem considers unknown sampling method.
problem Optimal stopping in urn sampling with unknown method.
method Continuous-time analog and bounds on value function.
result Optimal strategy same for balanced urn, surprising.
New BAM model connects tensor factorization and topic models using Polya Urns.
problem Efficiently modeling and analyzing nonnegative tensors and topic distributions.
method Dynamic generative model BAM based on Poisson process and Polya-Bayes process.
result Developed efficient simulation algorithms for NTF and topic models.
New MAB model incentivizes user arm-pulling with self-reinforcing preferences.
problem Balancing exploration and exploitation in recommender systems with incentivized user preferences.
method Proposes a new MAB model with random arm selection and two policies: At-Least-n Explore-Then-Commit and UCB-List. result Achieves O(logT) expected regret and O(logT) expected payment over a time horizon T. URN neural network dynamically generates various neural structures during training.
problem Creating neural networks with flexible, dynamic structures during training.
method Introduced Unstructured Recursive Network (URN) and used gradient descent on a single loss function.
result Different neural structures can emerge from a single URN during training.
We describe the combinatorial stochastic process underlying a sequence of conditionally independent Bernoulli processes with a shared beta process hazard measure. As shown by Thibaux and Jordan [TJ07], in the special case when the underlying beta process has a constant concentration function and a finite and nonatomic …
A new sampler speeds up LDA topic modeling for big data.
problem Training LDA on large corpora is slow and requires dense memory storage.
method Uses a Pólya-urn-based approximation in a sparse partially collapsed sampler.
result The new sampler is faster and asymptotically exact.
Researchers develop methods to identify diffusion sources in tree networks.
problem Identifying the source of a diffusion in regular tree networks.
method Construct confidence sets for the diffusion source with size independent of the number of infected nodes, using probabilistic analysis of Pólya urns.
result It is possible to construct confidence sets for the diffusion source with size independent of the number of infected nodes.
Self-poisoning in adaptive OOD detectors is explained with a sharp threshold theory and certified calibration.
problem Self-poisoning in adaptive OOD detectors.
method Modeling bank impurity as a generalized Pólya urn, proving almost-sure convergence to a mean-field equilibrium.
result A certified admission gate removes the transition at every contamination rate, controlling false positives label-free.
PPT optimizes transformer behavior by steering its latent posterior using prior samples.
problem Eliciting desired behavior from transformers without backpropagation.
method Posterior Prefix Tuning (PPT) uses predictive Monte Carlo (PMC) samples and importance sampling to optimize the latent posterior.
result PPT optimizes transformer behavior without backpropagation, achieving high utility across different utility functions.
This paper attempts to find out numerically the distribution of the queue-length ratio in the context of a model of preferential attachment. Here we consider two restaurants only and a large number of customers (agents) who come to these restaurants. Each day the same number of agents sequentially arrives and decides w…
We present a Bayesian nonparametric framework for multilevel clustering which utilizes group-level context information to simultaneously discover low-dimensional structures of the group contents and partitions groups into clusters. Using the Dirichlet process as the building block, our model constructs a product base-m…
The paper introduces the concept of a cluster structure to define a joint distribution of the sample size and its exchangeable random partitions. The cluster structure allows the probability distribution of the random partitions of a subset of the sample to be dependent on the sample size, a feature not presented in a …
This paper proposes a technique for the unsupervised detection and tracking of arbitrary objects in videos. It is intended to reduce the need for detection and localization methods tailored to specific object types and serve as a general framework applicable to videos with varied objects, backgrounds, and image qualiti…
Model predicts capital flow and product share dynamics in international trade.
problem Understanding how capital flows between different industrial sectors affects product shares in international trade.
method Stochastic transfer model based on observed scaling relations.
result Model accurately predicts the distribution of product shares and identifies capital condensation.
The main contribution of this article is a new prior distribution over directed acyclic graphs, which gives larger weight to sparse graphs. This distribution is intended for structured Bayesian networks, where the structure is given by an ordered block model. That is, the nodes of the graph are objects which fall into …
We investigate the Heston model with stochastic volatility and exponential tails as a model for the typical price fluctuations of the Brazilian São Paulo Stock Exchange Index (IBOVESPA). Raw prices are first corrected for inflation and a period spanning 15 years characterized by memoryless returns is chosen for the ana…
Statistical network modeling has focused on representing the graph as a discrete structure, namely the adjacency matrix, and considering the exchangeability of this array. In such cases, the Aldous-Hoover representation theorem (Aldous, 1981;Hoover, 1979} applies and informs us that the graph is necessarily either dens…
Survey on Bayesian inference for Gaussian mixture models.
problem Estimating parameters of Gaussian mixture models using Bayesian methods.
method Uses Bayesian inference to estimate parameters and uncertainty of Gaussian mixture models.
result Bayesian approach provides point estimates and associated uncertainty for mixture model parameters.
Enhances deep reinforcement learning with object recognition.
problem Few works consider object characteristics in deep reinforcement learning.
method Proposes a novel method to incorporate object recognition into deep reinforcement learning models.
result Shows state-of-the-art results on Atari games.
Enhanced reinforcement learning using ensemble methods.
problem Improving reinforcement learning performance.
method Distributional reinforcement learning with ensemble group-aided training.
result Ensemble methods lead to more robust and efficient learning.
Boosted trees improve reinforcement learning solutions that are easy to understand.
problem Creating accurate reinforcement learning solutions that are also easy to understand.
method Using boosted regression trees to combine multiple regression trees.
result Boosted regression trees produce solutions that are as accurate as other methods but are also easy to understand.
Survey explores how transfer learning improves deep reinforcement learning.
problem Challenges in reinforcement learning efficiency and effectiveness.
method Categorizes and analyzes transfer learning approaches.
result Transfer learning enhances reinforcement learning performance.
This paper analyzes generalization issues in deep reinforcement learning.
problem Understanding and improving generalization capabilities of deep reinforcement learning policies.
method Formalizing and categorizing solutions to address overfitting in deep reinforcement learning.
result A comprehensive analysis of generalization challenges and solutions in deep reinforcement learning.
Proof of convergence for multi-objective optimization using inverse reinforcement learning.
problem Proving convergence in multi-objective optimization problems.
method Wasserstein inverse reinforcement learning with projective subgradient method and gradient descent.
result Convergence of inverse reinforcement learning for multi-objective optimization.
Paper discusses flaws in traditional RL for lifelong learning.
problem Traditional RL fails to model lifelong learning systems.
method Simplified prototype of lifelong RL system.
result Insights into lifelong RL, showing traditional RL's limitations.
Unified framework for efficient exploration in reinforcement learning.
problem Challenging control tasks in reinforcement learning.
method Combines Bayesian parameter updates with deep reinforcement learning.
result Practical algorithm achieves efficient exploration.
Non-deterministic policy improvement stabilizes reinforcement learning methods.
problem Instability in greedy policy improvement in approximated reinforcement learning.
method Non-deterministic policy improvement and suitable value function representation.
result Non-deterministic policy improvement stabilizes LSPI and other reinforcement learning methods.
Deep reinforcement learning enhances AI systems' visual understanding.
problem Scaling reinforcement learning to complex visual tasks.
method Value-based and policy-based methods, deep neural networks.
result Deep reinforcement learning enables autonomous systems to learn from raw visual inputs.
Paper robustifies reinforcement learning with risk-averse methods.
problem Making predictions robust to changes in system dynamics or rewards.
method Approximates Robust Reinforcement Learning using Φ-divergence and Risk-Averse formulation. result Classical Reinforcement Learning can be robustified using standard deviation penalization.
Improved reinforcement learning in Minecraft with human demonstrations.
problem Sample inefficiency in reinforcement learning.
method Training policy networks on human demonstrations first, then fine-tuning with reinforcement learning.
result Best agent achieved a mean score of 48 in Minecraft.
Deep reinforcement learning combines deep learning with reinforcement learning for complex tasks.
problem Complex tasks with high-dimensional data.
method Combining deep learning architectures (autoencoders, CNN, RNN) with reinforcement learning.
result Successful learning of useful representations for high-dimensional data.
The abstract explores connections between reinforcement learning, scaling, and diffusion.
problem Aligning reinforcement learning with human feedback and scaling techniques.
method Clarifying connections between reinforcement learning, scaling, and diffusion.
result Introducing a resampling approach for alignment and reward-directed diffusion models.
Reinforcement learning optimizes market making decisions.
problem Optimizing market making decisions in cryptocurrency markets.
method Multi-agent reinforcement learning framework with two agents: macro-agent and micro-agent.
result The proposed framework leads to successful trades in the Bitcoin market.
Paper analyzes CDRL algorithms for reinforcement learning.
problem Understanding theoretical properties of CDRL algorithms.
method Introduces a framework to analyze CDRL algorithms, establishes the importance of the projected distributional Bellman operator, draws connections to Cramér distance, and proves convergence.
result Proof of convergence for sample-based categorical distributional reinforcement learning algorithms.
Review of latest DRL algorithms with theoretical and practical insights.
problem Challenges in reinforcement learning and deep learning.
method Theoretical justification and empirical analysis of DRL algorithms.
result Empirical properties and practical limitations of DRL algorithms are discussed.
Implicit Policy simplifies complex reinforcement learning policies.
problem Complex action distributions in reinforcement learning.
method Rich policy class with entropy regularization.
result Entropy regularization with rich policy class achieves desirable properties.
Deep RL for portfolio management shows poor robustness.
problem Robustness of Deep RL algorithms in online portfolio management.
method Proposed a training and evaluation process for assessing DRL algorithms.
result Most Deep RL algorithms are not robust, generalizing poorly and degrading quickly.
Non-stationary reinforcement learning is challenging due to the complexity of updating value functions.
problem Challenges in non-stationary reinforcement learning, especially in updating value functions.
method Proved a worst-case complexity result for modifying reinforcement learning problems.
result Modifying reinforcement learning problems requires an amount of time almost as large as the number of states.
This paper calibrates deep reinforcement learning models to improve planning and performance.
problem Inaccurate predictive uncertainties from deep learning systems hinder model-based reinforcement learning.
method Describes a simple method to calibrate uncertainties in model-based reinforcement learning agents.
result Calibrated model-based reinforcement learning agents achieve state-of-the-art performance with fewer samples.
Study on meta-reinforcement learning generalization in high-dimensional tasks.
problem Generalization performance of meta-reinforcement learning algorithms in high-dimensional tasks.
method High-dimensional, procedurally generated environments.
result Meta-reinforcement learning algorithms exhibit strong overfitting on challenging tasks.
Unsupervised meta-learning speeds up reinforcement learning tasks.
problem Efficiently solving new reinforcement learning tasks.
method Formulating unsupervised meta-reinforcement learning and using mutual information for task proposals.
result Unsupervised meta-reinforcement learning effectively acquires accelerated procedures without manual task design.
New method uses reinforced regression for solving optimal stopping problems.
problem Solving optimal stopping problems in mathematical finance.
method Reinforced regression based on previously estimated continuation values.
result Illustrated by a numerical example from mathematical finance.
New algorithm for multi-agent reinforcement learning scales with number of agents.
problem Existing methods for deep multi-agent reinforcement learning struggle with increasing number of agents.
method Proposes a distributed optimization approach assuming policies of agents are close in parameter space.
result Demonstrates superior performance on co-operative and competitive tasks compared to existing methods.
A new framework optimizes molecules using deep reinforcement learning.
problem Optimizing molecules while maintaining chemical validity and drug-likeness.
method Combining deep reinforcement learning with domain knowledge of chemistry, MolDQN directly modifies molecules.
result MolDQN achieves optimization of molecules without bias from pre-training datasets.
A new method learns disentangled macro actions from sequences for reinforcement learning.
problem Curse of dimensionality in reinforcement learning action space.
method Autonomously learns disentangled factor representation of actions to generate macro actions.
result Higher scores in complex environments compared to other reinforcement learning algorithms.
Meta-reinforcement learning enables causal reasoning in complex environments.
problem Discovering causal structure in complex environments.
method Training a recurrent network with model-free reinforcement learning to solve problems with causal structure.
result The trained agent can perform causal reasoning in novel situations, select informative interventions, draw causal inferences, and make counterfactual predictions.
Epicurus' philosophy aligns with Reinforcement Learning principles.
problem Misconceptions about Epicurean philosophy and its relation to RL.
method Analysis of Epicurus' letters and comparison with RL concepts.
result Epicurean hedonism objective function is equivalent to RL objective function.