New algorithm learns efficiently in multi-agent settings.
problem Efficient learning in multi-agent Markov decision processes.
method Cooperative Prioritized Sweeping: model-based reinforcement learning with sample efficiency.
result Outperforms state-of-the-art on SysAdmin and randomized environments.
Enhances cooperative multi-task SemCom for distributed users.
problem Performance degradation in cooperative multi-tasking due to negative information transfer.
method Federated learning (FL) with semantic-aware task clustering.
result Constructive cooperation across distributed users with semantic-aware task clustering.
Exploration efficiency is a challenging problem in multi-agent reinforcement learning (MARL), as the policy learned by confederate MARL depends on the collaborative approach among multiple agents. Another important problem is the less informative reward restricts the learning speed of MARL compared with the informative…
Modern reinforcement learning algorithms reach super-human performance on many board and video games, but they are sample inefficient, i.e. they typically require significantly more playing experience than humans to reach an equal performance level. To improve sample efficiency, an agent may build a model of the enviro…
Continuous Sweep improves binary quantifier performance.
problem Estimating class prevalence in datasets.
method Parametric binary quantifier inspired by Median Sweep, using parametric class distributions and mean of Adjusted Count estimates.
result Continuous Sweep outperforms other quantifiers in simulations and empirical data analysis.
A new MARL framework for community-based cooperation with transfer and active exploration.
problem Flexible coordination patterns in multi-agent systems with community structures.
method Community-based multi-agent reinforcement learning with transfer and active exploration.
result Provably convergent actor-critic algorithms for structured information sharing and transfer learning.
Two sweeps of the Brennan-Schwartz algorithm solve American options under negative rates.
problem Inability of the Brennan-Schwartz algorithm to solve American options under negative interest rates.
method Two sweeps of the Brennan-Schwartz algorithm in two directions.
result Recovery of the exact solution for American options under negative rates.
We prove the absence of a universal diameter bound on lengths of curves in a sweep-out of a Riemannian 2-sphere. If such bound existed it would yield a simple proof of existence of short geodesic segments and closed geodesics on a sphere of small diameter.
Trading floors need to be twice as deep as electronic markets to compete.
problem Informed traders prefer fast electronic markets over slow trading floors.
method Examined the performance of trading floors and electronic markets in a hybrid system.
result Trading floors need to be twice as deep as electronic markets to compete.
Paper estimates area covered by a line-sweep sensor in robotics.
problem Accurately estimating the area covered by a line-sweep sensor.
method Relies on coverage measure and topological degree in the plane.
result Guaranteed characterization of the explored area using interval analysis.
AI-driven sales prioritization boosts renewal bookings by 8.08%.
problem Manual sales account prioritization is inefficient and under-invested.
method Developed an AI-based Account Prioritizer using machine learning and explanation algorithms.
result Generated a +8.08% increase in renewal bookings.
Unified framework improves gene prioritization in disease studies.
problem Identifying genes involved in diseases using heterogeneous biological data.
method Network propagation-based gene prioritization with integrated biological information.
result Significant improvements in prioritizing genes not identified by traditional methods.
Develops methods for cooperative Bayesian inference.
problem Cooperation between learning agents.
method Sequential Bayesian inference approaches.
result Theoretical foundation for cooperative inference.
A new method prioritizes and recycles experiences for better reinforcement learning.
problem Improving reinforcement learning efficiency by prioritizing and recycling experiences.
method Double-prioritized state-recycled (DPSR) experience replay.
result DPSR achieved state-of-the-art results in Atari games, outperforming original and prioritized methods.
It is known that the almost-Kaehler anti-self-dual metrics on a given 4-manifold sweep out an open subset in the moduli space of anti-self-dual metrics. However, we show here by example that this subset is not generally closed, and so need not sweep out entire connected components in the moduli space. Our construction …
Multi-cell cooperative processing with limited backhaul traffic is studied for cellular uplinks. Aiming at reduced backhaul overhead, a sparsity-regularized multi-cell receive-filter design problem is formulated. Both unstructured distributed cooperation as well as clustered cooperation, in which base station groups ar…
ReaPER improves learning efficiency by prioritizing reliable experiences.
problem Inefficient sampling of past experiences in reinforcement learning.
method Introducing a novel measure of reliability to prioritize experiences in PER.
result ReaPER outperforms PER in various environments, including Atari-10.
Develops an equilibrium model for securities pricing in a mixed cooperative and non-cooperative market.
problem Equilibrium pricing of securities in a market with cooperative and non-cooperative agents.
method Conditional extended mean-field control for cooperative agents, mean-field model for both cooperative and non-cooperative agents.
result Existence of a unique equilibrium for both finite-agent and mean-field models under certain conditions.
RATE metrics evaluate treatment prioritization rules, subsuming existing methods.
problem Comparing and testing the quality of treatment prioritization rules.
method Rank-weighted average treatment effect (RATE) metrics.
result RATE metrics enable asymptotically exact inference in various study settings.
Stable cooperation emerges in fluctuating environments.
problem Evolutionary stability of cooperation in fluctuating conditions.
method Agents with fluctuating wealth share public goods, leading to cooperation.
result Agents with cooperation produce an advantage in fluctuating environments.
Decoupled PFNs improve sequential decision-making by separating epistemic and aleatoric uncertainties.
problem Sequential decision-making requires distinguishing between epistemic uncertainty about latent signals and irreducible aleatoric observation noise.
method Developed a decoupled PFN architecture that uses query-level labels to train separate heads for latent signal and aleatoric noise.
result Empirically, decoupled PFNs mitigate the failure mode of total-variance exploration in noisy and heteroscedastic settings.
The paper proves the existence of CMC surfaces with controlled topology in 3-manifolds.
problem Proving the existence of constant mean curvature surfaces with specific topological constraints.
method Min-max construction and convergence to a CMC-parametrized varifold.
result Existence of a non-trivial, branched immersion of a closed Riemann surface with constant mean curvature in a 3-manifold.
The paper tackles cooperative RL with function approximation, achieving near-optimal learning with limited communication.
problem Cooperative multi-agent reinforcement learning with function approximation.
method Careful message-passing and cooperative value iteration.
result Achieving near-optimal no-regret learning with limited communication in cooperative multi-agent settings.
Experience replay is widely used in deep reinforcement learning algorithms and allows agents to remember and learn from experiences from the past. In an effort to learn more efficiently, researchers proposed prioritized experience replay (PER) which samples important transitions more frequently. In this paper, we propo…
The cooperative hierarchical structure is a common and significant data structure observed in, or adopted by, many research areas, such as: text mining (author-paper-word) and multi-label classification (label-instance-feature). Renowned Bayesian approaches for cooperative hierarchical structure modeling are mostly bas…
We study the explore-exploit tradeoff in distributed cooperative decision-making using the context of the multiarmed bandit (MAB) problem. For the distributed cooperative MAB problem, we design the cooperative UCB algorithm that comprises two interleaved distributed processes: (i) running consensus algorithms for estim…
Proposes a method to cluster tasks for constructive cooperative multi-tasking.
problem Destructive cooperation in cooperative multi-tasking.
method Semantic clustering followed by end-to-end joint training within clusters.
result Effective mitigation of destructive cooperation and negative transfer.
CSAC enables cooperative reinforcement learning for multi-stage tasks.
problem Coordinating consecutive reinforcement learning agents for long-term multi-stage tasks.
method CSAC modifies each agent's policy to maximize both current and next agent's critic.
result CSAC outperforms uncooperative policies and single-agent training in multi-room maze domain.
Develops an algorithm to find the best subset of points for maximizing the coefficient of determination.
problem Finding the optimal subset of points for maximizing the coefficient of determination in robust correlation analysis.
method The extit{quadratic sweep} method, which involves projecting points into \(\mathbb{R}^5\) and iterating over linearly separable \(k\)-subsets.
result The method optimally finds the best subset of points for maximizing the coefficient of determination without error over several million trials up to \(n=30\).
Cooperation is a persistent behavioral pattern of entities pooling and sharing resources. Its ubiquity in nature poses a conundrum. Whenever two entities cooperate, one must willingly relinquish something of value to the other. Why is this apparent altruism favored in evolution? Classical solutions assume a net fitness…
Cooperative communication plays a central role in theories of human cognition, language, development, culture, and human-robot interaction. Prior models of cooperative communication are algorithmic in nature and do not shed light on why cooperation may yield effective belief transmission and what limitations may arise …
Paper tackles non-uniform coverage planning for robots.
problem Non-uniform coverage planning for robots that need to visit some points more frequently.
method Proposes a novel reinforcement learning approach in a Semi-Markov Decision Process.
result Significant improvement over existing greedy approach in simulations.
Study efficient power iteration for tensor models, proving convergence under specific conditions.
problem Simultaneous alternating power iteration for fixed-order asymmetric rank-one spiked tensor models.
method Finite-iteration local theory, geometrically decaying transient, fixed-order multilinear noise event, warm-start mechanism.
result Convergence to the unique informative local fixed point under specific conditions.
Facing a heavy task, any single person can only make a limited contribution and team cooperation is needed. As one enjoys the benefit of the public goods, the potential benefits of the project are not always maximized and may be partly wasted. By incorporating individual ability and project benefit into the original pu…
New method combines heuristics and search techniques to speed up cooperative planning for autonomous vehicles.
problem Efficient cooperative planning for autonomous vehicles in complex traffic scenarios.
method Combining learned heuristics with Monte Carlo Tree Search (MCTS) to guide search towards promising actions.
result Better solutions at lower computational costs achieved through accelerated planning.
Kernel method improves cooperative decision-making among agents.
problem Cooperative multi-agent decision making with contextual information.
method Proposed extsc{Coop-KernelUCB} algorithm for near-optimal per-agent regret.
result Near-optimal bounds on per-agent regret with efficient computation and communication.
Our goal is to generalize the Choe-Hoppe helicoid and Clifford cones in Euclidean space. By sweeping out L indpendent Clifford cones in R2N+2 via the multi-screw motion, we construct minimal submanifolds in RL(2N+2)+1. Also, we sweep out the L-rays Clifford cone (introduced in Sectio…
Improved cooperation between levels boosts reinforcement learning performance.
problem Training multi-level policies in hierarchical reinforcement learning.
method Modeling policy optimization as a multi-agent process and inducing cooperation between sub-policies.
result Inducing cooperation between sub-policies leads to stronger and more sample-efficient policies.
We consider in a market model the cooperative emergence of value due to a positive feedback between perception of needs and demand. Here we consider also a negative feedback from production of the traded products, and find that this cooperativity is robust, provided that the production rate is slow. Cooperativity is fo…
A new algorithm reduces regret in cooperative multi-agent bandits with heavy-tailed data.
problem Cooperative multi-agent bandits with heavy-tailed data.
method MP-UCB algorithm incorporating robust estimation with message-passing protocol.
result Optimal regret bounds for MP-UCB in various settings.
SMC analysis reveals key transient effects in macroeconomic ABM.
problem Analysis of complex ABMs is challenging and often relies on ad hoc methods.
method Statistical model checking (SMC) implemented through MultiVeStA.
result Clear contrast across parameter families in macro-financial and structural sweeps.
New algorithm reduces individual regret and communication costs in cooperative bandits.
problem Optimal individual and group regret in cooperative multi-agent bandits.
method Integrates a new communication policy into a learning algorithm.
result Achieves optimal individual regret and constant communication costs.
Cooperation information sharing is important to theories of human learning and has potential implications for machine learning. Prior work derived conditions for achieving optimal Cooperative Inference given strong, relatively restrictive assumptions. We relax these assumptions by demonstrating convergence for any disc…
Asynchronous cooperative learning rules ensure all agents converge to correct hypothesis.
problem Cooperative learning in networks with unreliable communication.
method Proposed robust cooperative learning rule for weak communication networks.
result All agents' beliefs exponentially decay to the correct hypothesis.
Most exact methods for k-nearest neighbour search suffer from the curse of dimensionality; that is, their query times exhibit exponential dependence on either the ambient or the intrinsic dimensionality. Dynamic Continuous Indexing (DCI) offers a promising way of circumventing the curse and successfully reduces the dep…
Study shows cooperation can improve everyone's market efficiency.
problem Understanding when collective cooperation improves market efficiency.
method General semimartingale framework, deriving necessary and sufficient conditions.
result Strict improvement in each agent's indirect utility when cooperation is beneficial.
A variety of cooperative multi-agent control problems require agents to achieve individual goals while contributing to collective success. This multi-goal multi-agent setting poses difficulties for recent algorithms, which primarily target settings with a single global reward, due to two new challenges: efficient explo…
Elucidating the genetic basis of human diseases is a central goal of genetics and molecular biology. While traditional linkage analysis and modern high-throughput techniques often provide long lists of tens or hundreds of disease gene candidates, the identification of disease genes among the candidates remains time-con…