This paper explores how environmental properties can simplify reinforcement learning in non-episodic settings.
problem Challenges in reinforcement learning with continuous interaction and sparse delayed rewards.
method Analysis of environment shaping and dynamism properties to simplify learning.
result Properties like environment shaping and dynamism can significantly ease learning in non-episodic, sparse reward settings.
New algorithms reduce reinforcement learning regret in factored MDPs.
problem Optimizing reinforcement learning in non-episodic factored MDPs.
method Proposed two near-optimal and oracle-efficient algorithms for FMDPs.
result Oracle-efficient algorithms achieve near-optimal regret bounds of O ( D S A T ) O(DS\sqrt{AT}) O ( D S A T ) . Paper proposes new γ γ γ -regret measure for non-episodic RL.
problem Measuring performance in non-episodic RL environments.
method Introduces γ γ γ -regret as a new performance measure and derives bounds. result Closed the gap between lower and upper bounds for γ γ γ -regret. We give a simple optimistic algorithm for which it is easy to derive regret bounds of O ~ ( t m i x S A T ) \tilde{O}(\sqrt{t_{\rm mix} SAT}) O ~ ( t mix S A T ) after T T T steps in uniformly ergodic Markov decision processes with S S S states, A A A actions, and mixing time parameter t m i x t_{\rm mix} t mix . These bounds are the first regret bounds in the general, non-epi…
New RL method learns K-step lookahead Q-functions for fixed-horizon MDPs.
problem Challenges in online reinforcement learning for non-episodic, finite-horizon MDPs.
method Introduces a K-step lookahead Q-function with a time-varying threshold for selecting actions.
result Achieves minimax optimal constant regret for K=1 and O ( max ( ( K − 1 ) , C K − 1 ) S A T log ( T ) ) \mathcal{O}(\max((K-1),C_{K-1})\sqrt{SAT\log(T)}) O ( max (( K − 1 ) , C K − 1 ) S A T log ( T ) ) regret for K ≥ 2. Restless bandit problems assume time-varying reward distributions of the arms, which adds flexibility to the model but makes the analysis more challenging. We study learning algorithms over the unknown reward distributions and prove a sub-linear, O ( T log T ) O(\sqrt{T}\log T) O ( T log T ) , regret bound for a variant of Thompson sampling. Our…
Control of non-episodic, finite-horizon dynamical systems with uncertain dynamics poses a tough and elementary case of the exploration-exploitation trade-off. Bayesian reinforcement learning, reasoning about the effect of actions and future observations, offers a principled solution, but is intractable. We review, then…
New algorithm for average reward learning with bounded hitting time assumption.
problem Minimizing regret in average reward reinforcement learning with bounded hitting time.
method Optimistic Q-learning with a novel L ‾ \overline{L} L operator for bounded hitting time. result Regret bound of i l d e O ( H 5 S A T ) ilde{O}(H^5 S\sqrt{AT}) i l d e O ( H 5 S A T ) for average reward learning. New algorithm reduces sample complexity for online reinforcement learning.
problem Reducing sample complexity for online reinforcement learning in nonlinear systems.
method Generalized algorithm for various dynamical systems, including neural networks.
result Achieves policy regret of O(Nε^2 + d_u ln(m(ε))/ε^2) in general settings.
This paper sets communication complexity bounds for distributed RL.
problem Establishing minimum communication requirements for distributed RL.
method Information-theoretic lower bounds and algorithm development.
result Developed algorithms achieving optimal risk up to logarithmic factors.
New algorithm learns FMDP structure while minimizing regret.
problem Regret minimization in FMDPs with unknown structure.
method Optimism in face of uncertainty principle combined with statistical structure learning.
result First algorithm to learn FMDP structure while minimizing regret.
Study designs incentives for adapting multi-agent systems without knowing their learning dynamics.
problem Designing incentives for an adapting population in multi-agent systems without prior knowledge of their learning dynamics.
method Introduces a model-based non-episodic Reinforcement Learning (RL) formulation for steering Markovian agents towards desired policies, focusing on history-dependent strategies to handle model uncertainty.
result Identifies conditions for the existence of steering strategies to guide agents to desired policies and provides empirical algorithms to approximately solve the objective.
A new reinforcement learning method improves Max-Cut solutions without needing training data.
problem Max-Cut problem is NP-hard, and existing methods struggle with generalizability and scalability.
method Training-data-free reinforcement learning approach to hyperplane rounding for Max-Cut optimization.
result Our method consistently achieves better Max-Cut solutions across various graph types.
A graph bandit algorithm learns optimal paths on unknown graphs.
problem Optimal path selection on unknown graphs under uncertainty.
method G-UCB algorithm based on offline graph planning and optimism principle.
result Achieves tight regret bound of O ( ∣ S ∣ T log ( T ) + D ∣ S ∣ log T ) O(\sqrt{|S|T\log(T)}+D|S|\log T) O ( ∣ S ∣ T log ( T ) + D ∣ S ∣ log T ) . UCRL-CMDP algorithm optimizes RL with constraints on average costs.
problem Optimizing RL in MDPs with average cost constraints.
method Model-based RL algorithms maximizing reward while keeping costs within bounds.
result UCRL-CMDP algorithm's expected regret is upper-bounded by $T^{2\slash 3}$ .
Study shows Julia sets and gasket limit sets are quasiconformally different.
problem Quasiconformal non-equivalence of Julia sets and gasket limit sets.
method Proved quasiconformal non-equivalence of Julia sets and gasket limit sets.
result Julia sets and gasket limit sets are quasiconformally different.
Study shows non-symmetric convex sets have full boundary limits.
problem Understanding boundaries of non-symmetric convex sets.
method Proved using proximal limit set analysis.
result Proximal limit set equals full projective boundary for non-symmetric irreducible divisible convex sets.
The paper analyzes set-to-set matching with neural networks, focusing on theoretical generalization.
problem Theoretical analysis of set-to-set matching with neural networks.
method Generalization error analysis of set-to-set matching with neural networks.
result Theoretical insights into the behavior of set-to-set matching models.
Generative model learns to autoencode and generate sets of images.
problem Learning to represent and generate sets of images with unknown number of sets.
method Set Distribution Networks (SDNs) learn set encoder, discriminator, generator, and prior.
result SDNs can reconstruct and generate sets of images with preserved attributes.
Study on cold and freezing sets in digital images.
problem Properties of cold sets in digital images.
method Analysis of properties and relationships between cold and freezing sets.
result Examined relationships between cold and freezing sets.
Paper solves whether zero sets are mapping degree sets.
problem Whether finite sets containing zero are mapping degree sets.
method Examined oriented closed connected manifolds of the same dimension.
result Affirmative answer given for both integer and rational settings.
Matching two different sets of items, called heterogeneous set-to-set matching problem, has recently received attention as a promising problem. The difficulties are to extract features to match a correct pair of different sets and also preserve two types of exchangeability required for set-to-set matching: the pair of …
New set-valued star-shaped risk measures introduced for better risk assessment.
problem Improving risk assessment in financial contexts.
method Developed new set-valued star-shaped risk measures and proved their representation theorems.
result Set-valued star-shaped risk measures can be represented as unions of set-valued convex risk measures.
We introduce the concept of hereditarily non uniformly perfect sets, compact sets for which no compact subset is uniformly perfect, and compare them with the following: Hausdorff dimension zero sets, logarithmic capacity zero sets, Lebesgue 2-dimensional measure zero sets, and porous sets. In particular, we give an exa…
Study dynamics and topology of flows near non-saddle sets or W-sets.
problem Understanding the dynamics and topology of flows near specific invariant sets.
method Cohomological relations and global properties analysis.
result Dynamical classification of surfaces and robustness of non-saddle-sets.
The study explores mapping degree sets and their properties for manifolds.
problem Understanding the structure and properties of mapping degree sets for manifolds.
method Analyzes the properties of mapping degree sets and their relationships with self-mapping degree sets.
result Not every multiplicative set containing 0,1 is a self-mapping degree set.
Current approaches for predicting sets from feature vectors ignore the unordered nature of sets and suffer from discontinuity issues as a result. We propose a general model for predicting sets that properly respects the structure of sets and avoids this problem. With a single feature vector as input, we show that our m…
This paper studies the geometry of minimum-volume confidence sets for multinomial parameters.
problem Determining if minimum-volume confidence sets for multinomial outcomes are disjoint.
method Enumerating and covering the continuous regions of the exact p-value function to study the geometry of minimum-volume confidence sets.
result The geometry of minimum-volume confidence sets for multinomial parameters is studied, providing insights into their structure and properties.
Consider a general machine learning setting where the output is a set of labels or sequences. This output set is unordered and its size varies with the input. Whereas multi-label classification methods seem a natural first resort, they are not readily applicable to set-valued outputs because of the growth rate of the o…
Deep Sets approximates functions on sets with high-dimensional latent space.
problem Modeling functions of sets (permutation-invariant functions).
method Deep Sets, a method known to be a universal approximator for continuous set functions.
result Deep Sets' universal approximation property is only guaranteed with a sufficiently high-dimensional latent space.
Study online learning with set-valued feedback, showing differences between deterministic and randomized approaches.
problem Online learning with set-valued feedback, where labels are sets rather than single labels.
method Introduced new combinatorial dimensions (Set Littlestone and Measure Shattering) to characterize learnability.
result Characterized deterministic and randomized online learnability, and established bounds for various learning settings.
A stability-based method selects the most desirable conformal prediction set.
problem Selecting the most desirable conformal prediction set from multiple valid sets invalidates coverage guarantees.
method A stability-based approach that ensures coverage for the selected prediction set.
result The stability-based approach maintains coverage guarantees for the selected prediction set.
The paper links set cuspidality to function regularity and flatness.
problem Linking set cuspidality to function regularity and flatness.
method Analyzes arc-smooth functions and their properties on various sets.
result Establishes a precise link between set cuspidality and function regularity.
This work establishes properties on diffeological structures for set-valued maps and measures.
problem Establish rigorous properties on diffeological structures for set-valued maps and measures.
method Using diffeologies, the authors link various structures including set-valued maps, relations, gradients, measures, and shape analysis.
result Established rigorous properties on sample diffeologies.
Causal Set Theory's Hauptvermutung is resolved in two ways, one of which is true.
problem Formulating and resolving the Hauptvermutung in Causal Set Theory.
method Two mathematically well-defined formulations of the Hauptvermutung, one of which is true.
result The Hauptvermutung is true when finite sets are replaced by countable sets.
This letter introduces an abstract learning problem called the "set embedding": The objective is to map sets into probability distributions so as to lose less information. We relate set union and intersection operations with corresponding interpolations of probability distributions. We also demonstrate a preliminary so…
Find limiting sets for digital cones and suspensions.
problem Digital topology cone and suspension constructions.
method Identify (m, n)-limiting sets, especially (0, 0)-freezing sets.
result Discover (0, 0)-limiting sets for digital cones and suspensions.
Analytic sets with unique infinite tangent cone are algebraic.
problem Characterizing analytic sets with unique infinite tangent cones.
method Analytic and algebraic set properties, degree of complex algebraic sets.
result Degree of Lipschitz normally embedded sets equals their infinite tangent cone degree.
Study freezing sets for digital images in a 2D grid.
problem Determine minimal freezing sets for digital images.
method Prove methods to obtain freezing sets for digital images (X, c_i) where X is a subset of Z^2.
result Examples show how methods can lead to the determination of minimal freezing sets.
Fuzzy prediction sets generalize binary predictions to include elements at varying confidence levels.
problem Binary prediction sets are limited; fuzzy prediction sets offer richer guarantees.
method Generalize prediction sets to fuzzy sets, showing they are e-values with merging properties.
result Optimal e-values lead to optimal fuzzy prediction sets, including optimal conformal prediction.
Representations of sets are challenging to learn because operations on sets should be permutation-invariant. To this end, we propose a Permutation-Optimisation module that learns how to permute a set end-to-end. The permuted set can be further processed to learn a permutation-invariant representation of that set, avoid…
Develops deep neural network techniques for sets as input and output.
problem Bottlenecks in set representation and discontinuity issues in set prediction.
method Techniques for set representation and prediction, addressing unordered nature and relations.
result Improvements in set prediction and representation across various experiments.
Proves a theorem for Assouad dimension with applications to distance sets and radial projections.
problem Problems related to Assouad dimension and distance sets.
method General nonlinear projection theorem for Assouad dimension.
result Sharp estimates for sets with Assouad dimension less than 1 and exceptional set estimates.
The paper explores connections between perimeter, area, and visual angle of convex sets.
problem Understanding geometric properties of convex sets through visual angle and related measurements.
method Establishing universal formulas and characterizing convex sets of constant width.
result Crofton's formula is the unique universal formula relating visual angle, length, and area.
The paper defines cyclic sets from ribbon string links and connects them to quantum invariants.
problem Defining and relating cyclic sets from ribbon string links.
method Endowing ribbon string links with cyclic and cocyclic structures, relating to coend of a ribbon category via quantum invariants.
result Established a relationship between ribbon string links and quantum invariants.
New tools for constructing fixed point sets in digital topology.
problem Constructing fixed point sets in digital topology.
method Defining excludable points and articulation points, and showing their exclusion from freezing sets.
result Excludable points and articulation points can be excluded from all freezing sets.
Set risk measures extend traditional risk measures to handle sets of positions.
problem Handling sets of positions with a single capital requirement.
method Developed an axiomatic framework for set risk measures, dual representation through topology and measures.
result Characterized worst-case set risk measures and provided examples.
Study contractibility of boundaries in convex sets and limit sets of subgroups.
problem Understanding contractibility of boundaries and wildness of limit sets in geometric structures.
method Use sufficient conditions for contractibility, study coarse upper curvature bounds, and analyze interpolation in geodesic metric spaces.
result Conditions for contractibility of boundaries and properties of limit sets are established.