We introduce a hierarchical architecture for video understanding that exploits the structure of real world actions by capturing targets at different levels of granularity. We design the model such that it first learns simpler coarse-grained tasks, and then moves on to learn more fine-grained targets. The model is train…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Hierarchical Modular Reinforcement Learning (HMRL), consists of 2 layered learning where Profit Sharing works to plan a prey position in the higher layer and Q-learning method trains the state-actions to the target in the lower layer. In this paper, we expanded HMRL to multi-target problem to take the distance between …
Generically learns movement control policies from exploration data.
We study a formulation of the standard Poisson sigma model in which the target space Poisson manifold carries the Hamilton action of some finite dimensional Lie algebra. We show that the structure of the action and the properties of the gauge invariant observables can be understood in terms of the associated target spa…
Optimizes predictions for specific tasks using parametrized decision analysis.
New approach to meaningful and robust algorithmic recourse.
Study Liouville action for harmonic maps between Riemann surfaces.
BiHermitian geometry, discovered long ago by Gates, Hull and Roceck, is the most general sigma model target space geometry allowing for (2,2) world sheet supersymmetry. By using the twisting procedure proposed by Kapustin and Li, we work out the type A and B topological sigma models for a general biHermtian target spac…
Study designs logging policies to minimize off-policy evaluation error.
Generalizes skyrmion theory to gauged maps with -action.
Policy-gradient method controls multiple non-cohesive targets.
Deep RL learns to construct objects from 2D images by avoiding brick overlaps.
The goal of task transfer in reinforcement learning is migrating the action policy of an agent to the target task from the source task. Given their successes on robotic action planning, current methods mostly rely on two requirements: exactly-relevant expert demonstrations or the explicitly-coded cost function on targe…
The simplest orientifolds of the WZW models are obtained by gauging a Z_2 symmetry group generated by a combined involution of the target Lie group G and of the worldsheet. The action of the involution on the target is by a twisted inversion g \mapsto (ζg)^{-1}, where ζis an element of the center of G. It reverses the …
UWM-JEPA predicts future scenarios in belief space, improving accuracy in partially observed environments.
We study a sigma-model with target space the flag manifold U(3)/U(1)^3. A peculiarity of the model is that the complex structure on the target space enters explicitly in the action. We describe the classical solutions of the model for the case when the worldsheet is a sphere CP^1.
This paper studies poisoning attacks in episodic RL and discovers their effectiveness depends on reward bounds.
Transfer reinforcement learning (RL) aims at improving the learning efficiency of an agent by exploiting knowledge from other source agents trained on relevant tasks. However, it remains challenging to transfer knowledge between different environmental dynamics without having access to the source environments. In this …
SHIFT framework identifies subgroups with large ML model performance decay.
Zamolodchikov's c-theorem type argument (and also string theory effective action constructions) imply that the RG flow in 2d sigma model should be gradient one to all loop orders. However, the monotonicity of the flow of the target-space metric is not obvious since the metric on the space of metric-dilaton couplings is…
We extend the stochastic Perron method to analyze the framework of stochastic target games, in which one player tries to find a strategy such that the state process almost surely reaches a given target no matter which action is chosen by the other player. Within this framework, our method produces a viscosity sub-solut…
Improves AI agents' 3D navigation by learning from failures and 3D spatial relationships.
Off-shell supermultiplets in 2-dimensions are formulated. These are used to construct sigma models whose target spaces are vector bundles over manifolds that are hyperkähler with torsion. The off-shell supersymmetry implies that the complex structures are simultaneously integrable and allows us to write actions…
In this paper, we propose an active perception method for recognizing object categories based on the multimodal hierarchical Dirichlet process (MHDP). The MHDP enables a robot to form object categories using multimodal information, e.g., visual, auditory, and haptic information, which can be observed by performing acti…
It is well established that humans decision making and instrumental control uses multiple systems, some which use habitual action selection and some which require deliberate planning. Deliberate planning systems use predictions of action-outcomes using an internal model of the agent's environment, while habitual action…
A new pricing controller handles resource constraints to infer target prices effectively.
CNT leverages noisy targets to guide model learning.
BiKaehler geometry is characterized by a Riemannian metric g_{ab} and two covariantly constant generally non commuting complex structures K_+^a_b, K_-^a_b, with respect to which g_{ab} is Hermitian. It is a particular case of the biHermitian geometry of Gates, Hull and Roceck, the most general sigma model target space …
Defines smooth actions of a group on manifolds and vector spaces.
New action poisoning attacks improve LinUCB's performance by changing action signals.
A new method for reinforcement learning that adapts to different domains using auxiliary classifiers.
Inter-Cell Interference Coordination (ICIC) is a promising way to improve energy efficiency in wireless networks, especially where small base stations are densely deployed. However, traditional optimization based ICIC schemes suffer from severe performance degradation with complex interference pattern. To address this …
We announce ultrametric analogues of the results of Kleinbock-Margulis for shrinking target properties of semisimple group actions on symmetric spaces. The main applications are S-arithmetic Diophantine approximation results and logarithm laws for buildings, generalizing the work of Hersonsky-Paulin on trees.
We study a stochastic game where one player tries to find a strategy such that the state process reaches a target of controlled-loss-type, no matter which action is chosen by the other player. We provide, in a general setup, a relaxed geometric dynamic programming principle for this problem and derive, for the case of …
MAXMINLCB optimizes unknown target functions with preference feedback using a Stackelberg game approach.
For a particular class of backgrounds, equations of motion for string sigma models targeted in mutually dual Poisson-Lie groups are equivalent. This phenomenon is called the Poisson-Lie T-duality. On the level of the corresponding string effective actions, the situation becomes more complicated due to the presence of t…
MAGE optimizes policies using action gradients from model-based learning.
We propose a tool-use model that can detect the features of tools, target objects, and actions from the provided effects of object manipulation. We construct a model that enables robots to manipulate objects with tools, using infant learning as a concept. To realize this, we train sensory-motor data recorded during a t…
We study the propagation of bosonic strings in singular target space-times. For describing this, we assume this target space to be the quotient of a smooth manifold by a singular foliation on it. Using the technical tool of a gauge theory, we propose a smooth functional for this scenario, such that the p…
Paper develops IV method for consistent OPE in confounded MDPs.
Study examines CSO algorithm for 3D swarming and tracking multiple targets.
Deep reinforcement learning agents have recently been successful across a variety of discrete and continuous control tasks; however, they can be slow to train and require a large number of interactions with the environment to learn a suitable policy. This is borne out by the fact that a reinforcement learning agent has…
We propose a targeted communication architecture for multi-agent reinforcement learning, where agents learn both what messages to send and whom to address them to while performing cooperative tasks in partially-observable environments. This targeting behavior is learnt solely from downstream task-specific reward withou…
FPGs use structure to improve policy learning in complex tasks.
Paper tackles RL with continuous actions and unmeasured confounders.
The Poisson--Weil sigma model, worked out by us recently, stems from gauging a Hamiltonian Lie group symmetry of the target space of the Poisson sigma model. Upon gauge fixing of the BV master action, it yields interesting topological field theories such as the 2--dimensional Donaldson-Witten topological gauge theory a…
Improved sample complexity for target Q-learning in finite MDPs with generative oracle.
DFL framework improves action and outcome fairness in policy learning.