This paper studies poisoning attacks in episodic RL and discovers their effectiveness depends on reward bounds.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New attack manipulates UCB algorithm, new defense algorithm reduces pseudo-regret.
State-only imitation learning improves dexterous manipulation learning from videos.
We propose a tool-use model that can detect the features of tools, target objects, and actions from the provided effects of object manipulation. We construct a model that enables robots to manipulate objects with tools, using infant learning as a concept. To realize this, we train sensory-motor data recorded during a t…
Proposes KStar Diffuser for kinematics-aware bimanual robotic manipulation.
HO2 learns options from data efficiently, improving robot manipulation tasks.
We provide direct evidence of market manipulation at the beginning of the financial crisis in November 2007. The type of manipulation, a "bear raid," would have been prevented by a regulation that was repealed by the Securities and Exchange Commission in July 2007. The regulation, the uptick rule, was designed to preve…
This work tackles force control for contact-rich manipulation tasks with rigid robots using RL.
SAVO actor improves reinforcement learning by avoiding local optima in complex Q-functions.
Robot learns to manipulate objects using multiple geometric representations.
Mastering robotic manipulation skills through reinforcement learning (RL) typically requires the design of shaped reward functions. Recent developments in this area have demonstrated that using sparse rewards, i.e. rewarding the agent only when the task has been successfully completed, can lead to better policies. Howe…
Paper detects pump and dump schemes in cryptocurrencies.
Goal-directed manipulation of representations is a key element of human flexible behaviour, while consciousness is often related to several aspects of higher-order cognition and human flexibility. Currently these two phenomena are only partially integrated (e.g., see Neurorepresentationalism) and this (a) limits our un…
When consequential decisions are informed by algorithmic input, individuals may feel compelled to alter their behavior in order to gain a system's approval. Models of agent responsiveness, termed "strategic manipulation," analyze the interaction between a learner and agents in a world where all agents are equally able …
Hierarchical reinforcement learning (HRL) has recently shown promising advances on speeding up learning, improving the exploration, and discovering intertask transferable skills. Most recent works focus on HRL with two levels, i.e., a master policy manipulates subpolicies, which in turn manipulate primitive actions. Ho…
We study finite-dimensional integrals in a way that elucidates the mathematical meaning behind the formal manipulations of path integrals occurring in quantum field theory. This involves a proper understanding of how Wick's theorem allows one to evaluate integrals perturbatively, i.e., as a series expansion in a formal…
Researchers create a flickering attack to fool video recognition networks.
For a social networking service to acquire and retain users, it must find ways to keep them engaged. By accurately gauging their preferences, it is able to serve them with the subset of available content that maximises revenue for the site. Without the constraints of an appropriate regulatory framework, we argue that a…
We study adversarial attacks that manipulate the reward signals to control the actions chosen by a stochastic multi-armed bandit algorithm. We propose the first attack against two popular bandit algorithms: -greedy and UCB, \emph{without} knowledge of the mean rewards. The attacker is able to spend only logarithmic …
New method learns robot actions from videos without explicit labels.
Learning predictive models from interaction with the world allows an agent, such as a robot, to learn about how the world works, and then use this learned model to plan coordinated sequences of actions to bring about desired outcomes. However, learning a model that captures the dynamics of complex skills represents a m…
AI learns market manipulation through simulation, suggesting regulation.
Collectives can manipulate learning platforms by coordinated data submission, requiring strategic assessments and algorithms.
Survey examines challenges and solutions in sim-to-real transfer for robotics.
Robotic systems are ever more capable of automation and fulfilment of complex tasks, particularly with reliance on recent advances in intelligent systems, deep learning and artificial intelligence. However, as robots and humans come closer in their interactions, the matter of interpretability, or explainability of robo…
The paper analyzes how leverage affects manipulation in event-linked markets, offering new insights into regulation.
New model improves neural network robustness against input manipulations.
Manipulating video content is easier than ever. Due to the misuse potential of manipulated content, multiple detection techniques that analyze the pixel data from the videos have been proposed. However, clever manipulators should also carefully forge the metadata and auxiliary header information, which is harder to do …
We adapt the ideas underlying the success of Deep Q-Learning to the continuous action domain. We present an actor-critic, model-free algorithm based on the deterministic policy gradient that can operate over continuous action spaces. Using the same learning algorithm, network architecture and hyper-parameters, our algo…
Modeling option market making with hedging-induced price impact.
Policy gradient methods have enjoyed great success in deep reinforcement learning but suffer from high variance of gradient estimates. The high variance problem is particularly exasperated in problems with long horizons or high-dimensional action spaces. To mitigate this issue, we derive a bias-free action-dependent ba…
Q-chunking improves RL for long tasks by chunking actions.
In this work we propose a model that can manipulate individual visual attributes of objects in a real scene using examples of how respective attribute manipulations affect the output of a simulation. As an example, we train our model to manipulate the expression of a human face using nonphotorealistic 3D renders of a f…
Prediction markets can be manipulated by traders who can move contract settlements, harming price discovery.
Most reinforcement learning algorithms are inefficient for learning multiple tasks in complex robotic systems, where different tasks share a set of actions. In such environments a compound policy may be learnt with shared neural network parameters, which performs multiple tasks concurrently. However such compound polic…
DIGIT is a low-cost tactile sensor for in-hand manipulation.
Predictive State Representations (PSRs) are an expressive class of models for controlled stochastic processes. PSRs represent state as a set of predictions of future observable events. Because PSRs are defined entirely in terms of observable data, statistically consistent estimates of PSR parameters can be learned effi…
Volunteer labor can temporarily yield lower benefits to charities than its costs. In such instances, organizations may wish to defer volunteer donations to a later date. Exploiting a discontinuity in blood donations' eligibility criteria, we show that deferring donors reduces their future volunteerism. In our setting, …
Manipulating data, such as weighting data examples or augmenting with new instances, has been increasingly used to improve model training. Previous work has studied various rule- or learning-based approaches designed for specific types of data manipulation. In this work, we propose a new method that supports learning d…
Paper presents a new port-Hamiltonian model for vehicle manipulators.
We envision that in the near future, humanoid robots would share home space and assist us in our daily and routine activities through object manipulations. One of the fundamental technologies that need to be developed for robots is to enable them to detect objects and recognize them for effective manipulations and take…
Optimal benchmark design varies based on costs in financial manipulation.
Study examines reasons for Nutek India's share price drop.
New method uses statistical physics to detect financial market manipulation.
KINet learns object interactions without supervision for robotic pushing.
Market manipulation is a strategy used by traders to alter the price of financial securities. One type of manipulation is based on the process of buying or selling assets by using several trading strategies, among them spoofing is a popular strategy and is considered illegal by market regulators. Some promising tools h…
We study a generalized setup for learning from demonstration to build an agent that can manipulate novel objects in unseen scenarios by looking at only a single video of human demonstration from a third-person perspective. To accomplish this goal, our agent should not only learn to understand the intent of the demonstr…
New foundation for Shapley value immune to coalitional manipulations.