Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,932 papers · 148 categories

Trend · papers per month

12.5%25.0%37.5%50.0% · Sep 199319922001200920172026
48 results for changing action set

RL agent learns to smoothly change lanes in a dynamic driving environment.

problem Challenging lane change control with safety and comfort.
method Formulated continuous action for lane change in DDPG algorithm, defined reward function for learning.
result Successfully changed lanes with 100% success rate in diverse driving situations.

New action poisoning attacks improve LinUCB's performance by changing action signals.

problem Improving understanding of adversarial attacks on contextual bandit algorithms.
method Proposed action poisoning attacks in white-box and black-box settings.
result Action poisoning attacks can force LinUCB to pull a target arm frequently with low cost.

Let ρ0ρ_0 be an action of a Lie group on a manifold with boundary that is transitive on the interior. We study the set of actions that are topologically conjugate to ρ0ρ_0, up to smooth or analytic change of coordinates. We show that in many cases, including the compactifications of negatively curved symmetric spaces, …

2008-04-15abs ↗pdf ↗

We introduce a new online learning framework where, at each trial, the learner is required to select a subset of actions from a given known action set. Each action is associated with an energy value, a reward and a cost. The sum of the energies of the actions selected cannot exceed a given energy budget. The goal is to…

2018-10-28abs ↗pdf ↗

Algorithm optimizes bandit decisions with changing action sets using Gaussian processes.

problem Optimizing decisions in a bandit problem with time-varying action sets.
method Proposes an algorithm called O'CLOK-UCB using Gaussian processes to handle changing action sets and contexts.
result Achieves regret bound of ildeO(λ(K)KTγKT(tTXt)) ilde{O}(\sqrt{λ^*(K)KTγ_{KT}(\cup_{t\leq T}\mathcal{X}_t)} ) with high probability.

Efficient algorithms for online learning with changing action sets, achieving no-approximate-regret guarantees.

problem Online learning with sleeping experts/bandits, where only a subset of actions are available each time.
method Developed computationally efficient algorithms providing no-approximate-regret guarantees for the general problem and better approximation ratios for special cases.
result Achieved no-approximate-regret guarantees for the general sleeping expert/bandit problems and better approximation ratios for specific cases.

New algorithm handles MDPs with unknown, changing rewards efficiently.

problem Handling MDPs with unknown, changing rewards in large state spaces.
method Developed an algorithm with O(τ(lnS+lnA)Tln(T))O(\sqrt{τ(\ln|S|+\ln|A|)T}\ln(T)) regret bound and a modified algorithm with polynomial complexity.
result Achieved state-of-the-art regret bounds for large scale MDPs with changing rewards.

Causal Bayesian networks interpret actions as interventions to connect models to real-world outcomes.

problem Connecting causal model predictions to real-world outcomes.
method Formal framework to interpret actions as interventions and prove impossibility results.
result No non-circular interpretation exists that satisfies natural desiderata without violating some.

Inverse classification is the process of perturbing an instance in a meaningful way such that it is more likely to conform to a specific class. Historical methods that address such a problem are often framed to leverage only a single classifier, or specific set of classifiers. These works are often accompanied by naive…

2016-10-05abs ↗pdf ↗

PRINCE provides interpretable explanations for recommender systems by removing minimal user actions.

problem Lack of interpretable explanations for recommender systems.
method PRINCE uses a polynomial-time optimal algorithm based on random walks over dynamic graphs to find minimal user actions that change recommendations.
result PRINCE produces more compact explanations than intuitive baselines and is viable for user understanding.

This work introduces a method to attribute model performance drops to distribution shifts.

problem Attributing performance drops of machine learning models to distribution shifts.
method Formulated as a cooperative game, value of a set of distributions is defined as the change in model performance when only that set of distributions changes. Importance weighting method for computing the value of an arbitrary set of distributions is derived. Quantifying the contribution of each distribution as its Shapley value.
result Demonstrated the effectiveness of the method on various case studies.

A successful response to climate change needs vast investments in low-carbon research, energy, and sustainable development. Governments can drive research, provide environmental regulation, and accelerate global development, but the necessary low-carbon investments of 2-3% GDP have yet to materialise. A new strategy to…

2018-07-09abs ↗pdf ↗

The paper develops a method to create robust control policies for robots using information bottlenecks.

problem Robotic control policies are sensitive to task-irrelevant state and sensor changes.
method Derives a policy gradient algorithm that creates an information bottleneck between states and task-relevant representations.
result Task-driven policies are more robust to sensor noise and environmental changes.

We study Cohen-Macaulay actions, a class of torus actions on manifolds, possibly without fixed points, which generalizes and has analogous properties as equivariantly formal actions. Their equivariant cohomology algebras are computable in the sense that a Chang-Skjelbred Lemma, and its stronger version, the exactness o…

2009-12-03abs ↗pdf ↗

We propose algorithms for online principal component analysis (PCA) and variance minimization for adaptive settings. Previous literature has focused on upper bounding the static adversarial regret, whose comparator is the optimal fixed action in hindsight. However, static regret is not an appropriate metric when the un…

2019-01-23abs ↗pdf ↗

Safe imitation learning with a safety layer for flexible training.

problem Flexible yet safe imitation learning for complex tasks.
method Theory and modular method with a safety layer for continuous policy, adversarial training, and worst-case safety guarantees.
result Robustness advantage of safety layer during training compared to test time.

Study identifies change points in piecewise constant reward functions with fixed exploration budget.

problem Locating abrupt changes in piecewise constant reward functions under bandit feedback.
method Fixed exploration budget, piecewise constant bandit problem, lower bounds, near optimal algorithms.
result Established lower bounds and near matching upper bounds for both small and large budgets.

Recourse explanations can become invalid if collective actions change statistical data.

problem Recourse explanations may become invalid due to collective behavior changing data statistics.
method Formal characterization of conditions under which recourse explanations remain valid under performativity.
result Recourse actions may become invalid if they are influenced by or intervene on non-causal variables.

Linear contextual bandit is an important class of sequential decision making problems with a wide range of applications to recommender systems, online advertising, healthcare, and many other machine learning related tasks. While there is a lot of prior research, tight regret bounds of linear contextual bandit with infi…

2019-05-04abs ↗pdf ↗

New framework guides resource usage to achieve sublinear regret in adversarial settings.

problem Achieving sublinear regret in online decision making with changing reward and cost distributions.
method General primal-dual methods guided by spending plans that ensure balanced resource usage.
result Achieves sublinear regret with respect to spending plans that balance resource usage.

Improved algorithm detects changes in RL environments with non-stationary MDPs.

problem Learning in non-stationary reinforcement learning environments.
method R-BOCPD-UCRL2 algorithm for MDPs with multinomial state transitions.
result Near-optimal theoretical guarantees in terms of false-alarm rate and detection delay.

A group action is called polar if there exists an immersed submanifold (a section) which intersects all orbits orthogonally. Such group actions have been studied extensively on symmetric spaces. We show how to construct a manifold admitting a polar group action by prescribing their isotropy groups along a fundamental d…

2012-08-05abs ↗pdf ↗

The paper tackles energy management in buildings with PCM using dynamic programming.

problem Optimal scheduling of HVAC systems in buildings with PCM is challenging due to nonlinear and non-convex characteristics.
method The paper uses dynamic programming to address the nonlinear nature of PCM, incorporating macro actions and multi-time scale Markov decision processes to reduce computational burden.
result The proposed method demonstrates a computational speed-up of up to 12,900 times compared to direct DP application.

Proposes a topological model for partial equivariance in neural networks.

problem Capturing partial equivariance in neural networks for data analysis.
method Introduces P-GENEOs and studies spaces of measurements and P-GENEOs between them.
result Spaces of measurements and P-GENEOs have convenient approximation and convexity properties.

Study variations of Riemannian submersions to maintain geodesic fibers and positive curvatures.

problem Maintain geodesic fibers and positive sectional curvatures in Riemannian submersions.
method Vary Riemannian metrics while keeping fibers totally geodesic and horizontal distribution fixed.
result Conditions for making sectional curvatures positive and existence of fat submersions.

We achieve a finite regret bound of O(dlogd) for online inverse linear optimization with M-convex action sets.

problem Online inverse linear optimization with M-convex action sets.
method Combining structural characterization of optimal solutions on M-convex sets with geometric volume argument.
result Finite regret bound of O(dlogd) for online inverse linear optimization with M-convex action sets.

Proposes an efficient method for ordered counterfactual explanations.

problem Insufficient explanation of perturbation vectors for executing actions.
method Mixed-Integer Linear Optimization (MILP) approach for evaluating and extracting optimal pairs of actions and orders.
result Demonstrated effectiveness of the proposed method on real datasets.

Differential invariants of a (pseudo)group action can vary when restricted to invariant submanifolds (differential equations). The algebra is still governed by the Lie-Tresse theorem, but may change a lot. We describe in details the case of the motion group O(n)RnO(n)\ltimes\R^n acting on the full (unconstraint) jet-space …

2007-12-20abs ↗pdf ↗

The main result of this paper asserts that if a Seifert fibered 4-manifold has nonzero Seiberg-Witten invariant, the homotopy class of regular fibers has infinite order. This is a nontrivial obstruction to smooth circle actions; as applications, we show how to destroy smooth circle actions on a 4-manifold by knot surge…

2011-03-29abs ↗pdf ↗

KeRNS tackles non-stationary reinforcement learning in metric spaces.

problem Non-stationary reinforcement learning in metric spaces.
method KeRNS uses time-dependent kernels to model non-stationary Markov Decision Processes (MDPs).
result KeRNS achieves a regret bound that scales with the covering dimension and total variation of the MDP.

One important effect of price shocks in the United States has been increased political attention paid to the structure and performance of oil and natural gas markets, along with some governmental support for energy conservation. This paper describes how price changes helped lead the emergence of a political agenda acco…

2015-02-25abs ↗pdf ↗