Study lenient regret and good-action identification in Gaussian process bandits.
problem Optimizing function values above a certain threshold in Gaussian process bandits.
method Study lenient regret notions and introduce algorithms for finding good actions.
result Upper and lower bounds on lenient regret for GP-UCB and elimination algorithms.
Paper improves action selection for accurate parameter estimation in linear bandits.
problem Best action identification in stochastic linear bandits with fixed confidence constraints.
method Designs a sequential adaptive policy to estimate underlying parameter efficiently.
result The designed policy achieves the same estimation error scaling as a lower bound.
New method for identifying best action in strategic games.
problem Identifying the best action in a game with random outcomes.
method Two strategies: Maximin-LUCB and Maximin-Racing.
result Both methods achieve good sample complexity and performance.
New algorithm identifies optimal subtrees in fixed-budget tree search.
problem Identifying optimal subtrees in fixed-budget Monte Carlo Tree Search.
method ε-agnostic algorithm for max-min action identification.
result Misidentification probability decays exponentially with sample size.
Active learning method estimates nonlinear systems efficiently.
problem Identifying nonlinear dynamical systems with continuous states and actions.
method Repeating three steps: trajectory planning, tracking, and re-estimation.
result Estimates nonlinear dynamical systems at a parametric rate.
Study robust best-arm identification in linear bandits with lower bounds and algorithms.
problem Identify a near-optimal robust arm in linear bandits with adversarial actions.
method Propose instance-dependent lower bounds and both static and adaptive bandit algorithms.
result Sample complexity matches the lower bound and algorithms effectively identify robust arms.
Simplifies causal effect identification by pruning unnecessary variables.
problem Complex expressions from causal models can lead to unnecessary variables and computational burden.
method Graphical criteria and improved identifiability algorithm to detect and prune redundant variables.
result Improved computational efficiency and reduced bias in estimating causal effects.
Paper tackles RL with continuous actions and unmeasured confounders.
problem Offline policy learning with continuous actions and unmeasured confounders.
method Developed a novel identification result and a minimax estimator for nonparametric policy value estimation.
result Introduced a policy-gradient-based algorithm to identify the optimal policy.
The Farrell-Jones Conjecture holds for groups acting acylindrically on trees.
problem Verifying the Farrell-Jones Conjecture for groups acting on trees.
method Analyzing acylindrical actions on simplicial trees and using the Farrell-Jones Conjecture.
result The Farrell-Jones Conjecture holds for groups acting acylindrically on trees.
The action of a Lie pseudogroup G on a smooth manifold M induces a prolonged pseudogroup action on the jet spaces Jn of submanifolds of M. We prove in this paper that both the local and global freeness of the action of G on Jn persist under prolongation in the jet order n. Our results underlie the const…
New algorithm tackles nonstationary linear bandits with latent dynamics.
problem Nonstationary bandit problem with latent states and unknown dynamics.
method Explore-then-commit algorithm with exploration and commitment phases.
result Achieves ildeO(T2/3) regret. The paper introduces metrics to rank potential outcomes for better decision-making.
problem Optimal action selection in uncertain situations using causal reasoning.
method Introducing two new metrics: probabilities of potential outcome ranking (PoR) and probability of achieving the best potential outcome (PoB). Establishing identification theorems and deriving bounds for these metrics, and presenting estimation methods.
result The estimators' finite-sample properties and their application to a real-world dataset are demonstrated.
Study best arm identification in restless bandits with unknown TPMs.
problem Identify the best arm with fixed confidence in restless bandits with unknown TPMs.
method Proposed a policy for best arm identification and proved its expected stopping time matches the lower bound.
result The state-action visitation proportions match the optimal proportions under any asymptotically optimal policy.
A new algorithm for identifying the best arm in linear feedback with safety constraints.
problem Identifying the best arm in linear feedback with safety constraints.
method A gap-based algorithm that ensures safety while minimizing sample complexity.
result The algorithm achieves meaningful sample complexity while ensuring safety.
Paper tackles best arm identification with cost consideration.
problem Best arm identification with cost consideration in product development.
method Derives a theoretical lower bound and proposes algorithms CTAS and CO.
result Simple algorithms can deliver near-optimal performance.
The paper provides a non-asymptotic error bound for linear system identification under nonlinear policies.
problem System identification for linear systems with nonlinear and/or time-varying policies under i.i.d. random excitation noises.
method Least square estimation with non-asymptotic error bound for bounded state and action trajectories.
result The error bound is consistent with linear policies and generalizes existing guarantees.
Paper uses sparse learning to estimate quasi-potential and drift components in stochastic systems.
problem Estimating quasi-potential and drift components in stochastic systems.
method Sparse identification of non-linear dynamics (SINDy) combined with action minimization methods.
result Evaluation of quasi-potential landscape from a single trajectory.
Paper presents a method for recognizing human actions using GLAC features from motion and static images.
problem Action recognition in 3D depth videos.
method 3D Motion Trail Model (3DMTM) for MHIs and SHIs, GLAC features extraction, l2-regularized Collaborative Representation Classifier (l2-CRC) for classification.
result The method outperforms other approaches in recognizing human actions.
Algorithm identifies best arm in combinatorial bandits with semi-bandit feedback.
problem Identifying the best arm in combinatorial bandits with semi-bandit feedback.
method Interpreted as a sequential zero-sum game, developed a CombGame meta-algorithm with finite time guarantees.
result First computationally efficient algorithm that is asymptotically optimal and has competitive empirical performance.
New RL method handles hidden actions in offline learning.
problem Learning from unseen actions in real-world RL datasets.
method LURE (Learning from the Unseen: Robust Estimator) method using next-state variable as proxy.
result Valid statistical inference and improved RL conclusions with hidden actions.
Bootstrap policies improve regret in continuous state-action reinforcement learning.
problem Improving regret in reinforcement learning for continuous state and action spaces.
method Bootstrap-based policies for stochastic linear systems with quadratic cost functions.
result Bootstrap policies achieve a square root scaling of regret with respect to time.
New RLHF algorithm identifies optimal policies from human feedback without explicit reward inference.
problem Training large language models with human feedback without reward inference.
method Model-free RLHF algorithm BSAD that identifies optimal policies directly from human preference. result Provable, instance-dependent sample complexity ildeO(cMSA3H3Mlogδ1). New algorithms identify best policies in discounted linear MDPs efficiently.
problem Identifying the best policy in discounted linear MDPs with limited samples.
method Derive lower bounds and devise simple yet near-optimal algorithms.
result Upper bound on sample complexity matches existing bounds.
The paper introduces pseudo-quotients for algebraic actions and applies them to character varieties.
problem Characterizing algebraic actions and their quotients.
method Introducing pseudo-quotients as a weak version of quotients for algebraic actions, focusing on purely topological properties.
result Pseudo-quotients are unique up to virtual class in characteristic zero and can be used to compute character varieties.
Discussing moving frames for curve and surface invariants.
problem Identifying differential invariants of curves and surfaces.
method Using moving frames for Euclidean, affine, and conformal transformations.
result Determine differential invariants of curves and surfaces.
Improved sample and time complexity for identifying mixtures of product distributions.
problem Identifying a mixture of k product distributions from statistics. method Combining robust tensor decomposition and Hadamard extensions to bound the condition number of key matrices.
result Achieved sample complexity and run-time complexity of (1/ζ)O(k) for n≥2k−1. MGpi model predicts social actions in group interactions.
problem Social interaction among multiple agents and groups.
method Deep neural network with Kinesic-Proxemic-Message gate for social signal gating.
result Achieves state-of-the-art performance in group identification.
Identifies spectral curves for SU(3) coadjoint orbits.
problem Understanding the geometry of coadjoint orbits in SU(3).
method Using Hitchin pairs and spectral curves, identifies a Hamiltonian circle action and finds Darboux coordinates.
result Identifies a differential equation for the Hamiltonian.
New method identifies latent variables with sparse perturbations.
problem Identifying latent variables with minimal supervision.
method Weakly supervised representation learning with sparse perturbations.
result Identification of latent variables up to specified blocks.
Groups acting on bifoliated planes are left-orderable.
problem Characterizing groups acting on bifoliated planes.
method Construction of a linear order on leaf space ends and identification with boundary circle subsets.
result Groups acting on bifoliated planes are left-orderable.
Researchers create a framework to value player actions in CSGO.
problem Lack of accessible data and analytical frameworks for esports players.
method Data model, graph distance measure, context-aware framework.
result Demonstrated framework's consistency and independence compared to existing methods.
Predictive State Representations (PSRs) are an expressive class of models for controlled stochastic processes. PSRs represent state as a set of predictions of future observable events. Because PSRs are defined entirely in terms of observable data, statistically consistent estimates of PSR parameters can be learned effi…
The paper proves a pseudo-Kähler structure on a torus's projective space.
problem Existence of a pseudo-Kähler structure on a torus's projective space.
method Proved the existence of a pseudo-Kähler structure using complex, symplectic, and Riemannian compatibility.
result Existence of a moment map for the SL(2, R) action over the deformation space.
The Knowledge Gradient policy is improved for MABs by avoiding dominated actions.
problem Weaknesses in KG policy for MABs, including taking dominated actions.
method Proposed variants of KG that avoid taking dominated actions, including an index heuristic.
result New policies perform well over a range of MABs, including those for which index policies are not optimal.
Differentiable MPC improves reinforcement learning efficiency.
problem Combining model-free and model-based reinforcement learning approaches.
method Differentiating through KKT conditions of a convex approximation of MPC.
result MPC policies are more data-efficient and superior to traditional system identification.
A new pricing controller handles resource constraints to infer target prices effectively.
problem Resource constraints prevent fixed-price inference, leading to support exclusion.
method Formalizes support-exclusion failure, designs a target-aware controller, and uses a realized information clock.
result The controller can certify feasible target bands and log continuous local densities, leading to polynomial rates of inference.
Renormalized circle diffeos with breaks converge to Moebius maps with a symplectic structure.
problem Analyzing the renormalization of circle diffeomorphisms with breaks.
method Proving convergence to invariant piecewise Moebius maps and identifying the renormalization operator with a sub-action of the mapping class group.
result Renormalization identifies with a symplectic form preserved by the mapping class group.
Study identifies change points in piecewise constant reward functions with fixed exploration budget.
problem Locating abrupt changes in piecewise constant reward functions under bandit feedback.
method Fixed exploration budget, piecewise constant bandit problem, lower bounds, near optimal algorithms.
result Established lower bounds and near matching upper bounds for both small and large budgets.
Study para-hyperKähler geometry of anti-de Sitter structures.
problem Understand the geometry of anti-de Sitter structures.
method Investigate para-hyperKähler structures and their relations.
result Found neutral pseudo-Riemannian metric and symplectic structures.
Study quantile reward identification with 1-bit feedback constraints.
problem Best arm identification with quantile reward and 1-bit communication.
method Proposes an algorithm using noisy binary search for quantile reward estimation.
result Derives upper and lower bounds on sample complexity for 1-bit feedback.
We use statistically validated networks, a recently introduced method to validate links in a bipartite system, to identify clusters of investors trading in a financial market. Specifically, we investigate a special database allowing to track the trading activity of individual investors of the stock Nokia. We find that …
Paper explores using EEG for better speaker identification, even in noisy environments.
problem Speaker identification performance degrades in background noise.
method Uses EEG signals to enhance speaker identification systems, comparing with acoustic features.
result Speaker identification system using only EEG features outperforms one using only acoustic features in high background noise.
IDS integrates physics engines into deep learning for efficient, interpretable system identification.
problem Lack of generalization and interpretability in learning-based models of physical systems.
method Interactive Differentiable Simulation (IDS) that allows efficient, accurate inference of physical properties.
result Automatic task-based robot design and parameter estimation for nonlinear dynamical systems.
Noiseless IO bounds inferred from demonstrations, matching adversarial settings.
problem Inferring decision-maker's objective from observed data.
method High-probability generalization bounds for induced action set.
result Generalization bound of O(Td) for consistent estimators. New method identifies critical states to improve RL agent explainability and speed.
problem Challenges in RL agent explainability and action selection timing.
method Identify critical states based on action-based variance in Q-function, prioritize exploitation on these states.
result Identified critical states accelerate RL in grid worlds and deep RL tasks.
StakeBench evaluates language understanding by linking comments to market commitments, improving model alignment with real-world outcomes.
problem Existing financial NLP benchmarks measure perceived language rather than market commitments.
method StakeBench uses observable market behavior to supervise models, testing their ability to detect commitments, identify sides, and project odds.
result Models partially recover position-side signals but struggle with later tasks, highlighting structural failures.
Optimization geometrodynamics simplifies adaptive optimizer dynamics.
problem Hidden states in adaptive optimizers complicate gradient-based learning.
method Develops a variational theory to eliminate hidden states and compose across hierarchies.
result Yields interaction curvature that integrates to finite contrasts.
CoCAI uses copulas for accurate multivariate time-series forecasting and anomaly detection.
problem Accurate multivariate time-series forecasting and robust anomaly detection.
method Copula-based conformal prediction for multivariate time-series analysis.
result CoCAI provides statistically valid predictive regions and robust anomaly scores.