The paper uses machine learning to optimize rework policies in semiconductor manufacturing.
problem Optimizing rework steps to increase yield without increasing costs.
method Applied double/debiased machine learning (DML) to estimate treatment effects.
result Derived optimal rework policies and estimated their value empirically.
The paper uses causal machine learning to optimize rework decisions in manufacturing.
problem Optimizing rework policies in manufacturing systems to balance yield improvement and rework costs.
method Proposes a causal model using double/debiased machine learning (DML) techniques to estimate conditional treatment effects and derive rework policies.
result Achieved a yield improvement of 2-3% during the color-conversion process of white LEDs.
These are course notes I wrote for my Fall 2013 graduate topics course on geometric structures, taught at ICERM. The notes rework many of proofs in William P. Thurston's beautiful but hard-to-understand paper, "Shapes of Polyhedra". A number of people, both in and out of the class, found these notes very useful and so …
We show how the theory of tangles is equivalent to that of well-connected tangles. These are drawn on a surface with boundary, and equivalent via Reidemeister moves of a restricted kind. This reworking of the graphical foundations for link and tangle theory can be expected to have a variety of applications, including o…
We demonstrate a conditional autoregressive pipeline for efficient music recomposition, based on methods presented in van den Oord et al.(2017). Recomposition (Casal & Casey, 2010) focuses on reworking existing musical pieces, adhering to structure at a high level while also re-imagining other aspects of the work. This…
This is the second of three papers that refine and extend portions of our earlier preprint, "The depth of a knot tunnel." Together, they rework the entire preprint. The theory of tunnel number 1 knots that we introduced in "The tree of knot tunnels" yields a parameterization in which each tunnel is described uniquely b…
In geometric group theory one uses group actions on spaces to gain information about groups. One natural space to use is the Cayley graph of a group. The Cayley graph arguments that one encounters tend to require local finiteness, and hence finite generation of the group. In this paper, I take the theory of intersectio…
This paper means to correct an error by the authors for the composite q case in the paper "Lens Spaces, Isospectral on Forms but not on Functions", published in LMS J. Comput. Math.} 9 (2006), 270-286. All calculations and examples presented in \cite{GM} for prime q remain valid, and we include detailed calculation…
This is the first of three papers that refine and extend portions of our earlier preprint, "Depth of a knot tunnel." Together, they rework the entire preprint. H. Goda, M. Scharlemann, and A. Thompson described a general construction of all tunnels of all tunnel number 1 knots using "tunnel moves". We apply the theory …
The monoids of simplicial endomorphisms, i.e. the monoids of endomorphisms in the simplicial category, are submonoids of monoids one finds in Temperley-Lieb algebras, and as the monoids of Temperley-Lieb algebras are linked to situations where an endofunctor is adjoint to itself, so the monoids of simplicial endomorphi…
NIFTy.re accelerates imaging models and expands Gaussian processes and variational inference.
problem Slow performance and limited inference strategies in NIFTy.
method Rewritten NIFTy with new modeling principles, inference strategies, and JAX integration.
result Dramatic acceleration of models and new inference capabilities.
The purpose of this paper is to synthesize the approaches taken by Chatterjee-Meckes and Reinert-Röllin in adapting Stein's method of exchangeable pairs for multivariate normal approximation. The more general linear regression condition of Reinert-Röllin allows for wider applicability of the method, while the method of…
Ozsvath, Rasmussen and Szabo constructed odd Khovanov homology. It is a link invariant which has the same reduction modulo 2 as (even) Khovanov homology. Szabo introduced a spectral sequence with mod 2 coefficients from mod 2 Khovanov homology to another link homology. He got his spectral sequence from a chain complex …
Torchmeta simplifies meta-learning evaluation across multiple datasets.
problem Inconsistent evaluation of meta-learning algorithms across different datasets.
method Introduces a library that provides data-loaders for standard benchmarks and simplifies model compatibility.
result Seamless and consistent evaluation of meta-learning algorithms on multiple datasets.
This is the third of three papers that refine and extend portions of our earlier preprint, "The depth of a knot tunnel." Together, they rework the entire preprint. In this paper, we use the theory of tunnel number 1 knots that we introduced in "The tree of knot tunnels" to strengthen the Tunnel Leveling Theorem of H. G…
Unified approach for estimating quantiles of potential outcomes using inverse estimating equations.
problem Estimating quantiles of potential outcomes for causal inference.
method Inverse estimating equations and moment function.
result Unified approach to estimate mean and quantiles of potential outcomes.
Adapts GRPO for off-policy RL, improving reward.
problem Improving training stability and efficiency in RL.
method Adapts GRPO to off-policy setting, uses clipped surrogate objectives.
result Off-policy GRPO outperforms on-policy GRPO in empirical tests.
Paper tackles efficient evaluation of natural stochastic policies in offline RL.
problem Efficiency issues in evaluating natural stochastic policies due to unknown evaluation policy.
method Derive efficiency bounds for tilting and modified treatment policies, propose nonparametric estimators.
result Proposed estimators attain efficiency bounds under lax conditions and enjoy partial double robustness.
New framework studies policy learning problems under data scarcity.
problem Learning improving policies when data is insufficient.
method Developed a mathematical framework for policy learning problems.
result Reduced policy learning problems to simpler ones in sample complexity.
New algorithms improve policy evaluation in reinforcement learning.
problem Off-policy stability and on-policy efficiency issues in policy evaluation.
method Introduced novel algorithms using oblique projection method.
result Demonstrated both off-policy stability and on-policy efficiency.
We study the problem of off-policy policy optimization in Markov decision processes, and develop a novel off-policy policy gradient method. Prior off-policy policy gradient approaches have generally ignored the mismatch between the distribution of states visited under the behavior policy used to collect data, and what …
Stabilizes policy optimization with off-policy data using divergence augmentation.
problem Premature convergence and instability in policy optimization with off-policy data.
method Incorporates Bregman divergence between behavior and current policies to ensure safe policy updates.
result Empirically shows better performance in data-scarce scenarios compared to other algorithms.
New method estimates state-action stationary distribution for better off-policy policy evaluation.
problem Accurately estimating state-action stationary distribution for off-policy policy evaluation.
method Estimated Mixture Policy (EMP) for state and state-action stationary distribution corrections.
result Empirical validation shows improved accuracy over state-of-the-art methods.
We consider the problem of off-policy evaluation in Markov decision processes. Off-policy evaluation is the task of evaluating the expected return of one policy with data generated by a different, behavior policy. Importance sampling is a technique for off-policy evaluation that re-weights off-policy returns to account…
New methods estimate policy value and gradients for deterministic policies from off-policy data.
problem Estimating policy value and gradients for deterministic policies from off-policy data.
method Proposed new doubly robust estimators based on kernelization approaches.
result Demonstrated a rate independent of horizon length for policy value and gradient estimation.
Study designs logging policies to minimize off-policy evaluation error.
problem Minimizing OPE error with logging policies for target policies.
method Characterizes reward-coverage tradeoff, proposes a unifying framework, derives optimal policies.
result Provides actionable guidance for firms choosing recommendation systems.
DSPI connects natural policy gradient to policy iteration, proving global convergence.
problem Optimizing policies in reinforcement learning.
method DSPI framework, combining smoothed policy iteration and natural policy gradient.
result DSPI achieves geometric convergence and optimal complexity for policy optimization.
POTEC tackles off-policy learning in large action spaces, improving effectiveness.
problem Existing OPL methods fail in large discrete action spaces due to bias or variance issues.
method Two-stage algorithm: cluster selection via policy-based approach, action selection via regression-based approach.
result POTEC provides substantial improvements in off-policy learning effectiveness, especially in large and structured action spaces.
Monotonic policy improvement and off-policy learning are two main desirable properties for reinforcement learning algorithms. In this paper, by lower bounding the performance difference of two policies, we show that the monotonic policy improvement is guaranteed from on- and off-policy mixture samples. An optimization …
Memory-efficient algorithm reduces variance in off-policy RL.
problem High variance in off-policy policy optimization.
method Memory-efficient, stochastically variance-reduced algorithm using off-policy samples.
result Empirically validated effectiveness of the proposed algorithm.
In this work, we consider the problem of estimating a behaviour policy for use in Off-Policy Policy Evaluation (OPE) when the true behaviour policy is unknown. Via a series of empirical studies, we demonstrate how accurate OPE is strongly dependent on the calibration of estimated behaviour policy models: how precisely …
Protects proprietary policies from imitation learning by training adversarial policy ensembles.
problem Protecting policies from external observers cloning them.
method Introduces a reinforcement learning framework that trains an ensemble of near-optimal policies, making demonstrations useless for external observers.
result Demonstrates the existence of 'non-clonable' ensembles and provides a solution to the optimization problem.
Extends OPE to evaluate policies using diverse logging data.
problem Evaluate policies using log data from different policies.
method Develops an OPE method for various logging policies.
result Method's predictions converge to true performance as sample size increases.
PS framework selects best policy from library for CSO problems.
problem Policy selection in CSO with heterogeneous performance across covariate space.
method PS framework constructs library of candidate policies and learns a meta-policy to select the best one.
result PS consistently outperforms best single policy in heterogeneous CSO problems.
Entropy regularization improves policy optimization in reinforcement learning.
problem Improving policy optimization in reinforcement learning.
method Entropy regularization is introduced to soften the greedy policy towards a more diverse softmax policy, leading to a continuously parameterized algorithm that interpolates between policy gradient and Q-learning.
result An intermediate algorithm can improve performance in reinforcement learning.
New method optimizes treatment policies to avoid winner's curse.
problem Winner's curse in treatment policy optimization.
method Inference-aware policy optimization.
result Optimizes for both estimated performance and downstream evaluation.
PBVFs generalize across policies using learned value functions.
problem RL algorithms forget information about old policies when updating value functions to track the learned policy.
method Introduce Parameter-Based Value Functions (PBVFs) that include policy parameters in their inputs, enabling them to generalize across different policies.
result PBVFs enable zero-shot learning of new policies that outperform any policy seen during training.
New method improves off-policy critic evaluation in reinforcement learning.
problem High variance and instability in off-policy policy evaluation.
method Doubly robust estimators applied to actor-critic algorithms.
result Doubly robust estimation significantly improves performance in continuous control tasks.
Policy gradient aims to maximize expected return using gradient ascent.
problem Finding a policy that maximizes expected return in a given class of policies.
method Gradient ascent applied to a differentiable model of the policy, estimating the gradient of expected return.
result Policy gradient methods require on-policy data for gradient estimation, limiting sample efficiency.
Optimizes Thompson sampling policies using policy gradient methods.
problem Improving Thompson sampling in bandit problems.
method Applies policy gradient algorithms to optimize Thompson sampling policies.
result Direct policy search on Thompson sampling improves performance.
Study optimizes portfolio allocation policies using off-policy data and constraints.
problem Optimizing portfolio allocation policies under constraints using off-policy data.
method Solves a minimax objective with off-policy estimators and online learning to control constraint violations.
result Constructs near-optimal allocation policies for various regimes of operation and constraints.
Paper introduces a new policy optimization method using importance sampling.
problem Stable and low variance policy learning with small policy updates.
method Derives an alternative objective using importance sampling and introduces an approximation to balance bias and variance.
result The new algorithm improves on-policy policy optimization on continuous control benchmarks.
This paper extends off-policy reinforcement learning to the multi-agent case in which a set of networked agents communicating with their neighbors according to a time-varying graph collaboratively evaluates and improves a target policy while following a distinct behavior policy. To this end, the paper develops a multi-…
This paper introduces a new method to evaluate multiple policies simultaneously.
problem Estimating the value of many policies for a single set of states.
method Developed a scalable, differentiable fingerprinting mechanism to represent complex policies.
result The method can produce policies that outperform those that generated the training data, in zero-shot manner.
The paper interprets policy-gradient algorithms using continuation theory.
problem Optimizing nonconvex functions in reinforcement learning.
method Formulates policy optimization as optimization by continuation, interprets policy-gradient algorithms as implicitly optimizing deterministic policies.
result Exploration in policy-gradient algorithms is seen as computing a continuation of the return of the policy.
New method reduces state distribution mismatch in off-policy RL.
problem State distribution mismatch in off-policy RL algorithms.
method Develops a novel constrained off-policy gradient objective to minimize state distribution shift.
result Minimizing state distribution shift improves performance in off-policy RL algorithms.
This work analyzes the gap between off-policy and on-policy policy gradient methods and provides conditions to reduce this gap.
problem The gap between off-policy and on-policy policy gradient methods and conditions to reduce it.
method Theoretical analysis and empirical evidence of conditions to reduce the on-off gap.
result Conditions to reduce the on-off gap between off-policy and on-policy policy gradient methods.
We make policy optimization algorithms batch size-invariant by decoupling proximal and behavior policies.
problem Some policy optimization algorithms do not have batch size-invariance, leading to inefficiencies.
method We decouple the proximal policy from the behavior policy to achieve batch size-invariance.
result Our approach makes policy optimization algorithms more efficient and allows them to use stale data more effectively.