The paper uses machine learning to optimize rework policies in semiconductor manufacturing.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The paper uses causal machine learning to optimize rework decisions in manufacturing.
These are course notes I wrote for my Fall 2013 graduate topics course on geometric structures, taught at ICERM. The notes rework many of proofs in William P. Thurston's beautiful but hard-to-understand paper, "Shapes of Polyhedra". A number of people, both in and out of the class, found these notes very useful and so …
We show how the theory of tangles is equivalent to that of well-connected tangles. These are drawn on a surface with boundary, and equivalent via Reidemeister moves of a restricted kind. This reworking of the graphical foundations for link and tangle theory can be expected to have a variety of applications, including o…
We demonstrate a conditional autoregressive pipeline for efficient music recomposition, based on methods presented in van den Oord et al.(2017). Recomposition (Casal & Casey, 2010) focuses on reworking existing musical pieces, adhering to structure at a high level while also re-imagining other aspects of the work. This…
This is the second of three papers that refine and extend portions of our earlier preprint, "The depth of a knot tunnel." Together, they rework the entire preprint. The theory of tunnel number 1 knots that we introduced in "The tree of knot tunnels" yields a parameterization in which each tunnel is described uniquely b…
In geometric group theory one uses group actions on spaces to gain information about groups. One natural space to use is the Cayley graph of a group. The Cayley graph arguments that one encounters tend to require local finiteness, and hence finite generation of the group. In this paper, I take the theory of intersectio…
This paper means to correct an error by the authors for the composite case in the paper "Lens Spaces, Isospectral on Forms but not on Functions", published in LMS J. Comput. Math.} 9 (2006), 270-286. All calculations and examples presented in \cite{GM} for prime remain valid, and we include detailed calculation…
This is the first of three papers that refine and extend portions of our earlier preprint, "Depth of a knot tunnel." Together, they rework the entire preprint. H. Goda, M. Scharlemann, and A. Thompson described a general construction of all tunnels of all tunnel number 1 knots using "tunnel moves". We apply the theory …
The monoids of simplicial endomorphisms, i.e. the monoids of endomorphisms in the simplicial category, are submonoids of monoids one finds in Temperley-Lieb algebras, and as the monoids of Temperley-Lieb algebras are linked to situations where an endofunctor is adjoint to itself, so the monoids of simplicial endomorphi…
NIFTy.re accelerates imaging models and expands Gaussian processes and variational inference.
The purpose of this paper is to synthesize the approaches taken by Chatterjee-Meckes and Reinert-Röllin in adapting Stein's method of exchangeable pairs for multivariate normal approximation. The more general linear regression condition of Reinert-Röllin allows for wider applicability of the method, while the method of…
Ozsvath, Rasmussen and Szabo constructed odd Khovanov homology. It is a link invariant which has the same reduction modulo 2 as (even) Khovanov homology. Szabo introduced a spectral sequence with mod 2 coefficients from mod 2 Khovanov homology to another link homology. He got his spectral sequence from a chain complex …
The constant introduction of standardized benchmarks in the literature has helped accelerating the recent advances in meta-learning research. They offer a way to get a fair comparison between different algorithms, and the wide range of datasets available allows full control over the complexity of this evaluation. Howev…
This is the third of three papers that refine and extend portions of our earlier preprint, "The depth of a knot tunnel." Together, they rework the entire preprint. In this paper, we use the theory of tunnel number 1 knots that we introduced in "The tree of knot tunnels" to strengthen the Tunnel Leveling Theorem of H. G…
Unified approach for estimating quantiles of potential outcomes using inverse estimating equations.
Adapts GRPO for off-policy RL, improving reward.
Paper tackles efficient evaluation of natural stochastic policies in offline RL.
New framework studies policy learning problems under data scarcity.
New algorithms improve policy evaluation in reinforcement learning.
We study the problem of off-policy policy optimization in Markov decision processes, and develop a novel off-policy policy gradient method. Prior off-policy policy gradient approaches have generally ignored the mismatch between the distribution of states visited under the behavior policy used to collect data, and what …
Stabilizes policy optimization with off-policy data using divergence augmentation.
We consider the problem of off-policy evaluation in Markov decision processes. Off-policy evaluation is the task of evaluating the expected return of one policy with data generated by a different, behavior policy. Importance sampling is a technique for off-policy evaluation that re-weights off-policy returns to account…
New methods estimate policy value and gradients for deterministic policies from off-policy data.
Study designs logging policies to minimize off-policy evaluation error.
DSPI connects natural policy gradient to policy iteration, proving global convergence.
POTEC tackles off-policy learning in large action spaces, improving effectiveness.
Monotonic policy improvement and off-policy learning are two main desirable properties for reinforcement learning algorithms. In this paper, by lower bounding the performance difference of two policies, we show that the monotonic policy improvement is guaranteed from on- and off-policy mixture samples. An optimization …
Memory-efficient algorithm reduces variance in off-policy RL.
In this work, we consider the problem of estimating a behaviour policy for use in Off-Policy Policy Evaluation (OPE) when the true behaviour policy is unknown. Via a series of empirical studies, we demonstrate how accurate OPE is strongly dependent on the calibration of estimated behaviour policy models: how precisely …
Extends OPE to evaluate policies using diverse logging data.
PS framework selects best policy from library for CSO problems.
Entropy regularization improves policy optimization in reinforcement learning.
New method optimizes treatment policies to avoid winner's curse.
PBVFs generalize across policies using learned value functions.
Optimizes Thompson sampling policies using policy gradient methods.
Study optimizes portfolio allocation policies using off-policy data and constraints.
This paper extends off-policy reinforcement learning to the multi-agent case in which a set of networked agents communicating with their neighbors according to a time-varying graph collaboratively evaluates and improves a target policy while following a distinct behavior policy. To this end, the paper develops a multi-…
This paper introduces a new method to evaluate multiple policies simultaneously.
The paper interprets policy-gradient algorithms using continuation theory.
We study the problem of off-policy critic evaluation in several variants of value-based off-policy actor-critic algorithms. Off-policy actor-critic algorithms require an off-policy critic evaluation step, to estimate the value of the new policy after every policy gradient update. Despite enormous success of off-policy …
This work analyzes the gap between off-policy and on-policy policy gradient methods and provides conditions to reduce this gap.
We make policy optimization algorithms batch size-invariant by decoupling proximal and behavior policies.
Improves reinforcement learning stability and efficiency.
RPO uses past and future state-action info for better policy optimization.
New algorithm stabilizes RL policy learning through divergence regularization.
Framework improves policy generalizability under biased training data.
Hybrid RL algorithm combines offline and online data for robust and efficient policy learning.