Framework for differentiating WFSTs for structured loss functions.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Develops STC for sequential data with missing labels.
This paper formalizes Uniswap v3 using PTA and FST for rigorous analysis.
Improved neural transducer model outperforms attention model on longer sequences.
Improved speech recognition model with better performance.
Conditional random fields (CRFs) have been shown to be one of the most successful approaches to sequence labeling. Various linear-chain neural CRFs (NCRFs) are developed to implement the non-linear node potentials in CRFs, but still keeping the linear-chain hidden structure. In this paper, we propose NCRF transducers, …
A keyword spotting (KWS) system determines the existence of, usually predefined, keyword in a continuous speech stream. This paper presents a query-by-example on-device KWS system which is user-specific. The proposed system consists of two main steps: query enrollment and testing. In query enrollment step, phonetic pos…
A new method calculates optimal decisions from classifier outputs, improving predictions in drug discovery.
Requirements elicitation can be very challenging in projects that require deep domain knowledge about the system at hand. As analysts have the full control over the elicitation process, their lack of knowledge about the system under study inhibits them from asking related questions and reduces the accuracy of requireme…
We propose a method for modeling and learning turn-taking behaviors for accessing a shared resource. We model the individual behavior for each agent in an interaction and then use a multi-agent fusion model to generate a summary over the expected actions of the group to render the model independent of the number of age…
Phylogenetic tree reconstruction is traditionally based on multiple sequence alignments (MSAs) and heavily depends on the validity of this information bottleneck. With increasing sequence divergence, the quality of MSAs decays quickly. Alignment-free methods, on the other hand, are based on abstract string comparisons …
Requirements elicitation requires extensive knowledge and deep understanding of the problem domain where the final system will be situated. However, in many software development projects, analysts are required to elicit the requirements from an unfamiliar domain, which often causes communication barriers between analys…
Acoustic Neighbor Embeddings map speech and text to fixed dimensions for phonetic confusability.
Having a sequence-to-sequence model which can operate in an online fashion is important for streaming applications such as Voice Search. Neural transducer is a streaming sequence-to-sequence model, but has shown a significant degradation in performance compared to non-streaming models such as Listen, Attend and Spell (…
Streamable model improves speech recognition performance.
The paper simplifies multi-agent RL dynamics in finite-state Markov games using homogenization.
A new layer learns abstract relations from graph structure using finite-state automata.
New methods improve integration of external LMs with AED models.
In this paper, we first establish the reflected backward stochastic difference equations with finite state (FS-RBSDEs for short). Then we explore the Existence and Uniqueness Theorem as well as the Comparison Theorem by "one step" method. The connections between FS-RBSDEs and optimal stopping time problems are investig…
In his 2011 work, Maas has shown that the law of any time-reversible continuous-time Markov chain with finite state space evolves like a gradient flow of the relative entropy with respect to its stationary distribution. In this work we show the converse to the above by showing that if the relative law of a Markov chain…
We extend a recent synchronization analysis of exact finite-state sources to nonexact sources for which synchronization occurs only asymptotically. Although the proof methods are quite different, the primary results remain the same. We find that an observer's average uncertainty in the source state vanishes exponential…
The paper explores geometric calculations on probability manifolds derived from master equations.
A new framework improves ASR alignment accuracy via optimal transport.
New approach to understand recurrent policies as FSMs without minimization.
A fixed point theorem is proved for inverse transducers, leading to an automata-theoretic proof of the fixed point subgroup of an endomorphism of a finitely generated virtually free group being finitely generated. If the endomorphism is uniformly continuous for the hyperbolic metric, it is proved that the set of regula…
Two neural network methods solve the master equation for MFGs.
A framework solves parametric families of MFGs efficiently.
We generalize to the finite-state case the notion of the extreme effect variable that accumulates all the effect of a variant variable observed in changes of another variable . We conduct theoretical analysis and turn the problem of finding of an effect variable into a problem of a simultaneous decomposition…
Transformers simulate finite-state automata with fewer layers.
This paper develops a Hoeffding inequality for the partial sums , where is an irreducible Markov chain on a finite state space , and is a real-valued function. Our bound is simple, general, since it only assumes irreducibility and finiteness…
The smallest eigenvectors of the graph Laplacian are well-known to provide a succinct representation of the geometry of a weighted graph. In reinforcement learning (RL), where the weighted graph may be interpreted as the state transition process induced by a behavior policy acting on the environment, approximating the …
We analyze how an observer synchronizes to the internal state of a finite-state information source, using the epsilon-machine causal representation. Here, we treat the case of exact synchronization, when it is possible for the observer to synchronize completely after a finite number of observations. The more difficult …
Stochastic gradient methods are the workhorse (algorithms) of large-scale optimization problems in machine learning, signal processing, and other computational sciences and engineering. This paper studies Markov chain gradient descent, a variant of stochastic gradient descent where the random samples are taken on the t…
We stabilize the activations of Recurrent Neural Networks (RNNs) by penalizing the squared distance between successive hidden states' norms. This penalty term is an effective regularizer for RNNs including LSTMs and IRNNs, improving performance on character-level language modeling and phoneme recognition, and outperfor…
Conformal Prediction Regions match Imprecise Highest Density Regions under consonance.
The potential approach is a general and simple method for modelling interest rates, foreign exchange rates, and in principle other types of financial assets. This paper takes data on some liquid interest rate derivatives, and fits potential models using a small finite-state Markov chain as the base Markov process.
The paper tackles restless bandits with limited observation, proposing a method to analyze and approximate their optimal strategies.
We study an open problem of risk-sensitive portfolio allocation in a regime-switching credit market with default contagion. The state space of the Markovian regime-switching process is assumed to be a countably infinite set. To characterize the value function, we investigate the corresponding recursive infinite-dimensi…
We consider the estimation of the policy gradient in partially observable Markov decision processes (POMDP) with a special class of structured policies that are finite-state controllers. We show that the gradient estimation can be done in the Actor-Critic framework, by making the critic compute a "value" function that …
Recurrent neural networks trained on regular languages exhibit stable states that can recover from noise.
Study shows challenges in converting RNNs to FSMs due to computational complexity.
In this paper, a finite-state mean-reverting model for the short-rate, based on the continuous time Ehrenfest process, will be examined. Two explicit pricing formulae for zero-coupon bonds will be derived in the general and the special symmetric cases. Its limiting relationship to the Vasicek model will be examined wit…
Recurrent neural networks are a widely used class of neural architectures. They have, however, two shortcomings. First, it is difficult to understand what exactly they learn. Second, they tend to work poorly on sequences requiring long-term memorization, despite having this capacity in principle. We aim to address both…
This work optimizes MCMC algorithms for modern accelerators without synchronization overheads.
Most of the parameters in large vocabulary models are used in embedding layer to map categorical features to vectors and in softmax layer for classification weights. This is a bottle-neck in memory constraint on-device training applications like federated learning and on-device inference applications like automatic spe…
We propose a new way of thinking about deep neural networks, in which the linear and non-linear components of the network are naturally derived and justified in terms of principles in probability theory. In particular, the models constructed in our framework assign probabilities to uncertain realizations, leading to Ku…
Reservoir computers (RCs) and recurrent neural networks (RNNs) can mimic any finite-state automaton in theory, and some workers demonstrated that this can hold in practice. We test the capability of generalized linear models, RCs, and Long Short-Term Memory (LSTM) RNN architectures to predict the stochastic processes g…
Study on natural actor-critic for POMDPs with finite memory.