End-to-end ASR error detection using audio-transcript entailment.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
In training a deep learning system to perform audio transcription, two practical problems may arise. Firstly, most datasets are weakly labelled, having only a list of events present in each recording without any temporal information for training. Secondly, deep neural networks need a very large amount of labelled train…
The application of deep recurrent networks to audio transcription has led to impressive gains in automatic speech recognition (ASR) systems. Many have demonstrated that small adversarial perturbations can fool deep neural networks into incorrectly predicting a specified target with high confidence. Current work on fool…
Neural language models are a critical component of state-of-the-art systems for machine translation, summarization, audio transcription, and other tasks. These language models are almost universally autoregressive in nature, generating sentences one token at a time from left to right. This paper studies the influence o…
The paper proposes methods to control errors in language generation models using textual entailment.
Word embeddings provide point representations of words containing useful semantic information. We introduce multimodal word distributions formed from Gaussian mixtures, for multiple word meanings, entailment, and rich uncertainty information. To learn these distributions, we propose an energy-based max-margin objective…
Training modern deep learning models requires large amounts of computation, often provided by GPUs. Scaling computation from one GPU to many can enable much faster training and research progress but entails two complications. First, the training library must support inter-GPU communication. Depending on the particular …
Learning graph representations via low-dimensional embeddings that preserve relevant network properties is an important class of problems in machine learning. We here present a novel method to embed directed acyclic graphs. Following prior work, we first advocate for using hyperbolic spaces which provably model tree-li…
In this article we extend cutting and blowing up to the nonrational symplectic toric setting. This entails the possibility of cutting and blowing up for symplectic toric manifolds and orbifolds in nonrational directions.
By representing words with probability densities rather than point vectors, probabilistic word embeddings can capture rich and interpretable semantic information and uncertainty. The uncertainty information can be particularly meaningful in capturing entailment relationships -- whereby general words such as "entity" co…
Learning word representations has garnered greater attention in the recent past due to its diverse text applications. Word embeddings encapsulate the syntactic and semantic regularities of sentences. Modelling word embedding as multi-sense gaussian mixture distributions, will additionally capture uncertainty and polyse…
ICP improves text infilling and POS tagging with valid confidence sets.
The Trek Separation Theorem (Sullivant et al. 2010) states necessary and sufficient conditions for a linear directed acyclic graphical model to entail for all possible values of its linear coefficients that the rank of various sub-matrices of the covariance matrix is less than or equal to n, for any given n. In this pa…
R package for multi-objective model selection in statistics.
This paper is an attempt at understanding the quantum-like dynamics of financial markets in terms of non-differentiable price-time continuum having fractal properties. The main steps of this development are the statistical scaling, the non-differentiability hypothesis, and the equations of motion entailed by this hypot…
We prove that a wide class of correlated stochastic volatility models exactly measure an empirical fact in which past returns are anticorrelated with future volatilities: the so-called ``leverage effect''. This quantitative measure allows us to fully estimate all parameters involved and it will entail a deeper study on…
Modern neural networks are often augmented with an attention mechanism, which tells the network where to focus within the input. We propose in this paper a new framework for sparse and structured attention, building upon a smoothed max operator. We show that the gradient of this operator defines a mapping from real val…
We study online prediction of bounded stationary ergodic processes. To do so, we consider the setting of prediction of individual sequences and build a deterministic regression tree that performs asymptotically as well as the best L-Lipschitz constant predictors. Then, we show why the obtained regret bound entails the …
We reformulate Lehmer's question from 1933 and a question due to Schinzel and Zassenhaus from 1965 in terms of a comparison of the Mahler measures and the houses, respectively, of monic integer reciprocal and skew-reciprocal polynomials of the same degree. This entails that understanding the difference between orientat…
The Killing operator on a Riemannian manifold is a linear differential operator on vector fields whose kernel provides the infinitesimal Riemannian symmetries. The Killing operator is best understood in terms of its prolongation, which entails some simple tensor identities. These simple identities can be viewed as aris…
In this paper we prove explicit formulas for all Willmore surfaces of revolution and demonstrate their use in the discussion of the associated Dirichlet boundary value problems. It is shown by an explicit example that symmetric Dirichlet boundary conditions do in general not entail the symmetry of the surface. In addit…
Embedding methods which enforce a partial order or lattice structure over the concept space, such as Order Embeddings (OE) (Vendrov et al., 2016), are a natural way to model transitive relational data (e.g. entailment graphs). However, OE learns a deterministic knowledge base, limiting expressiveness of queries and the…
We discuss the construction of Sp(2)Sp(1)-structures whose fundamental form is closed. In particular, we find 10 new examples of 8-dimensional nilmanifolds that admit an invariant closed 4-form with stabiliser Sp(2)Sp(1). Our constructions entail the notion of SO(4)-structures on 7-manifolds. We present a thorough inve…
Design of reliable systems must guarantee stability against input perturbations. In machine learning, such guarantee entails preventing overfitting and ensuring robustness of models against corruption of input data. In order to maximize stability, we analyze and develop a computationally efficient implementation of Jac…
Proposes a new neural head for asymmetric representation learning.
The mean field methods, which entail approximating intractable probability distributions variationally with distributions from a tractable family, enjoy high efficiency, guaranteed convergence, and provide lower bounds on the true likelihood. But due to requirement for model-specific derivation of the optimization equa…
Taylor expansions improve reinforcement learning policies.
Paper characterizes causal graphs from hard interventions and proposes a learning algorithm.
We study the first uniformly finite homology group of Block and Weinberger for uniformly locally finite graphs, with coefficients in and . When the graph is a tree, or coefficients are in , a characterisation of the group is obtained. In the general case, we describe three pheno…
Multi-hop inference is necessary for machine learning systems to successfully solve tasks such as Recognising Textual Entailment and Machine Reading. In this work, we demonstrate the effectiveness of adaptive computation for learning the number of inference steps required for examples of different complexity and that l…
Combines curvature descriptors with TDA for graph model evaluation.
Jacobi solved geodesics on triaxial ellipsoids.
Stable solutions to a specific equation are one-dimensional.
New proof of harmonic map uniqueness with analytic targets.
We show that the metrical connection can be introduced in the two-dimensional Finsler space such that entailed parallel transports along curves joining points of the underlying manifold keep the two-vector angle as well as the length of the tangent vector, thereby realizing isometries of tangent spaces under the parall…
Geodesic algorithms extended to arbitrary ellipsoids.
VAEs struggle with surjective multimodal data, especially class labels describing images.
Temporal Point Processes (TPP) with partial likelihoods involving a latent structure often entail an intractable marginalization, thus making inference hard. We propose a novel approach to Maximum Likelihood Estimation (MLE) involving approximate inference over the latent variables by minimizing a tight upper bound on …
TabPFN quickly classifies small tabular data without tuning.
We study vector fields generating a local flow by automorphisms of a parabolic geometry with higher order fixed points. We develop general tools extending the techniques of [1], [2], and [3]. We apply these tools to almost Grassmannian, almost quaternionic, and contact parabolic geometries, including CR structures, to …
The pseudo-Finsleroid relativistic metric was constructed upon assuming that the involved vector field is time-like. In the present paper it is shown that the metric admits just the alternative counterpart in which the field is space-like. The entailed pseudo-Finsleroid-spatial framework is systematically describ…
Current machine learning systems operate, almost exclusively, in a statistical, or model-free mode, which entails severe theoretical limits on their power and performance. Such systems cannot reason about interventions and retrospection and, therefore, cannot serve as the basis for strong AI. To achieve human level int…
SNN architecture shows gradient descent converges to regularized solution in matrix sensing problems.
Paper summarizes unsupervised learning challenges for disentangled representations.
We formalize in the proof assistant Isabelle essential basic notions and results in financial mathematics. We provide generic formal definitions of concepts such as markets, portfolios, derivative products, arbitrages or fair prices, and we show that, under the usual no-arbitrage condition, the existence of a replicati…
In this note we consider homogeneous Willmore surfaces in . The main result is that a homogeneous Willmore two-sphere is conformally equivalent to a homogeneous minimal two-sphere in , i.e., either a round two-sphere or one of the Borůvka-Veronese 2-spheres in . This entails a classification o…
Obtaining continuous representations of structural data such as directed acyclic graphs (DAGs) has gained attention in machine learning and artificial intelligence. However, embedding complex DAGs in which both ancestors and descendants of nodes are exponentially increasing is difficult. Tackling in this problem, we de…
Machine learning tasks entail the use of complex computational pipelines to reach quantitative and qualitative conclusions. If some of the activities in a pipeline produce erroneous or uninformative outputs, the pipeline may fail or produce incorrect results. Inferring the root cause of failures and unexpected behavior…