Forward translation improves neural machine translation for sentences originally in source language.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Microsoft Research Asia won first place in 8 out of 11 WMT19 language directions.
This paper uses LLMs and cycle consistency for better machine translation evaluation.
Paper proposes a new neural machine translation method for wave data.
The bits-back argument suggests that latent variable models can be turned into lossless compression schemes. Translating the bits-back argument into efficient and practical lossless compression schemes for general latent variable models, however, is still an open problem. Bits-Back with Asymmetric Numeral Systems (BB-A…
Let denote the identity connected component of the real orthogonal group with signature . We give a complete description of the spaces of continuous and generalized translation- and -invariant valuations, generalizing Hadwiger's classification of Euclidean isometry-invari…
The authors examine the concept of probability of default for asset-backed loans. In contrast to unsecured loans it is shown that probability of default can be defined as either a measure of the likelihood of the borrower failing to make required payments, or as the likelihood of an insufficiency of collateral value on…
BBRT improves molecular properties through iterative translation.
CERT improves language understanding by contrastively learning sentence-level semantics.
The paper explores dual learning, a technique that improves machine translation and image transformation.
We present BART, a denoising autoencoder for pretraining sequence-to-sequence models. BART is trained by (1) corrupting text with an arbitrary noising function, and (2) learning a model to reconstruct the original text. It uses a standard Tranformer-based neural machine translation architecture which, despite its simpl…
Given datasets from multiple domains, a key challenge is to efficiently exploit these data sources for modeling a target domain. Variants of this problem have been studied in many contexts, such as cross-domain translation and domain adaptation. We propose AlignFlow, a generative modeling framework that models each dom…
We generalize the natural cross ratio on the ideal boundary of a rank one symmetric spaces, or even space, to higher rank symmetric spaces and (non-locally compact) Euclidean buildings - we obtain vector valued cross ratios defined on simplices of the building at infinity. We show several properties …
In this work, we address the problem of modifying textual attributes of sentences. Given an input sentence and a set of attribute labels, we attempt to generate sentences that are compatible with the conditioning information. To ensure that the model generates content compatible sentences, we introduce a reconstruction…
Let be a real-analytic manifold and a proper triangulable subanalytic map. Given a subanalytic -form on whose pull-back to every non singular fiber of is exact, we show tha has a relative primitive: there is a subanalytic -form such that . The p…
Neural language models have been widely used in various NLP tasks, including machine translation, next word prediction and conversational agents. However, it is challenging to deploy these models on mobile devices due to their slow prediction speed, where the bottleneck is to compute top candidates in the softmax layer…
Generalising well in supervised learning tasks relies on correctly extrapolating the training data to a large region of the input space. One way to achieve this is to constrain the predictions to be invariant to transformations on the input that are known to be irrelevant (e.g. translation). Commonly, this is done thro…
This paper proposes an alternating back-propagation algorithm for learning the generator network model. The model is a non-linear generalization of factor analysis. In this model, the mapping from the continuous latent factors to the observed signal is parametrized by a convolutional neural network. The alternating bac…
The vast majority of successful deep neural networks are trained using variants of stochastic gradient descent (SGD) algorithms. Recent attempts to improve SGD can be broadly categorized into two approaches: (1) adaptive learning rate schemes, such as AdaGrad and Adam, and (2) accelerated schemes, such as heavy-ball an…
Kernel Multigrid accelerates Back-fitting for additive Gaussian Processes.
A new method combines energy-based models and entropy-regularized optimal transport.
We prove a Bochner type vanishing theorem for compact complex manifolds in Fujiki class , with vanishing first Chern class, that admit a cohomology class which is numerically effective (nef) and has positive self-intersection (meaning , where $n\,=\,\di…
Recurrent Networks are one of the most powerful and promising artificial neural network algorithms to processing the sequential data such as natural languages, sound, time series data. Unlike traditional feed-forward network, Recurrent Network has a inherent feed back loop that allows to store the temporal context info…
As traditional neural network consumes a significant amount of computing resources during back propagation, \citet{Sun2017mePropSB} propose a simple yet effective technique to alleviate this problem. In this technique, only a small subset of the full gradients are computed to update the model parameters. In this paper …
Generalizes bits back coding for time-series models with latent Markov structures.
We show how to restructure the counterparty risk faced by the originator of a securitization or covered bond arising from an interest rate hedging swap assisted by a "one-way" collateral agreement. This risk emerges when the swap is negotiated between the special purpose vehicle and a third party that covers itself thr…
The back-propagation algorithm is the cornerstone of deep learning. Despite its importance, few variations of the algorithm have been attempted. This work presents an approach to discover new variations of the back-propagation equation. We use a domain specific lan- guage to describe update equations as a list of primi…
We define the pull-back of a smooth principal fibre bundle, and show that it has a natural principal fibre bundle structure. Next, we analyse the relationship between pull-backs by homotopy equivalent maps. The main result of this article is to show that for a principal fibre bundle over a paracompact manifold, there i…
Empirical law predicts accuracy of Google Translate's translation chains.
Interpretable representations improve explainable AI by translating complex data into understandable concepts.
The authors of (Cho et al., 2014a) have shown that the recently introduced neural network translation systems suffer from a significant drop in translation quality when translating long sentences, unlike existing phrase-based translation systems. In this paper, we propose a way to address this issue by automatically se…
We present a voice conversion solution using recurrent sequence to sequence modeling for DNNs. Our solution takes advantage of recent advances in attention based modeling in the fields of Neural Machine Translation (NMT), Text-to-Speech (TTS) and Automatic Speech Recognition (ASR). The problem consists of converting be…
Deep learning predicts back-pain risk during manual lifting.
Paper proposes a method to learn word translations bidirectionally.
Under a pulled-back approach given in [1] and firstly presented in [2], we introduce, in this paper, the concepts of almost contact and normal almost contact Finsler structures on the pulled-back bundle. Properties of structures partly Sasakians are studied. Using the hh-curvature tensor of Chern connection given in [2…
Task-agnostic data augmentation shows little benefit for pretrained transformers.
The paper constructs a multi-valued inverse of quasiregular maps and develops pull-back theory for differential forms.
We develop a new approach to the pulling back fixed point theorem of W. Browder and use it in order to prove various generalizations of this result.
Stochastic gradient descent (SGD) has achieved great success in training deep neural network, where the gradient is computed through back-propagation. However, the back-propagated values of different layers vary dramatically. This inconsistence of gradient magnitude across different layers renders optimization of deep …
It is shown by several authors going back to Huisken-Yau that asymptotically Schwarzschildean time-slices possess a unique foliation by stable constant mean curvature (CMC) spheres defining the so-called CMC center of mass. We analyze how the leaves of this foliation evolve in time under the Einstein equations. More pr…
Study on rigidity of translating hypersurfaces not in graphical direction.
Narasihman and Ramanan proved that an arbitrary connection in a vector bundle over a base space B can be obtained as the pull-back (via a correctly chosen classifying map from B into the appropriate Grassmannian) of the universal connection in the universal bundle over the Grassmannian. The purpose of this paper is to …
This study examines the execution phase of corporate share buy-backs, highlighting inefficiencies and costs.
The natural automorphism group of a translation surface is its group of translations. For finite translation surfaces of genus g > 1 the order of this group is naturally bounded in terms of g due to a Riemann-Hurwitz formula argument. In analogy with classical Hurwitz surfaces, we call surfaces which achieve the maxima…
Paper introduces a new risk measure for multivariate residual estimation.
Study on stable translation lengths of surface homeomorphisms and their approximations.
Researchers classify and describe -translators in Euclidean space.
The quality of machine translation is rapidly evolving. Today one can find several machine translation systems on the web that provide reasonable translations, although the systems are not perfect. In some specific domains, the quality may decrease. A recently proposed approach to this domain is neural machine translat…