Forward translation improves neural machine translation for sentences originally in source language.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Let denote the identity connected component of the real orthogonal group with signature . We give a complete description of the spaces of continuous and generalized translation- and -invariant valuations, generalizing Hadwiger's classification of Euclidean isometry-invari…
This paper uses LLMs and cycle consistency for better machine translation evaluation.
Demographic projections of future mortality rates involve a high level of uncertainty and require stochastic mortality models. The current paper investigates forward mortality models driven by a (possibly infinite dimensional) Wiener process and a compensated Poisson random measure. A major innovation of the paper is t…
Transformer networks have lead to important progress in language modeling and machine translation. These models include two consecutive modules, a feed-forward layer and a self-attention layer. The latter allows the network to capture long term dependencies and are often regarded as the key ingredient in the success of…
End-to-end optimization has achieved state-of-the-art performance on many specific problems, but there is no straight-forward way to combine pretrained models for new problems. Here, we explore improving modularity by learning a post-hoc interface between two existing models to solve a new task. Specifically, we take i…
Recurrent neural networks (RNNs) sequentially process data by updating their state with each new data point, and have long been the de facto choice for sequence modeling tasks. However, their inherently sequential computation makes them slow to train. Feed-forward and convolutional architectures have recently been show…
LayerNorm improves training and generalization by re-centering and re-scaling backward gradients.
We tackle unsupervised anomaly detection (UAD), a problem of detecting data that significantly differ from normal data. UAD is typically solved by using density estimation. Recently, deep neural network (DNN)-based density estimators, such as Normalizing Flows, have been attracting attention. However, one of their draw…
New MIF architecture improves posterior approximations in Bayesian models.
Ancient grain boundaries resemble atoms in their formation and properties.
A new QHR model extends HR model with a quadratic variance function.
AAMDRL uses DRL to manage assets in noisy, changing environments.
A new method uses normalizing flows to approximate optimal transport between empirical distributions.
The theory of convex risk functions has now been well established as the basis for identifying the families of risk functions that should be used in risk averse optimization problems. Despite its theoretical appeal, the implementation of a convex risk function remains difficult, as there is little guidance regarding ho…
We introduce a new and highly tractable structural model for spot and derivative prices in electricity markets. Using a stochastic model of the bid stack, we translate the demand for power and the prices of generating fuels into electricity spot prices. The stack structure allows for a range of generator efficiencies p…
Generative model uses DDPMs for risk-neutral derivative pricing.
We introduce dropout compaction, a novel method for training feed-forward neural networks which realizes the performance gains of training a large model with dropout regularization, yet extracts a compact neural network for run-time efficiency. In the proposed method, we introduce a sparsity-inducing prior on the per u…
Paper forecasts stock correlations using a hybrid model combining graph neural networks and transformers.
The capacity of a neural network to absorb information is limited by its number of parameters. Conditional computation, where parts of the network are active on a per-example basis, has been proposed in theory as a way of dramatically increasing model capacity without a proportional increase in computation. In practice…
We introduce two models of taxation, the latent and natural tax processes, which have both been used to represent loss-carry-forward taxation on the capital of an insurance company. In the natural tax process, the tax rate is a function of the current level of capital, whereas in the latent tax process, the tax rate is…
Deep learning's success requires vast computing power, making future progress unsustainable.
Empirical law predicts accuracy of Google Translate's translation chains.
The authors of (Cho et al., 2014a) have shown that the recently introduced neural network translation systems suffer from a significant drop in translation quality when translating long sentences, unlike existing phrase-based translation systems. In this paper, we propose a way to address this issue by automatically se…
Paper proposes a method to learn word translations bidirectionally.
The vast majority of successful deep neural networks are trained using variants of stochastic gradient descent (SGD) algorithms. Recent attempts to improve SGD can be broadly categorized into two approaches: (1) adaptive learning rate schemes, such as AdaGrad and Adam, and (2) accelerated schemes, such as heavy-ball an…
Study on rigidity of translating hypersurfaces not in graphical direction.
Bayesian deep learning improves seismic imaging uncertainty.
The natural automorphism group of a translation surface is its group of translations. For finite translation surfaces of genus g > 1 the order of this group is naturally bounded in terms of g due to a Riemann-Hurwitz formula argument. In analogy with classical Hurwitz surfaces, we call surfaces which achieve the maxima…
Study on stable translation lengths of surface homeomorphisms and their approximations.
Researchers classify and describe -translators in Euclidean space.
The quality of machine translation is rapidly evolving. Today one can find several machine translation systems on the web that provide reasonable translations, although the systems are not perfect. In some specific domains, the quality may decrease. A recently proposed approach to this domain is neural machine translat…
Study measures gender bias in machine translation using multiple reference points.
Classifies and constructs translators for curvature flows.
Paper finds a non-existence theorem for certain translators in high dimensions.
Constructing translating solitons from Lagrangian Grim Reapers.
Study classifies translators for mean curvature flow in 3D.
Study on singular points of translation surfaces under linearly dependent conditions.
Proves uniqueness of translators in 3D space.
An NMT system for Indic languages outperforms Google Translate.
New families of translation surfaces with multiple oblivious points discovered.
A fundamental problem in geophysical modeling is related to the identification and approximation of causal structures among physical processes. However, resolving the bidirectional mappings between physical parameters and model state variables (i.e., solving the forward and inverse problems) is challenging, especially …
In this article we prove two non-existence results for translating solitons of the mean curvature flow (translators for short) in . We also obtain an upper bound to the maximum height that a compact embedded translator in can achieve. On the other hand, we study graphical perturbation…
This paper classifies grim reapers in a specific product space.
Neural machine translation is a recently proposed approach to machine translation. Unlike the traditional statistical machine translation, the neural machine translation aims at building a single neural network that can be jointly tuned to maximize the translation performance. The models proposed recently for neural ma…
Recurrent Networks are one of the most powerful and promising artificial neural network algorithms to processing the sequential data such as natural languages, sound, time series data. Unlike traditional feed-forward network, Recurrent Network has a inherent feed back loop that allows to store the temporal context info…
Study invariant -translators in Lorentz-Minkowski space.
Unique minimal translation surfaces found in Heisenberg group.