Improves ASR accuracy for domain mismatch using machine translation.
problem Domain mismatch in ASR systems leads to suboptimal results.
method Machine translation to map out-of-domain ASR errors to in-domain terms.
result 7% absolute improvement in word error rate, 4 point BLEU score improvement.
Paper tackles NAT translation issues with auxiliary regularization.
problem Improves NAT translation quality by addressing repeated and incomplete translations.
method Improves decoder hidden representations via two auxiliary regularization terms.
result Significant improvement in NAT model accuracy with better inference efficiency.
The paper proves impossibilities and positive results for universal machine translation.
problem Learning shared sentence representations across multiple language pairs.
method Formal proofs and analysis of natural generative processes.
result Lower bound on translation error and positive results under natural structure.
SPoC uses search to translate pseudocode into correct programs with error localization.
problem Mapping pseudocode to functionally correct long programs.
method Search-based approach guided by compilation errors for credit assignment.
result Search improves synthesis success rate from 25.6% to 44.7%.
Improves NMT by sampling context from predicted sequence during training.
problem Error accumulation and overcorrection in NMT due to mismatched training and inference contexts.
method Samples context words from both ground truth and predicted sequences during training.
result Significant improvements on multiple datasets, including Chinese->English and WMT'14 English->German.
Paper identifies and solves a 'scrambled translation' issue in UNMT models.
problem Scrambled translation in UNMT models using word shuffle noise.
method Retraining UNMT models without noise after a pre-defined number of iterations.
result Retraining strategy improves BLEU scores by 10-20% on various language pairs.
Optimal square matrices for image approximation under translation and rotation.
problem Approximating images with translation and rotation invariant subspaces.
method Abstract harmonic analysis for constructing optimal square matrices.
result Optimal approximation of images with minimal quadratic error.
This handbook translates lead time analysis into R code.
problem Tracking divergence in booking lead times over time.
method Translated original article's methodology into R code.
result Demonstrated reproducibility and error bounds in lead time forecasts.
A new loss function speeds up sequence-to-sequence models for continuous outputs.
problem Slow and memory-intensive softmax layer limits vocabulary size and translation quality.
method Proposes a probabilistic loss and continuous embedding layer training/inference procedure.
result Models achieve up to 2.5x speed-up in training time with similar translation quality.
New method certifies images against transformations like rotations and translations.
problem Certifying robustness of images against transformations like rotations and translations.
method Randomized smoothing with three different kinds of defenses.
result Individual certificates can be obtained via statistical error bounds or efficient online inverse computation.
This paper has been withdrawn by the author due to a crucial sign error in equation 1. An isometry ρ of a connected Finsler space (M,F) is called bounded if the function d(x,ρ(x)) is bounded on M. It is called a Clifford-Wolf translation if the function d(x,ρ(x)) is constant on M. In this paper, we prove…
Improved GEC models use scored data from large pretraining to outperform.
problem Addressing data sparsity in Grammatical Error Correction.
method Derive example-level scores from a smaller, higher-quality dataset and incorporate delta-log-perplexity into training schedules.
result Models trained on scored data achieve state-of-the-art results.
A technique to quickly fix mistakes in neural networks.
problem Fixing model errors in neural networks quickly and without affecting other samples.
method Editable Training, a model-agnostic training technique.
result Effectiveness demonstrated on large-scale image classification and machine translation tasks.
New neural network processes 3D volumes with improved equivariance.
problem Improving neural network performance on 3D volumes with symmetries.
method Equivariant neural network using moving frames approach.
result Trained model outperforms benchmarks in medical volume classification.
New method uses non-translation invariant risk measures for fair financial derivative pricing.
problem Inequalities in financial derivative pricing under traditional risk measures.
method Deep reinforcement learning with modified deep hedging algorithm.
result Effective pricing of financial derivatives without price inflation.
The paper offers error bounds for quantized dynamical models.
problem Accuracy of dynamical models from dependent data sequences.
method Developed uniform error bounds for quantized models and imperfect optimization algorithms.
result Unified bounds for slow and fast rates, scaling with model encoding bits.
We propose SEARNN, a novel training algorithm for recurrent neural networks (RNNs) inspired by the "learning to search" (L2S) approach to structured prediction. RNNs have been widely successful in structured prediction applications such as machine translation or parsing, and are commonly trained using maximum likelihoo…
New algorithm computes Schrödinger Bridge for unpaired data translation.
problem Computing optimal transport maps for unpaired data translation.
method Schrödinger Bridge Flow, a discretization of a flow of path measures.
result Eliminates the need to train multiple DDM-like models.
Grammatical error correction, like other machine learning tasks, greatly benefits from large quantities of high quality training data, which is typically expensive to produce. While writing a program to automatically generate realistic grammatical errors would be difficult, one could learn the distribution of naturally…
Paper advances sparse regularisation theory for measures with new kernel insights.
problem Estimating sparse measures from noisy observations using continuous sparse regularisation.
method Develops new continuous sparse regularisation theory on measures with Beurling-LASSO, introduces kernel switch analysis.
result Proves the ``sinc-4'' kernel satisfies a technical LPC assumption for error bounds.
Deep learning models have lately shown great performance in various fields such as computer vision, speech recognition, speech translation, and natural language processing. However, alongside their state-of-the-art performance, it is still generally unclear what is the source of their generalization ability. Thus, an i…
Introduces new Wasserstein distances for more intrinsic metrics.
problem Improve metric for comparing distributions.
method Introduces RWp distances, designs algorithms for computation. result New distances are more intrinsic and computable.
Two methods generate parallel data for GEC, improving neural models' performance.
problem Lack of parallel data for GEC.
method Two approaches to generate large parallel datasets from Wikipedia data.
result Neural GEC models trained on generated corpora perform similarly and surpass state-of-the-art.
Unified framework detects natural and adversarial errors in image classifications.
problem Detecting both unintentional and intentional errors in image classifications.
method Detects errors using invariance to image transformations.
result Our approach surpasses previous methods by a large margin.
The study analyzes prediction errors in systems with memory kernels, providing bounds and stability results.
problem Prediction errors in stochastic dynamical systems with memory kernels.
method Analysis of generalized Langevin equations (GLEs) with Volterra equations, integrating synchronized noise coupling and weighted norms.
result Prediction discrepancies decay at a rate determined by the memory kernel's decay, quantitatively bounded by kernel estimation errors.
Empirical law predicts accuracy of Google Translate's translation chains.
problem Predicting accuracy in machine translation with multiple hops.
method Empirical testing of Google Translate's sequential translation.
result Accuracy decreases with the number of translating hops, following a power law.
The paper provides bounds on estimation error in a distributed online learning setting.
problem Estimating an unknown parameter in a distributed and online manner with finite sample guarantees.
method Proposes a distributed online estimation algorithm that improves accuracy through communication, providing non-asymptotic bounds on estimation error.
result Demonstrates a trade-off between estimation error and communication costs, and determines a stopping time for communication based on desired accuracy.
We present an autoencoder that leverages learned representations to better measure similarities in data space. By combining a variational autoencoder with a generative adversarial network we can use learned feature representations in the GAN discriminator as basis for the VAE reconstruction objective. Thereby, we repla…
The authors of (Cho et al., 2014a) have shown that the recently introduced neural network translation systems suffer from a significant drop in translation quality when translating long sentences, unlike existing phrase-based translation systems. In this paper, we propose a way to address this issue by automatically se…
Paper proposes a method to learn word translations bidirectionally.
problem Word translation between languages.
method Jointly learns translations in both directions with minimal supervision.
result Improves accuracy of translations over previous methods.
Sharp bounds for approximating Sobolev functions by ridge functions and networks.
problem Approximating Sobolev functions with multivariate ridge functions and networks.
method Proving sharp upper and lower bounds for approximation order.
result Order of approximation asymptotically behaves as n−r/(d−ℓ). Study on rigidity of translating hypersurfaces not in graphical direction.
problem Rigidity of translating hypersurfaces not in graphical direction.
method Proved rigidity results for complete graphical translating hypersurfaces under specific conditions.
result Entire graphical translating surfaces are flat under certain conditions.
Forward translation improves neural machine translation for sentences originally in source language.
problem Improving neural machine translation quality using synthetic data.
method Case study with French-English news translation, separating test sets by original language, analyzing domains, translationese, and noise.
result Forward translation delivers superior gains on sentences originally in source language, complementing back-translation on target language sentences.
Transformer models align words through attention weights, closely approximating Optimal Transport.
problem Understanding the internal mechanism of transformer models in language processing.
method Empirical evidence and theoretical analysis of attention weights and their relation to Optimal Transport.
result Transformer models can simulate gradient descent on the dual of entropy-regularized OT problem, providing a theoretical foundation for token alignment.
The natural automorphism group of a translation surface is its group of translations. For finite translation surfaces of genus g > 1 the order of this group is naturally bounded in terms of g due to a Riemann-Hurwitz formula argument. In analogy with classical Hurwitz surfaces, we call surfaces which achieve the maxima…
Bayesian deep learning improves seismic imaging uncertainty.
problem Uncertainty in seismic imaging due to data noise and linearization errors.
method Combines Bayesian inference and deep neural networks to quantify uncertainty in horizon tracking.
result Uncertainty in automatically tracked horizons can be quantified and visualized.
Study on stable translation lengths of surface homeomorphisms and their approximations.
problem Understanding stable translation lengths of homeomorphisms and their finite approximations.
method Comparing stable translation lengths of homeomorphisms and their finite approximations on curve graphs.
result Stable translation length of homeomorphisms with dense periodic points equals the supremum of their approximations.
Researchers classify and describe Kα-translators in Euclidean space.
problem Classifying and describing Kα-translators in Euclidean space. method Rotationally symmetric and helicoidal motions.
result For each α, there is a Kα-translator intersecting orthogonally the rotation axis. The quality of machine translation is rapidly evolving. Today one can find several machine translation systems on the web that provide reasonable translations, although the systems are not perfect. In some specific domains, the quality may decrease. A recently proposed approach to this domain is neural machine translat…
ResNet-type CNNs achieve optimal error rates in function classes with block-sparse structures.
problem Optimal approximation and estimation in function classes with sparse constraints.
method Developed ResNet-type CNNs that can approximate and estimate functions with block-sparse structures.
result ResNet-type CNNs attain minimax optimal error rates in Hölder and Barron classes.
Neural machine translation models trained for 5 South African languages.
problem Lack of resources and research for machine translation in African languages.
method Training neural machine translation models for 5 South African languages using modern techniques.
result Promises of neural machine translation for African languages.
Study measures gender bias in machine translation using multiple reference points.
problem Measuring and identifying gender bias in machine translation.
method Used an optimal non-biased translator, reference points from occupational statistics and survey.
result Found bias against both genders, but more against women, and found occupations have a greater effect than adjectives.
Classifies and constructs translators for curvature flows.
problem Understanding translating solitons in curvature flows.
method Developed rotational theory, introduced signed-neck framework.
result Classified and constructed catenoidal-type translators.
This paper uses LLMs and cycle consistency for better machine translation evaluation.
problem Evaluating translation quality and LLM capabilities without ground truth.
method Generate translation candidates, back-translate, and evaluate cycle consistency.
result Larger LLMs or more inference passes improve cycle consistency.
Neural machine translation used to convert CUDA to OpenCL.
problem Translating CUDA to OpenCL programs.
method Training input set generation, pre/post processing, case study.
result Improved accuracy in translating CUDA to OpenCL.
Paper finds a non-existence theorem for certain translators in high dimensions.
problem Non-existence of certain translators in high-dimensional spaces.
method Developed a non-existence theorem and found an example of a translator.
result Non-existence of entire Qn−1-translators in Rn+1. Constructing translating solitons from Lagrangian Grim Reapers.
problem Creating Lagrangian translating solitons from intersections of Grim Reapers.
method Desingularizing intersections with special Lagrangian Lawlor necks.
result Constructing Lagrangian translating solitons with multiple ends and loops.
Improved speech recognition with cumulative adaptation methods.
problem Robust speech recognition in varying environments and speakers.
method Used a bidirectional LSTM neural network and i-vectors for adaptation.
result Achieved 13% relative improvement in word error rate.