Improved spoken English intelligibility with computer recognition and feature extraction.
problem Improving spoken English pronunciation and intelligibility.
method Automatic speech recognition using PocketSphinx alignment and feature extraction with SVM classifier probability prediction.
result SVM models achieve 82 percent agreement with human transcriptions, up from 75 percent.
This research explores the effects of various training settings on a Polish to English Statistical Machine Translation system for spoken language. Various elements of the TED, Europarl, and OPUS parallel text corpora were used as the basis for training of language models, for development, tuning and testing of the tran…
This research explores the effects of various training settings from Polish to English Statistical Machine Translation system for spoken language. Various elements of the TED parallel text corpora for the IWSLT 2013 evaluation campaign were used as the basis for training of language models, and for development, tuning …
End-to-end Sanskrit TTS developed with limited data, achieving good quality.
problem Developing natural-sounding speech for Sanskrit with scarce data.
method Fine-tuning Tacotron2 model with WaveGlow and transfer learning.
result Achieved an overall MOS of 3.38 from 37 evaluators.
Enhances spoken speech quality using EEG signals.
problem Improves speech clarity in noisy environments.
method Generative adversarial network (GAN), gated recurrent unit (GRU), temporal convolutional network (TCN) regression models.
result Significant improvement in speech enhancement quality compared to traditional methods.
System diagnoses Alzheimer's disease from spoken language using multi-modal features.
problem Early diagnosis of Alzheimer's disease from spoken language.
method Classification system based on spoken language using three approaches (N-gram, i-vector, x-vector).
result Accuracy of 83.6% on the cookie picture description task from Pitt Corpus dementia bank.
ClovaCall introduces a new Korean call speech corpus for contact centers.
problem Lack of large-scale call-based speech corpora for Korean dialog scenarios.
method Development of a new large-scale Korean call-based speech corpus (ClovaCall) in a restaurant reservation domain.
result Validation of the dataset with ASR models shows its effectiveness.
Expanding spoken language understanding to handle complex entities and intents.
problem Handling compound entities and intents in spoken language understanding.
method Introducing a domain-agnostic shallow parser that handles linguistic coordination, learning domain-independent and slot-independent features.
result The model learns to segment conjunct boundaries of various phrasal categories and improves generalization across different slot types using adversarial training.
New framework learns phoneme metric from perception data.
problem Learn metric for phoneme similarity from data.
method Learning algorithms to derive metric from behavioral data.
result Framework outperforms previous metrics in phoneme prediction.
Benchmark tests spoken language models for infant language learning.
problem Understanding how infants learn language from speech.
method Developed a language-acquisition-friendly benchmark.
result Benchmarking shows models' strengths and weaknesses.
End-to-end neural network improves QbE-STD in multilingual speech search.
problem Query by example spoken term detection in zero-resource scenarios.
method Use of multilingual bottleneck features and CNN for pattern matching, integrated into a fully neural network framework.
result CNN-based matching outperforms DTW-based matching using bottleneck features.
Alternative method proposed for handling uncertainty in i-vector extraction.
problem Uncertainty in i-vector extraction for spoken language recognition.
method Proposes an alternative method to propagate uncertainty into the Gaussian back-end.
result Alternative method effectively handles uncertainty in i-vector extraction.
MPSA-DenseNet improves accent classification accuracy.
problem Accurate English accent identification.
method Combines multi-task learning and PSA attention mechanism with DenseNet.
result MPSA-DenseNet outperforms other models in accent classification.
Improves domain classification across multiple locales with shared language.
problem Improves domain classification accuracy in Spoken Language Understanding across multiple locales with shared language.
method Selective multi-task learning to create a joint representation of utterances over locales with different sets of domains.
result The proposed approach outperforms other baselines models especially when classifying locale-specific domains and low-resourced domains.
ACI converts call center conversations into actionable data.
problem Real-time spoken language understanding for call center conversations.
method Combines speech recognition, entity and intent recognition, and a business rules engine.
result ACI converts live audio into structured events for real-time supervision and assistance.
Improved ASR for English-isiZulu code-switched speech with semi-supervised training.
problem Improving ASR for code-switched speech between English and isiZulu.
method Semi-supervised training using automatic transcription of multilingual speech data.
result Semi-supervised training achieved significant WER reduction in ASR performance.
New model learns coupled representations for domains, intents, and slots.
problem Representation learning for domains, intents, and slots in spoken language understanding.
method Proposes a model that learns coupled representations by aggregating slot and intent representations based on their hierarchical relationships.
result Improved performance on contextual cross-domain reranking task.
Paper explores unsupervised transfer learning for SLU, improving model performance with unlabeled data.
problem Improving SLU model performance with limited labeled data.
method Uses ELMo embeddings for unsupervised pre-training and ELMo-Light for faster pre-training. Combines unsupervised and supervised transfer techniques.
result Unsupervised pre-training on unlabeled data significantly improves SLU performance, even outperforming conventional supervised transfer.
Spoken language translation (SLT) has become very important in an increasingly globalized world. Machine translation (MT) for automatic speech recognition (ASR) systems is a major challenge of great interest. This research investigates that automatic sentence segmentation of speech that is important for enriching speec…
Paper explores using hand gestures to type English letters.
problem Creating easy-to-remember, non-cumbersome gestures for typing.
method Statistical approach to handle randomness in hand movements.
result Achieved 97.33% accuracy with entire English alphabet.
Optimizes dialogue success and length using multi-objective reinforcement learning.
problem Balancing multiple reward components in spoken dialogue systems.
method Structured multi-objective reinforcement learning to find optimal reward weights.
result Optimized reward weights significantly improve dialogue performance across six domains.
Neural User Simulator outperforms traditional ABUS in training dialogue systems.
problem Limited diversity and lack of natural language in ABUS.
method NUS learns user behavior from a corpus and generates natural language.
result NUS trained policies outperform ABUS in real user evaluations.
RUSLAN is a large Russian speech corpus for text-to-speech.
problem Lack of high-quality annotated Russian speech data for text-to-speech.
method Developed a large annotated Russian speech corpus and trained a neural network for text-to-speech synthesis.
result Synthesized speech quality evaluated with MOS scores: 4.05 for naturalness, 3.78 for intelligibility.
Study examines bias in language models across multiple languages.
problem Assessing bias in language models across different languages.
method Semi-automatically translated data sets into multiple languages, analyzed mono- and multilingual models.
result Notable differences in bias across languages, with Turkish models showing least stereotypes.
PBN combines generative and discriminative capabilities in a neural network.
problem Combining generative and discriminative capabilities in neural networks.
method Convolutional PBN, sharing FF-NN embodiment, combining generative and discriminative qualities.
result PBN shows excellent qualities from either generative or discriminative viewpoint.
Improves E2E ASR performance on numeric sequences with additional training data and denormalization.
problem Challenges in recognizing numeric sequences out-of-vocabulary in ASR systems.
method Uses text-to-speech for additional numeric training data and a small-footprint neural network for denormalization.
result Reduction of WER by up to a factor of 8 in the longest numeric sequences.
The goal of this modern presentation, followed by an English translation from the German, is to make available some parts of Lie's very systematic mathematical thought which deserve to join the contemporary literature, and above all also, to be read.
Evolved Transformer improves on Transformer architecture for language tasks.
problem Improving Transformer architecture for sequence tasks.
method Evolutionary architecture search with warm starting and dynamic resource allocation.
result Evolved Transformer achieves state-of-the-art BLEU scores and reduces parameter count.
The paper improves dialogue quality estimation using a novel user satisfaction model.
problem Improving dialogue quality estimation in spoken dialogue systems.
method Proposes a novel user satisfaction estimator based on BiLSTMs and reinforcement learning.
result The novel user satisfaction estimator outperforms previous models in terms of user satisfaction and task success.
In this paper, we attempt to improve Statistical Machine Translation (SMT) systems on a very diverse set of language pairs (in both directions): Czech - English, Vietnamese - English, French - English and German - English. To accomplish this, we performed translation model training, created adaptations of training sett…
Memory-augmented neural networks improve machine translation performance.
problem Improving machine translation accuracy and flexibility.
method Evaluation of Neural Turing Machines and Differentiable Neural Computers for machine translation tasks.
result Memory-augmented neural networks perform similarly to attentional encoders on Vietnamese to English tasks but have lower BLEU scores on Romanian to English tasks.
In this paper, we present Neural Phrase-based Machine Translation (NPMT). Our method explicitly models the phrase structures in output sequences using Sleep-WAke Networks (SWAN), a recently proposed segmentation-based sequence modeling method. To mitigate the monotonic alignment requirement of SWAN, we introduce a new …
Paper shows continuous speech recognition with EEG features, no speech input.
problem Continuous speech recognition with limited vocabulary and noisy/no speech input.
method Connectionist temporal classification (CTC) model, EEG features, new deep learning architecture.
result Continuous speech recognition achieved on limited vocabulary with noisy/no speech input.
This study improves NMT using reinforcement learning, overcoming its instability.
problem Stability issues in reinforcement learning for neural machine translation.
method Systematic study on reinforcement learning factors and a new method for monolingual data.
result Competitive results on WMT17 Chinese-English translation task, setting a state-of-the-art performance.
Proposes a multilingual email segmentation benchmark and model.
problem Lack of multilingual email zoning corpora and models.
method Analysis of existing corpora, development of multilingual benchmark, introduction of OKAPI model.
result OKAPI model achieves state-of-the-art performance in English and generalizes well to unseen languages.
Hybrid ASR systems can model graphemes effectively using chenones, outperforming traditional methods.
problem Traditional hybrid ASR systems struggle with English's poor grapheme-phoneme correspondence.
method Leveraging tied context-dependent graphemes (chenones) to model graphemes directly.
result Chenone-based systems significantly outperform senone baselines by 4.5% to 11.1% on English datasets.
We show that the predictability of letters in written English texts depends strongly on their position in the word. The first letters are usually the least easy to predict. This agrees with the intuitive notion that words are well defined subunits in written languages, with much weaker correlations across these units t…
Tensor Train layer improves BLEU scores in NMT models.
problem Improving Neural Machine Translation (NMT) models' performance.
method Implemented Tensor Train layer in TensorFlow for NMT training.
result Higher learning rates and more 'rectangular' core dimensions improve BLEU scores.
Continuous speech recognition from brain activity without vocalization.
problem Recognizing silent speech from EEG signals.
method Implemented a CTC ASR model using EEG signals.
result Demonstrated feasibility of EEG for continuous silent speech recognition.
This paper closes the accuracy gap between A2W models and sub-word models using English conversational speech data.
problem Closing the accuracy gap between direct acoustics-to-word models and sub-word models.
method Training an A2W model with orders of magnitude more data, optimizing model initialization, training data order, and regularization.
result Achieved word error rates of 8.8%/13.9% on Hub5-2000 Switchboard/CallHome test sets.
WEEND uses a neural network to recognize speech and assign speakers to words.
problem End-to-end neural diarization without additional ASR and orchestration.
method Multi-task learning with an auxiliary network for ASR and speaker diarization.
result WEEND outperforms turn-based diarization and can handle 5-minute audio.
RNN models classify poem meters from plain text with high accuracy.
problem Classifying poem meters from plain text.
method Character-level encoding, RNN models, no feature handcrafting.
result 96.38% accuracy for Arabic poems, 82.31% for English poems.
Graphemes outperform phonemes in end-to-end models for English Voice-search and multi-dialect tasks.
problem Comparing phoneme-based and grapheme-based sub-word units in end-to-end models.
method Detailed experiments comparing phoneme-based and grapheme-based end-to-end models on large vocabulary English Voice-search and multi-dialect tasks.
result Graphemes outperform phonemes in end-to-end models for English Voice-search and multi-dialect tasks.
Dataset for measuring reading levels in India's children.
problem Measuring reading levels in India's vast population.
method Developed ASER dataset with 5,301 subjects in Hindi, Marathi, and English.
result Achieved 86% accuracy in English language classification.
We describe a unified and coherent syntactic framework for supporting a semantically-informed syntactic approach to statistical machine translation. Semantically enriched syntactic tags assigned to the target-language training texts improved translation quality. The resulting system significantly outperformed a linguis…
This is the less official, English version of the proof of the fact that every closed atoroidal 3-manifold carries finitely many isotopy classes of tight contact structures.
Comment classification on cookery channels using BERT and traditional models.
problem Volume of multilingual comments, variable lengths, slang, symbols, and abbreviations make comment classification challenging.
method Evaluated traditional machine learning models (Naive Bayes, KNN, SVM, Random Forest, Decision Trees) and BERT-based models (BERT, DISTILBERT, XLM) for multilingual comment classification.
result XLM was the top-performing BERT model with an accuracy of 67.31, while Random Forest with Term Frequency Vectorizer was the best traditional model with 63.59 accuracy.
Project develops lip reading algorithm for limited English.
problem Limited audio information for lip reading.
method Extract lip positions from video frames, classify visemes and phonemes, use HMMs to predict words.
result Algorithm predicts words from lip movements for a subset of English.