We present Listen, Attend and Spell (LAS), a neural network that learns to transcribe speech utterances to characters. Unlike traditional DNN-HMM models, this model learns all the components of a speech recognizer jointly. Our system has two components: a listener and a speller. The listener is a pyramidal recurrent ne…
End-to-end ASR system uses context n-grams for better speech recognition.
problem Contextual information impacts speech recognition accuracy.
method Jointly optimizes ASR components with context embeddings during inference.
result Proposed CLAS system outperforms traditional methods by 68% relative WER.
End-to-end model detects articulatory features from speech data.
problem Detecting articulatory features from speech data for various applications.
method Apply Listen, Attend and Spell (LAS) architecture and attention models.
result End-to-end training of manners and places of articulation detectors.
Having a sequence-to-sequence model which can operate in an online fashion is important for streaming applications such as Voice Search. Neural transducer is a streaming sequence-to-sequence model, but has shown a significant degradation in performance compared to non-streaming models such as Listen, Attend and Spell (…
AV-ASR system improves speech recognition with visual context.
problem Improving speech recognition accuracy with visual information.
method Transformer-based architecture with multiresolution and multimodal training.
result Multiresolution training speeds up convergence and improves WER by 18%.
SpecAugment improves speech recognition with simple feature augmentation.
problem Improving automatic speech recognition accuracy.
method Applying warping, frequency channel masking, and time step masking to feature inputs of neural networks.
result Achieved state-of-the-art performance on LibriSpeech and Switchboard tasks.
Attention-based encoder-decoder architectures such as Listen, Attend, and Spell (LAS), subsume the acoustic, pronunciation and language model components of a traditional automatic speech recognition (ASR) system into a single neural network. In previous work, we have shown that such architectures are comparable to stat…
A transformer model improves spell correction with hierarchical attention.
problem Improving spell correction accuracy and speed.
method Multi encoder-single decoder transformer architecture with hierarchical attention.
result Significant improvement in CER, WER, and SER error rates.
Seq2seq ASR adapts to speakers, improving performance by 25%.
problem Speaker adaptation for seq2seq ASR systems to match conventional methods.
method Applied Kullback-Leibler divergence and Linear Hidden Network adaptation to seq2seq models.
result 25% relative word error rate improvement with seq2seq model adaptation.
Real-time spell checker adapts to new languages.
problem No real-time, language-adaptable spell checkers for non-English languages.
method Used Wikipedia and subtitles data to generate dictionaries, created noisy channel datasets, compared with industry tools.
result System performs well across 24 languages, outperforming existing tools.
This study analyses the duration dependence of events that trigger volatility persistence in stock markets. Such events, in our context, are monthly spells of contiguous price decline or negative returns for the S&P500 stock market index over the last 145 years. Factors known to affect the duration of these spells are …
Novel approach to estimate P300 BCI efficiency using SNR.
problem Improving the accuracy of P300 BCI for severely disabled people.
method Introduced a novel approach considering P300 SNR for estimating efficiency, using a Gaussian noise model.
result P300 SNR significantly correlates with spelling accuracy, improving BCI efficiency.
Optimizes explanations for better listener understanding.
problem Insufficient consideration of listener preferences in concept-based explanations.
method Iterative training procedure based on direct preference optimization.
result Pragmatic explanations improve both model accuracy and user understanding.
We further study the incidence relations that arise from the various subtowers, known as Baby Monster, which exist within the R3-Monster Tower. This allows us to complete the RVT class spelling rules. We also present a method of calculating the various Baby Monster that appear within the Monster Tower.
Synthetic noise training improves machine translation robustness to spelling mistakes.
problem Making machine translation robust to spelling mistakes and natural noise.
method Training on synthetic noise to improve robustness to natural noise.
result Training on synthetic noise improves robustness to natural noise without diminishing performance on clean text.
This research investigates reliable local explanations for machine listening models.
problem Generating reliable local explanations for machine listening models.
method Investigates the sensitivity of SoundLIME explanations to input perturbations and proposes a novel method for identifying suitable content types.
result SoundLIME explanations are sensitive to the content in occluded input regions, and the average magnitude of input mel-spectrogram bins is the most suitable content type for temporal explanations.
Preschool attendance correlates with lower developmental vulnerabilities in Queensland, Australia.
problem Understanding the relationship between preschool attendance and developmental vulnerabilities in different regions.
method Data Analysis and Machine Learning to identify clusters of socio-demographic variables.
result Identified three clusters with varying socio-demographic variables affecting the relationship between preschool attendance and developmental vulnerabilities.
Predicting event attendance using social influence from social networks.
problem Predicting people's participation in real-world events.
method Modeling social influence, using non-geotagged posts and social group structures, applying graph embedding techniques, and training a neural network.
result The proposed classifier achieves 89% accuracy on the VFestival dataset, outperforming state-of-the-art methods.
OBJECTIVE: We aim to extract and denoise the attended speaker in a noisy, two-speaker acoustic scenario, relying on microphone array recordings from a binaural hearing aid, which are complemented with electroencephalography (EEG) recordings to infer the speaker of interest. METHODS: In this study, we propose a modular …
Chatbot uses BERT to handle financial investment questions, improving accuracy and decision-making.
problem Improving accuracy and decision-making in financial investment customer service.
method Deep Bidirectional Transformer (BERT) model, uncertainty measure comparison, mixed-integer programming, automatic spelling correction.
result Chatbot can recognize 381 intents and decide when to escalate questions.
Deep learning improves singing processing tasks.
problem Lack of data and computing resources for singing processing.
method State-of-the-art deep learning techniques.
result Advances in accuracy and sound quality.
Area attention allows dynamic area-based attention in memory.
problem Fixed attention granularity limits model performance.
method Area attention dynamically determines area shape and size via learning.
result Area attention improves performance on neural machine translation and image captioning.
CNN improves spatiotemporal emotion recognition from EEG during music listening.
problem Improving emotion recognition from EEG signals during music listening.
method Conducted a study on CNN and its spatiotemporal feature extraction for emotion recognition.
result CNN outperforms SVM in leave-one-subject-out cross validation.
Industrial-scale podcast recommender system optimizes long-term listening journeys.
problem Optimizing long-term listening experiences in podcast recommendation systems.
method Reinforcement learning approach to optimize user listening journeys over months.
result Significantly improved long-term performance in A/B tests compared to short-term metrics.
OtoWorld helps agents learn to navigate by listening in interactive environments.
problem Training agents to navigate using auditory information.
method Interactive environment with agents learning to listen and navigate.
result Agents can win at OtoWorld, a navigation game with sound sources.
We present Sequential Attend, Infer, Repeat (SQAIR), an interpretable deep generative model for videos of moving objects. It can reliably discover and track objects throughout the sequence of frames, and can also generate future frames conditioning on the current frame, thereby simulating expected motion of objects. Th…
Trains word embeddings from music and text data to link music contexts.
problem Varying vocabulary size and musical relevance in word embeddings.
method Combines general text and music-specific data to train word embeddings.
result Trained embeddings better associate music contexts with compositions.
Winterization of Texas power system profitable but risky, estimated at $11.74bn over 30 years.
problem Profitability and risk of winterizing Texas power system infrastructure.
method Combined temperature-dependent load and outage estimates over 71 years of climate data.
result Large-scale winterization of gas infrastructure and power plants is profitable, but risks are high due to low-frequency of cold spells.
Unified ML approach predicts ED attendances with high accuracy.
problem Managing hospital demand at emergency departments efficiently.
method Ensemble of time series and machine learning approaches with hyperparameter tuning.
result Predictions with mean absolute error of +/- 14 and +/- 10 patients, MAE of 6.8% and 8.6%.
Improves diversity of text-to-image models without sacrificing FID.
problem Lack of diversity and tendency to recreate training set images.
method Adds sparse repellency terms to diffusion SDE to guide trajectories away from a reference set.
result Improves diversity of diffusion models with minimal impact on FID.
Podcast recommendations improved by analyzing user listening paths.
problem Challenges in recommending podcasts effectively.
method Analyzes user listening paths as sequential trajectories for recommendations.
result 450% increase in effectiveness over baseline.
End-to-end DA method for domain-invariant CNNs using parallel audio recordings.
problem Distribution mismatches between training and application data in machine listening.
method Enforcing equal hidden layer representations for domain-parallel samples.
result Learn domain-invariant classifiers without requiring classification labels.
The Monster tower, also known as the Semple tower, is a sequence of manifolds with distributions of interest to both differential and algebraic geometers. Each manifold is a projective bundle over the previous. Moreover, each level is a fiber compactified jet bundle equipped with an action of finite jets of the diffeom…
Population-based learning improves representation of unstructured data.
problem Improving representation of unstructured data.
method Instantiating Lewis signaling games within a population of agents.
result Population-based learning produces better representations than single-agent learning.
CNNs improve generalization to unseen audio devices with increased width, not depth.
problem CNNs are sensitive to specific audio recording devices in acoustic scene classification.
method Investigated the relationship between over-parameterization and generalization in CNNs for audio classification.
result Increasing width improves generalization to unseen devices without increasing the number of parameters.
DSA improves sentence embedding by dynamically attending to words.
problem Efficiently capturing the importance of words in sentences for embedding.
method DSA modifies dynamic routing from capsule networks for self-attention in sentences.
result DSA achieves state-of-the-art results in SNLI with fewer parameters.
Grapheme ASR improves with G2G model that corrects spelling errors.
problem Rare long-tail words in non-phonemic languages like English.
method Train G2G model on text-to-speech data to rewrite character sequences into phonetically consistent forms.
result Reduces Word Error Rate by 3% to 11% over a strong graphemic baseline.
We review (non-abelian) extensions of a given Lie algebra, identify a 3-dimensional cohomological obstruction to the existence of extensions. A striking analogy to the setting of covariant exterior derivatives, curvature, and the Bianchi identity in differential geometry is spelled out. In the new version references ad…
Deep reinforcement learning (DRL) has shown incredible performance in learning various tasks to the human level. However, unlike human perception, current DRL models connect the entire low-level sensory input to the state-action values rather than exploiting the relationship between and among entities that constitute t…
Paper shows how LSTM can remember long sequences by attending to persisted information.
problem LSTMs struggle with long sequences due to fading information and bias towards recent data.
method The paper introduces a mechanism that allows LSTMs to attend to information in memory based on how long it was persisted by the gating mechanism.
result The method improves LSTM's ability to process long sequences by retrieving information proportionally to its persistence in memory.
GA-Net selectively attends to parts of a sequence for text classification.
problem Inefficient global attention mechanisms on long sequences.
method Gated Attention Network (GA-Net) using an auxiliary network to dynamically select and attend to important parts of the sequence.
result GA-Net achieves better performance with less computation and interpretability.
Discrete-AIR model identifies objects in images with interpretable latent codes.
problem Identifying objects in images without labeled data.
method Recurrent Auto-Encoder with structured latent distributions for discrete, continuous, and spatial attention.
result Discrete-AIR model uses minimal latent variables for efficient inference.
These notes were originally written for the Stochastic Analysis Seminar in the Department of Operations Research and Financial Engineering at Princeton University, in February of 2011. The seminar was attended and supported by members of the Research Training Group, with the author being partially supported by NSF gran…
Enhances speech quality in noisy environments using symbolic sequential modeling.
problem Improving speech quality in noisy conditions.
method Incorporates symbolic sequential modeling into speech enhancement framework.
result Significant improvement in speech quality metrics (PESQ, STOI) on TIMIT dataset.
Recent studies in the field of human vision science suggest that the human responses to the stimuli on a visual display are non-deterministic. People may attend to different locations on the same visual input at the same time. Based on this knowledge, we propose a new stochastic model of visual attention by introducing…
SurvBESA predicts survival times using ensemble methods with self-attention.
problem Challenges in survival analysis due to censored data and unstable predictions.
method SurvBESA combines Beran estimators with a self-attention mechanism to predict survival times.
result SurvBESA outperforms state-of-the-art models in predicting survival times.
From a machine learning perspective, the human ability localize sounds can be modeled as a non-parametric and non-linear regression problem between binaural spectral features of sound received at the ears (input) and their sound-source directions (output). The input features can be summarized in terms of the individual…
New neural network layer handles OOV words in NLP tasks without pre-training.
problem Handling out-of-vocabulary words in natural language processing.
method Contextual-compositional neural network layer that attends to character sequence and context.
result Improves performance on 23 languages in joint tagging tasks.