New framework learns phoneme metric from perception data.
problem Learn metric for phoneme similarity from data.
method Learning algorithms to derive metric from behavioral data.
result Framework outperforms previous metrics in phoneme prediction.
Representation mixing combines character and phoneme inputs for flexible TTS synthesis.
problem Limited control over pronunciation in character or phoneme-based TTS systems.
method Representation mixing combines multiple linguistic inputs in a single encoder.
result Flexibility in choosing between character, phoneme, or mixed representations during inference.
Neural model detects phoneme boundaries from speech, outperforming baselines.
problem Phoneme boundary detection for speech processing applications.
method Learnable segmental features with a structured loss function.
result Model achieves state-of-the-art performance on TIMIT and Buckeye corpora.
Study investigates predictive coding models for phonemic learning.
problem Understanding how predictive coding models generalize to different languages and dataset sizes.
method Investigated Autoregressive Predictive Coding and Contrastive Predictive Coding models in phoneme discrimination tasks for two languages with varying dataset sizes.
result Contrastive Predictive Coding model converges rapidly and outperforms Autoregressive Predictive Coding on both languages.
End-to-end transformer model improves lexical stress detection accuracy.
problem Inaccurate phoneme boundaries and limited features for stress classification.
method End-to-end sequence to sequence model using transformer trained on feature sequences and phoneme sequences with stress marks.
result End-to-end model achieves better performance and lower phoneme error rate (6.36%) compared to syllable segmentation methods.
For decades, context-dependent phonemes have been the dominant sub-word unit for conventional acoustic modeling systems. This status quo has begun to be challenged recently by end-to-end models which seek to combine acoustic, pronunciation, and language model components into a single neural network. Such systems, which…
Unsupervised speech recognition without labeled data using novel cost function and MAP refinement.
problem Training speech recognition systems without labeled data.
method Alternates between phoneme classifier learning and boundary refinement using Segmental Empirical Output Distribution Matching and MAP approach.
result Achieves phone error rate (PER) of 41.6% on TIMIT dataset.
KS-algebra consists of expressions constructed with four kinds operations, the minimum, maximum, difference and additively homogeneous generalized means. Five families of Z-classifiers are investigated on binary classification tasks between English phonemes. It is shown that the classifiers are able to reflect well…
Self-supervised model detects phoneme boundaries without annotations.
problem Unsupervised phoneme segmentation without manual annotations.
method Convolutional neural network trained with Noise-Contrastive Estimation.
result Model outperforms baselines on TIMIT and Buckeye corpora.
The goal of this project is to develop a limited lip reading algorithm for a subset of the English language. We consider a scenario in which no audio information is available. The raw video is processed and the position of the lips in each frame is extracted. We then prepare the lip data for processing and classify the…
We replace the Hidden Markov Model (HMM) which is traditionally used in in continuous speech recognition with a bi-directional recurrent neural network encoder coupled to a recurrent neural network decoder that directly emits a stream of phonemes. The alignment between the input and output sequences is established usin…
WaveNet reconstructs speech from brain activity, revealing acoustic features.
problem Reconstructing speech from brain activity with limited data.
method WaveNet model applied to STG intracranial recordings.
result WaveNet models reveal phoneme-level acoustic features.
We stabilize the activations of Recurrent Neural Networks (RNNs) by penalizing the squared distance between successive hidden states' norms. This penalty term is an effective regularizer for RNNs including LSTMs and IRNNs, improving performance on character-level language modeling and phoneme recognition, and outperfor…
We use automatic speech recognition to assess spoken English learner pronunciation based on the authentic intelligibility of the learners' spoken responses determined from support vector machine (SVM) classifier or deep learning neural network model predictions of transcription correctness. Using numeric features produ…
In this project we further investigate the idea of reducing the dimensionality of datasets using a Borel isomorphism with the purpose of subsequently applying supervised learning algorithms, as originally suggested by my supervisor V. Pestov (in 2011 Dagstuhl preprint). Any consistent learning algorithm, for example kN…
PCI combines perception and control using Bayesian inference with object-based representations.
problem Separate perception and control in reinforcement learning.
method Joint Perception and Control as Inference (PCI) framework with Object-based Perception Control (OPC).
result OPC achieves good perceptual grouping quality and outperforms baselines in accumulated rewards.
This paper benchmarks speech LVMs against deterministic models and adapts a video model to speech.
problem Speech generation models are inferior to deterministic models.
method Developed a speech benchmark of LVMs and compared them against deterministic models.
result The Clockwork VAE outperforms previous LVMs and reduces the gap to deterministic models.
Study the tradeoff between signal distortion and human perception over finite channels.
problem Characterize the distortion-perception tradeoff for finite channels with arbitrary metrics.
method Solve linear programming problems to compute the distortion-perception function and optimal reconstructions.
result DP function is piecewise linear in the perception index.
Recurrent sequence generators conditioned on input data through an attention mechanism have recently shown very good performance on a range of tasks in- cluding machine translation, handwriting synthesis and image caption gen- eration. We extend the attention-mechanism with features needed for speech recognition. We sh…
A framework isolates VQA reasoning from perception for better model evaluation.
problem Improper separation of visual perception and reasoning in VQA models.
method Introducing a framework and a top-down calibration technique to decouple reasoning from perception.
result Improved evaluation of VQA models by separating reasoning from perception.
End-to-end autonomous driving perception learns latent features for better performance.
problem Current autonomous driving systems are complex and require human engineering.
method Sequential latent representation learning for end-to-end perception.
result End-to-end perception model solves detection, tracking, localization, and mapping problems.
Geometric model explains music perception combining neuroscience and acoustics.
problem Rationalize and predict psycho-acoustic phenomena in music perception.
method Combining neuroscientific theories with acoustic observations, a geometric model of the space of all chords is created.
result The geometric model allows for rigorous studies of psychoacoustic quantities like roughness and harmonicity.
Safe control for vehicles using learned perception from images.
problem Controlling autonomous vehicles with partial state information from images.
method Learned perception map and safe set design for a closed loop system.
result Generalization properties of the perception-control loop are favorable.
Novel BCI system classifies imagined speech with high accuracy.
problem Classifying imagined speech from brain signals.
method Hierarchical deep learning with CNN and autoencoder.
result Achieved 83.42% average accuracy across six phonological tasks.
RETR improves indoor radar perception with a novel transformer model.
problem Indoor radar perception lacks models tailored for multi-view radar settings.
method RETR extends DETR architecture with depth-prioritized feature similarity, tri-plane loss, and radar-to-camera transformation.
result RETR outperforms state-of-the-art methods by 15.38+ AP for object detection and 11.91+ IoU for instance segmentation.
Grapheme ASR improves with G2G model that corrects spelling errors.
problem Rare long-tail words in non-phonemic languages like English.
method Train G2G model on text-to-speech data to rewrite character sequences into phonetically consistent forms.
result Reduces Word Error Rate by 3% to 11% over a strong graphemic baseline.
Researchers study how teachers' advising relationships influence their perceptions of satisfaction and students, not policy influence.
problem Understanding the relationship between teachers' advising relationships and their perceptions of satisfaction and students.
method Proposed a novel joint model of network and item responses (JNIRM) with correlated latent variables.
result Teachers' advising relationships contribute more to satisfaction and students than to influence over educational policies.
Recently, the connectionist temporal classification (CTC) model coupled with recurrent (RNN) or convolutional neural networks (CNN), made it easier to train speech recognition systems in an end-to-end fashion. However in real-valued models, time frame components such as mel-filter-bank energies and the cepstral coeffic…
Study how actions affect perception in embodied agents using group theory.
problem Understanding how actions influence perception in autonomous agents.
method Mathematical formalism of group theory applied to sensory commutativity of action sequences.
result Introduced Sensory Commutativity Probability (SCP) to measure action effects on perception.
Reconstruction-based learning produces uninformative features for perception tasks.
problem Misalignment between reconstruction-based learning and perception tasks.
method Investigated the impact of input space reconstruction on feature learning for perception tasks.
result Reconstruction-based learning allocates model capacity to a subspace with uninformative features for perception tasks.
Paper extends ICA to ISA with auxiliary variables for better speech representation learning.
problem Learning unsupervised speech representations with independent subspaces.
method Theoretical framework of nonlinear ISA with auxiliary variables.
result Proposes an algorithm to learn speech representations with independent subspaces.
Study uses deep learning to predict gender and analyze HPV vaccine perceptions on Twitter.
problem Analyzing gender differences in public perceptions on HPV vaccine using social media data.
method Convolutional neural network model trained on Twitter text for gender prediction, then applied to HPV vaccine related tweets.
result Identified gender differences in public perceptions on HPV vaccine, consistent with previous studies.
The study estimates how changing words in sentences affects audience perception.
problem Estimating the causal effect of lexical choice on audience perception.
method Two classes of methods: quasi-experimental designs and classification problems.
result Algorithmic estimates align with randomized-control trials and can be transferred across domains.
Tensor models decode human perception and memory using SPO triples.
problem Understanding implicit and explicit perception and memory in the brain.
method Tensor models with SPO triples, dual representations, and four layers.
result Semantic memory is crucial for explicit perception and declarative memories.
Approach to develop visual perception in robots through sensorimotor interactions.
problem Developing autonomous perception in robots.
method Sensorimotor contingencies theory applied to robot exploration and learning.
result Captured sensorimotor regularities in a predictive model for visual field discovery.
Study examines perceptions and attitudes about breast cancer on Twitter.
problem Understanding public perceptions and attitudes towards breast cancer on social media.
method Identified and collected tweets, used topic modeling and sentiment analysis.
result Identified themes and quantified users' perceptions and emotions about breast cancer.
PeL separates sensory interface optimization from decision learning.
problem Optimizing sensory interfaces without task-specific information.
method Formal separation of perception and decision learning, using metrics for stability, informativeness, and geometry.
result Updates preserving invariants are orthogonal to decision gradients.
Study investigates how simple speech sounds can form abstract categories.
problem How do abstract categories like phonemes emerge from speech exposure?
method Used modeling techniques to test Memory-Based Learning and Error-Correction Learning.
result Error-Correction Learning models can learn abstractions, identifying phone inventory and grouping.
Logical scaffolds enhance AI software quality.
problem Improving AI component quality in software.
method Logical scaffolds as a method to improve AI components.
result Logical scaffolds can improve AI beyond perception systems.
PERCEPT detects changes in high-dimensional data streams using topological data analysis.
problem Detecting changes in high-dimensional data streams, especially when embedded in a low-dimensional space.
method Leverages topological data analysis to learn embedded topology as a point cloud via persistence diagrams, then applies non-parametric monitoring for detecting changes.
result Demonstrates efficient detection of online changes from high-dimensional data streams.
While perception tasks such as visual object recognition and text understanding play an important role in human intelligence, the subsequent tasks that involve inference, reasoning and planning require an even higher level of intelligence. The past few years have seen major advances in many perception tasks using deep …
Default-ERM shortcut learning persists even without additional information.
problem Default-ERM shortcut learning in perception tasks despite stable feature sufficiency.
method Studied linear perception task; developed margin control (MARG-CTRL) loss functions.
result Margin control mitigates shortcut learning on various tasks.
DGP learns speech recognition by modeling complex relationships between utterances.
problem Modeling complex relationships in speech recognition without relational data.
method Bayesian nonparametric deep learning method (DGP) that generates infinite probabilistic graphs.
result DGP successfully infers relationships among utterances without relational data during training.
The study improves deep learning models for safer autonomous vehicles.
problem Robustness of deep neural network models in autonomous driving.
method Analyzes and proposes solutions for deep learning model robustness.
result Enhanced deep learning models for safer autonomous vehicles.
Paper tackles robust spatial perception by handling outliers efficiently.
problem Robust spatial perception is challenged by incorrect data association (outliers).
method Proposes adaptive trimming algorithm to remove outliers efficiently.
result Adaptive trimming algorithm outperforms state-of-the-art methods across applications.
Unsupervised learning of speech representations using WaveNet autoencoders.
problem Extract meaningful latent representations of speech signals.
method Applying autoencoding neural networks to speech waveforms, using a high capacity WaveNet decoder, and comparing three variants of latent representations.
result Comparable performance with top entries in the ZeroSpeech 2017 unsupervised acoustic unit discovery task.
New method allows a robot to perceive space dimensions without prior knowledge.
problem Limitation of previous methods in perceiving space dimensions with small movements.
method Non-linear dimension estimation method.
result Robots can now perceive space dimensions with larger movements.
A technique finds adversarial examples for deep neural networks using human perception.
problem Finding adversarial examples for deep neural networks without access to internal structure.
method Covariance Matrix Adaptation Evolution Strategy (CMA-ES) with perception-in-the-loop.
result CMA-ES can find adversarial examples with human feedback, showing favorable performance.