Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

491317 · Apr 201919922001200920182026
48 results for TIMIT corpus

Improved phone classification accuracy using graph-based regularization.

problem Phone classification with limited labeled data.
method Graph-based semi-supervised learning with stochastic entropic regularization.
result Significantly improved phone classification accuracy with low labeled data.

Bayesian SHMM discovers acoustic units from unlabeled speech.

problem Discovering language-specific acoustic units from unlabeled speech.
method Bayesian Subspace Hidden Markov Model (SHMM) trained on labeled data to find new acoustic units on target language.
result Significantly outperforms previous HMM-based systems and compares favorably with Variational Auto Encoder-HMM.

Unsupervised speech recognition without labeled data using novel cost function and MAP refinement.

problem Training speech recognition systems without labeled data.
method Alternates between phoneme classifier learning and boundary refinement using Segmental Empirical Output Distribution Matching and MAP approach.
result Achieves phone error rate (PER) of 41.6% on TIMIT dataset.

Many machine learning tasks can be expressed as the transformation---or \emph{transduction}---of input sequences into output sequences: speech recognition, machine translation, protein secondary structure prediction and text-to-speech to name but a few. One of the key challenges in sequence transduction is learning to …

2012-11-14abs ↗pdf ↗

Quaternion CNNs improve speech recognition with fewer parameters.

problem Efficient end-to-end speech recognition with minimal parameters.
method Integrating quaternion algebra into CNNs for speech feature processing.
result Quaternion CNNs achieve lower phoneme error rates with fewer parameters.

Improved phone classification using semi-supervised learning with autoencoders.

problem Phone classification accuracy with limited labeled data.
method Semi-supervised learning with sparse autoencoders, using both labeled and unlabelled data.
result The method outperforms standard supervised training and provides competitive error rates.

End-to-end transformer model improves lexical stress detection accuracy.

problem Inaccurate phoneme boundaries and limited features for stress classification.
method End-to-end sequence to sequence model using transformer trained on feature sequences and phoneme sequences with stress marks.
result End-to-end model achieves better performance and lower phoneme error rate (6.36%) compared to syllable segmentation methods.

RUSLAN is a large Russian speech corpus for text-to-speech.

problem Lack of high-quality annotated Russian speech data for text-to-speech.
method Developed a large annotated Russian speech corpus and trained a neural network for text-to-speech synthesis.
result Synthesized speech quality evaluated with MOS scores: 4.05 for naturalness, 3.78 for intelligibility.

Improved speech recognition with audio-visual fusion.

problem Enhance speech recognition accuracy in noisy conditions.
method Proposes an attention-based audio-visual fusion strategy to align and learn from acoustic and lip motion data.
result Significant improvements in recognition accuracy (7-30%) on TCD-TIMIT dataset.

This paper improves speech recognition by distilling knowledge from acoustic models.

problem Improving speech recognition accuracy using ensemble models.
method Proposes multi-teacher distillation strategies for joint CTC-attention end-to-end ASR systems, integrating error rate metric for optimization.
result Reports state-of-the-art error rates on various datasets and languages.

Researchers solve word2vec's Corpus Replication Task to understand relational word similarities.

problem Understanding how word2vec captures relational similarities in word embeddings.
method Propose a Corpus Replication Task to generate input text for word2vec that outputs specific target relations.
result Demonstrates that word2vec can capture relational similarities in word embeddings.

ClovaCall introduces a new Korean call speech corpus for contact centers.

problem Lack of large-scale call-based speech corpora for Korean dialog scenarios.
method Development of a new large-scale Korean call-based speech corpus (ClovaCall) in a restaurant reservation domain.
result Validation of the dataset with ASR models shows its effectiveness.

Corpus poisoning can manipulate word meanings in word embeddings, affecting natural language processing tasks.

problem Controlling word meanings via corpus modifications.
method Developed an explicit expression over corpus features to control word embeddings.
result Demonstrated the ability to manipulate word meanings in word embeddings, affecting various downstream tasks.

Prob-PIT improves speech separation by considering output-label permutations as random variables.

problem Overconfident output-label assignment in PIT leads to unreliable speech separation.
method Prob-PIT treats output-label permutations as a discrete latent random variable with a uniform prior distribution and maximizes the log-likelihood function.
result Prob-PIT significantly outperforms PIT in terms of Signal to Distortion Ratio and Signal to Interference Ratio.

The UN General Debate Corpus analyzes speeches from UN member states to reveal their political positions.

problem Lack of data on state preferences in international politics.
method Text analysis of over 7,700 speeches from 1970-2016.
result Demonstrates how the UN General Debate Corpus can reveal country positions on various policy dimensions.

Improved embeddings by topic-sensitive attention on large corpora.

problem Capturing sense of words in limited corpora using pretrained embeddings.
method Topic-sensitive attention on large topic-rich corpora to correct sense drift in pretrained embeddings.
result Limited corpus augmentation is more effective than adapting pretrained embeddings.

Unified framework improves cross-corpus EEG emotion recognition by aligning prototypes and refining decision boundaries.

problem Cross-corpus EEG emotion recognition suffers from performance degradation due to physiological variability and device inconsistencies.
method Prototype-driven Adversarial Alignment (PAA) framework with three configurations: local, contrastive, and boundary-aware.
result State-of-the-art performance improvements across four cross-corpus evaluation protocols.

Paper proposes continual learning for sentence encoders.

problem Optimize sentence encoders for new corpora while maintaining old corpus accuracy.
method Initialize encoders with corpus-independent features, update using Boolean operations of conceptor matrices.
result Proposed sentence encoder can continually learn features from new corpora.

Modeling lead-lag relationship between two text corpora for improved topic modeling.

problem Recognizing the relationship between multiple text corpora for better topic modeling.
method Proposed a jointly dynamic topic model and embedding extension for large-scale text corpus.
result The proposed model can well recognize the lead-lag relationship between two text corpora and improve topic learning.

Proposes a tree-based method to efficiently predict user interests in large recommender systems.

problem Efficiently predicting user-item preferences in large recommender systems with high calculation costs.
method Predicts user interests from coarse to fine using a tree structure, which can incorporate deep neural networks.
result Significantly outperforms traditional methods in both training and prediction.

Neural User Simulator outperforms traditional ABUS in training dialogue systems.

problem Limited diversity and lack of natural language in ABUS.
method NUS learns user behavior from a corpus and generates natural language.
result NUS trained policies outperform ABUS in real user evaluations.

Proposes new methods for interpreting document classification models.

problem Interpretation fragility of attention-based neural networks.
method Corpus-level and concept-based explanation methods using attention weights.
result Extracts semantically meaningful keywords and concepts for model predictions.

In this paper, we propose and study random maxout features, which are constructed by first projecting the input data onto sets of randomly generated vectors with Gaussian elements, and then outputing the maximum projection value for each set. We show that the resulting random feature map, when used in conjunction with …

2015-06-11abs ↗pdf ↗

Synthetic continued pretraining enhances model performance with synthetic data.

problem Data inefficiency in pretrained models when adapting to domain-specific documents.
method Synthetic data augmentation using EntiGraph to create a large synthetic corpus.
result Language models can answer questions and follow instructions without access to domain-specific documents.

Deep learning improves seizure detection in EEGs.

problem Challenges in automated seizure detection in EEGs due to low signal-to-noise ratio and confusion with artifacts.
method Evaluation of hybrid deep structures including Convolutional Neural Networks and Long Short-Term Memory Networks on the TUH EEG Seizure Corpus.
result 30% sensitivity at 7 false alarms per 24 hours using a novel recurrent convolutional architecture.

We stabilize the activations of Recurrent Neural Networks (RNNs) by penalizing the squared distance between successive hidden states' norms. This penalty term is an effective regularizer for RNNs including LSTMs and IRNNs, improving performance on character-level language modeling and phoneme recognition, and outperfor…

2015-11-26abs ↗pdf ↗

Enhances persona-based conversation model for multi-turn dialogue.

problem Improving persona-based conversation models for multi-turn dialogue.
method Introduced additional input modality into hredGAN to capture external attributes.
result Persona hredGAN (phredGANphredGAN) outperforms existing models in multi-turn dialogue corpora.

The paper creates a language evolution tree using word vectors from historical novels.

problem Exploring the evolution of language through historical texts.
method Constructed word vectors from novels, combined them, and used hierarchical clustering.
result Discovered a specific language evolution tree that reflects the year of the corpus.

Deep Boltzmann Machines analyze NeuroSynth's text data for brain function insights.

problem Analyzing text data from NeuroSynth for understanding brain function.
method Unsupervised analysis of NeuroSynth's text corpus using Deep Boltzmann Machines.
result DBMs learn embeddings with a clear semantic structure, facilitating machine learning.