Large open EEG seizure corpus created for clinical research.
problem Creating an accurate representation of clinical seizure EEG data.
method Developed a large open EEG seizure corpus, described techniques, and evaluated their effectiveness.
result Presented descriptive statistics on the resulting large open EEG seizure corpus.
Deep learning improves seizure detection in EEGs.
problem Challenges in automated seizure detection in EEGs due to low signal-to-noise ratio and confusion with artifacts.
method Evaluation of hybrid deep structures including Convolutional Neural Networks and Long Short-Term Memory Networks on the TUH EEG Seizure Corpus.
result 30% sensitivity at 7 false alarms per 24 hours using a novel recurrent convolutional architecture.
Improved EEG event classification using differential energy.
problem Automatic classification of EEG signals from time frequency representations.
method Comparison of feature extraction techniques, including differential energy and derivatives.
result 24% absolute reduction in error rate, improved discrimination between signal events and noise.
Tool interprets EEGs with high sensitivity and low false alarms.
problem Improving real-time diagnosis of EEGs for clinicians.
method Hybrid machine learning system combining HMM and deep learning.
result Delivers sensitivity above 90% with specificity below 5%.
SeizureNet classifies EEG seizures with high accuracy.
problem Challenges in classifying epileptic seizures due to signal quality and patient variability.
method Deep learning framework using multi-spectral feature embeddings and knowledge distillation.
result SeizureNet achieves high F1 scores for seizure and patient-wise classification.
Paper benchmarks machine learning for multi-class seizure type classification.
problem Accurate classification of seizure types in epileptic patients.
method Used machine learning algorithms on the TUH EEG seizure corpus.
result Achieved weighted F1 scores of up to 0.901 for seizure-wise cross validation. Convolutional LSTM networks outperform GRU in EEG seizure detection.
problem Seizure detection in EEG signals.
method Comparison of LSTM and GRU units, hybrid CNN-RNN architecture, various initialization and regularization methods.
result Convolutional LSTM networks achieve 30% sensitivity at 6 false alarms per 24 hours.
Study compares two EEG reference points and finds LE montage improves machine learning performance.
problem Variability in EEG data affects machine learning performance.
method Comparison of Linked Ear (LE) and Averaged Reference (AR) montages in machine learning performance.
result A system trained on Linked Ear data outperforms one trained only on Averaged Reference data (77.2% vs. 61.4%).
Neural memory networks improve seizure type classification.
problem Automating the classification of seizure type for clinical and research purposes.
method Introduced a novel approach using neural memory networks (NMNs) enhanced with external memory modules and trainable neural plasticity.
result Achieved a state-of-the-art weighted F1 score of 0.945 for seizure type classification.
New metrics improve evaluation of EEG event detection algorithms.
problem Lack of standard evaluation metrics for EEG event detection.
method Proposed and demonstrated new metrics: ATWV and TAES.
result Deep learning algorithms need improvement for strict user acceptance.
Machine learning improves EEG pathology classification.
problem Automating clinical EEG analysis using machine learning.
method Developed a comprehensive feature-based framework and compared it to deep neural networks.
result Feature-based framework achieves accuracies similar to deep neural networks.
RUSLAN is a large Russian speech corpus for text-to-speech.
problem Lack of high-quality annotated Russian speech data for text-to-speech.
method Developed a large annotated Russian speech corpus and trained a neural network for text-to-speech synthesis.
result Synthesized speech quality evaluated with MOS scores: 4.05 for naturalness, 3.78 for intelligibility.
Researchers solve word2vec's Corpus Replication Task to understand relational word similarities.
problem Understanding how word2vec captures relational similarities in word embeddings.
method Propose a Corpus Replication Task to generate input text for word2vec that outputs specific target relations.
result Demonstrates that word2vec can capture relational similarities in word embeddings.
Study examines market risks on pension system sustainability.
problem Impact of market risks on pension corpus sustainability.
method Monte Carlo simulations with historical data.
result Market risks significantly impact pension corpus sustainability.
Quantum analysis tags news for sentiment and entities.
problem Identifying bias in news reporting.
method Continuous data collection, NER and sentiment analysis.
result A corpus of tagged news articles for public use.
ClovaCall introduces a new Korean call speech corpus for contact centers.
problem Lack of large-scale call-based speech corpora for Korean dialog scenarios.
method Development of a new large-scale Korean call-based speech corpus (ClovaCall) in a restaurant reservation domain.
result Validation of the dataset with ASR models shows its effectiveness.
Corpus poisoning can manipulate word meanings in word embeddings, affecting natural language processing tasks.
problem Controlling word meanings via corpus modifications.
method Developed an explicit expression over corpus features to control word embeddings.
result Demonstrated the ability to manipulate word meanings in word embeddings, affecting various downstream tasks.
Challenge evaluates semantic code search using annotated corpus.
problem Evaluating relevant code from natural language queries.
method Release of CodeSearchNet Corpus and expert annotations.
result 99 queries with 4k relevance annotations for evaluation.
The UN General Debate Corpus analyzes speeches from UN member states to reveal their political positions.
problem Lack of data on state preferences in international politics.
method Text analysis of over 7,700 speeches from 1970-2016.
result Demonstrates how the UN General Debate Corpus can reveal country positions on various policy dimensions.
Improved embeddings by topic-sensitive attention on large corpora.
problem Capturing sense of words in limited corpora using pretrained embeddings.
method Topic-sensitive attention on large topic-rich corpora to correct sense drift in pretrained embeddings.
result Limited corpus augmentation is more effective than adapting pretrained embeddings.
New Polish word embeddings improve temporal expression recognition.
problem Recognizing temporal expressions in Polish text.
method Created KGR10 corpus, used BiLSTM-CRF model with custom embeddings.
result Custom embeddings enhance BiLSTM-CRF model's performance.
Unified framework improves cross-corpus EEG emotion recognition by aligning prototypes and refining decision boundaries.
problem Cross-corpus EEG emotion recognition suffers from performance degradation due to physiological variability and device inconsistencies.
method Prototype-driven Adversarial Alignment (PAA) framework with three configurations: local, contrastive, and boundary-aware.
result State-of-the-art performance improvements across four cross-corpus evaluation protocols.
End-to-end models classify composers from musical scores.
problem Classifying composers from musical scores using machine learning.
method Pooled and convolutional architectures for feature extraction.
result Models achieve high accuracy on a large corpus of scores.
FairGround offers a diverse dataset corpus for fair ML research.
problem Lack of diverse, well-annotated datasets in fair ML research.
method Unified framework and Python package for reproducible fair ML research.
result Advances reproducibility and generalizability of fair ML research.
Paper proposes continual learning for sentence encoders.
problem Optimize sentence encoders for new corpora while maintaining old corpus accuracy.
method Initialize encoders with corpus-independent features, update using Boolean operations of conceptor matrices.
result Proposed sentence encoder can continually learn features from new corpora.
Modeling lead-lag relationship between two text corpora for improved topic modeling.
problem Recognizing the relationship between multiple text corpora for better topic modeling.
method Proposed a jointly dynamic topic model and embedding extension for large-scale text corpus.
result The proposed model can well recognize the lead-lag relationship between two text corpora and improve topic learning.
Proposes a tree-based method to efficiently predict user interests in large recommender systems.
problem Efficiently predicting user-item preferences in large recommender systems with high calculation costs.
method Predicts user interests from coarse to fine using a tree structure, which can incorporate deep neural networks.
result Significantly outperforms traditional methods in both training and prediction.
Neural User Simulator outperforms traditional ABUS in training dialogue systems.
problem Limited diversity and lack of natural language in ABUS.
method NUS learns user behavior from a corpus and generates natural language.
result NUS trained policies outperform ABUS in real user evaluations.
Proposes new methods for interpreting document classification models.
problem Interpretation fragility of attention-based neural networks.
method Corpus-level and concept-based explanation methods using attention weights.
result Extracts semantically meaningful keywords and concepts for model predictions.
New method to understand bias in word embeddings.
problem Understanding and mitigating bias in word embeddings.
method Developed a technique to trace bias origins back to training documents.
result Accurate approximations of bias reduction can be made.
Synthetic continued pretraining enhances model performance with synthetic data.
problem Data inefficiency in pretrained models when adapting to domain-specific documents.
method Synthetic data augmentation using EntiGraph to create a large synthetic corpus.
result Language models can answer questions and follow instructions without access to domain-specific documents.
Large-scale automated meta-analysis of neuroimaging data has recently established itself as an important tool in advancing our understanding of human brain function. This research has been pioneered by NeuroSynth, a database collecting both brain activation coordinates and associated text across a large cohort of neuro…
Paper proposes FOFE for efficient WSD.
problem Word sense disambiguation (WSD) problem.
method Fixed-size ordinally forgetting encoding (FOFE) combined with FFNN.
result FOFE-based FFNN achieves comparable performance to state-of-the-art at lower cost.
Enhances persona-based conversation model for multi-turn dialogue.
problem Improving persona-based conversation models for multi-turn dialogue.
method Introduced additional input modality into hredGAN to capture external attributes.
result Persona hredGAN (phredGAN) outperforms existing models in multi-turn dialogue corpora. The paper creates a language evolution tree using word vectors from historical novels.
problem Exploring the evolution of language through historical texts.
method Constructed word vectors from novels, combined them, and used hierarchical clustering.
result Discovered a specific language evolution tree that reflects the year of the corpus.
In the probabilistic topic models, the quantity of interest---a low-rank matrix consisting of topic vectors---is hidden in the text corpus matrix, masked by noise, and the Singular Value Decomposition (SVD) is a potentially useful tool for learning such a low-rank matrix. However, the connection between this low-rank m…
Bayesian SHMM discovers acoustic units from unlabeled speech.
problem Discovering language-specific acoustic units from unlabeled speech.
method Bayesian Subspace Hidden Markov Model (SHMM) trained on labeled data to find new acoustic units on target language.
result Significantly outperforms previous HMM-based systems and compares favorably with Variational Auto Encoder-HMM.
Imaging neuroscience links brain activation maps to behavior and cognition via correlational studies. Due to the nature of the individual experiments, based on eliciting neural response from a small number of stimuli, this link is incomplete, and unidirectional from the causal point of view. To come to conclusions on t…
A model for entity relatedness considering time and context.
problem Time and context influence entity relatedness.
method Flexible model using time-aware and corpus-specific word embeddings.
result The model generates accurate entity relatedness lists.
New topic modeling method scales to large datasets.
problem Large-scale topic modeling with high co-occurrence data.
method Introduced Full Dependence Mixture (FDM) model for direct topic learning.
result FDM model performs comparably or better than benchmarks on large datasets.
Project develops lip reading algorithm for limited English.
problem Limited audio information for lip reading.
method Extract lip positions from video frames, classify visemes and phonemes, use HMMs to predict words.
result Algorithm predicts words from lip movements for a subset of English.
Study reduces gender bias in web data used for image recognition.
problem Gender bias in web data amplifies in machine learning models.
method Inject corpus-level constraints for calibrating structured prediction models.
result Bias amplification decreased by 47.5% and 40.5% for multilabel classification and visual semantic role labeling.
New corpus improves coreference resolution by removing gender and number cues.
problem Challenges in resolving ambiguous pronoun references.
method Developed a new annotated corpus, introduced antecedent switching technique.
result Models perform poorly on ambiguous pronoun references, but antecedent switching improves performance.
ProxiModel extracts high-quality news events from news corpora.
problem Mining high-quality structured event knowledge from noisy news data.
method ProxiModel uses a proximity-network to model event correlation within and across news corpora.
result ProxiModel efficiently and effectively extracts high-quality event descriptors and attributes.
Paper uses NLP to cluster patient visits for diagnosis validation.
problem Validating if similar patients receive similar diagnoses.
method Representation of medical visits using word embeddings, clustering patients' visits.
result Stable and separated segments of visits positively validated against diagnoses.
Improved GEC with weakly supervised data and iterative decoding.
problem Grammatical error correction using limited labeled data.
method Transformer model trained on weakly supervised bitext, iterative decoding.
result Iterative decoding improves GEC performance on CoNLL'14 benchmark.
Naive Bayes model performs best in classifying seismological articles about precursory seismicity.
problem Classifying seismological articles about precursory seismicity using machine learning.
method Various supervised machine learning classifiers (Naive Bayes, k-Nearest Neighbors, Support Vector Machines, Random Forests) were tested on a seismological corpus of 100 articles.
result Naive Bayes model performs best with cross-validation accuracies of 86% for binary classification and up to 78% for multiclass classification.
Nested Chinese Restaurant Process (nCRP) topic models are powerful nonparametric Bayesian methods to extract a topic hierarchy from a given text corpus, where the hierarchical structure is automatically determined by the data. Hierarchical Latent Dirichlet Allocation (hLDA) is a popular instance of nCRP topic models. H…