TF-MoDISco finds transcription factor motifs from genomic data.
problem Identifying transcription factor motifs from genomic sequence data.
method Algorithm for motif discovery from basepair-level importance scores.
result Improved version v0.5.6.5 of TF-MoDISco.
siRF identifies transcription factor binding near enhancers in flies.
problem Identifying functional transcription factor binding near enhancers.
method Signed iterative random forests (siRF) for machine learning.
result Infers regulatory interactions among transcription factors and enhancers.
Computational approaches to transcription factor binding site identification have been actively researched for the past decade. Negative examples have long been utilized in de novo motif discovery and have been shown useful in transcription factor binding site search as well. However, understanding of the roles of nega…
Improves audio transcription on scarce data with factorized tasks.
problem Weakly labelled data and lack of training samples.
method Factorizing audio transcription into multiple tasks and training a stacked CNN-RNN model.
result Different training methods for intermediate tasks have varying advantages and disadvantages.
Improved music transcription accuracy using Particle Filtering for PLCA model.
problem Limited performance of EM-based PLCA models in automatic music transcription.
method Employed Particle Filtering (PF) to overcome EM algorithm's limitations.
result Achieved 61.8% and 59.5% note-level transcription accuracy on two instrument repertoires.
To survive environmental conditions, cells transcribe their response activities into encoded mRNA sequences in order to produce certain amounts of protein concentrations. The external conditions are mapped into the cell through the activation of special proteins called transcription factors (TFs). Due to the difficult …
Recent advances in high-throughput cDNA sequencing (RNA-Seq) technology have revolutionized transcriptome studies. A major motivation for RNA-Seq is to map the structure of expressed transcripts at nucleotide resolution. With accurate computational tools for transcript reconstruction, this technology may also become us…
End-to-end ASR error detection using audio-transcript entailment.
problem Detecting transcription errors in ASR systems to prevent error propagation.
method Proposes a novel end-to-end approach using audio-transcript entailment, with acoustic and linguistic encoders.
result Achieves CER of 26.2% on all transcription errors and 23% on medical errors specifically, improving by 12% and 15.4% respectively over a strong baseline.
Adversarial learning improves music transcription accuracy.
problem Conditional independence of labels in deep learning models limits transcription performance.
method Adversarial training scheme operating on time-frequency representations to reduce inter-label dependencies.
result Adversarial learning reduces error rate and increases model confidence.
This study identifies sentence relationships in legal transcripts.
problem Improving understanding of legal case proceedings through sentence relationships.
method Combining machine learning and rule-based approach to classify sentence relationships.
result First study to use discourse relationships for legal court case transcripts.
New models analyze stability of gene regulation networks with coregulation.
problem Stability and structure of gene regulation networks with shared regulatory motifs.
method Developed formalism for modeling coregulation rules in RBN, analyzed stability through mean-field approach.
result Coregulation can increase network stability, especially in autoregulated multi-gene modules and hierarchical gene complexes.
Inverse Drum Machine separates drum mixes using transcription and synthesis.
problem Separating individual drum tracks from mixed recordings.
method Analysis-by-synthesis framework combining deep learning and automatic transcription.
result Separation quality comparable to supervised methods requiring isolated stems.
Improved piano transcription by predicting onsets and frames together.
problem Polyphonic piano music transcription accuracy.
method Deep convolutional and recurrent neural network trained to predict pitch onsets and frames.
result Over 100% relative improvement in note F1 score on MAPS dataset.
Develops probabilistic models for gene regulatory network inference.
problem Challenges in reconstructing gene regulatory networks from genome-wide data.
method Two complementary frameworks: PMF-GRN and GLM-Prior.
result Probabilistic inference refines regulatory estimates with quantified uncertainty.
Prototype Matching Network (PMN) improves genomic TFBS prediction.
problem Predicting Transcription Factor Binding Sites (TFBSs) with hundreds of TFs as labels.
method Prototype Matching Network (PMN) that learns motif-like features and TF-TF interactions.
result PMN significantly outperforms baselines on a large TFBS dataset.
New method synthesizes piano training data, improving transcription performance.
problem Lack of large piano datasets limits note onset transcription models.
method Synthesizes arbitrary training data, models piano dynamics, avoids disentanglement problem.
result Achieves good transcription performance on MAPS dataset and excellent generalization.
Physically-inspired Gaussian process models study post-transcriptional regulation in Drosophila.
problem Understanding spatiotemporal interactions between mRNAs and gap proteins in post-transcriptional regulation.
method Two physically-inspired Gaussian process models based on reaction-diffusion equations, tested with mRNA expression data.
result Novel GP model requires only kernel function differentiation, simplifying spatial discretisation.
Paper improves music transcription models with invariance and data augmentation.
problem Improving accuracy of frame-based music transcription models.
method Translation-invariant network combining filterbank and CNN, trained with pitch-shift augmented data.
result Top-performing model in MIREX evaluation, reducing model complexity and avoiding overfitting.
Motivation: Cell-biological processes are regulated through a complex network of interactions between genes and their products. The processes, their activating conditions, and the associated transcriptional responses are often unknown. Organism-wide modeling of network activation can reveal unique and shared mechanisms…
Paper proposes M2H-GAN to improve speech theme identification.
problem Limited ASR transcripts for speech theme identification.
method Uses M2H-GAN, a GAN-based approach, to generate TRS-like ASR transcripts.
result Improves speech theme identification performance close to human levels.
HMMs improve music transcription accuracy.
problem Improving automatic transcription of music.
method Employed PLCA for multi-pitch estimation and integrated HMMs for note segmentation and post-processing.
result HMMs enhance transcription accuracy on different instruments.
Scattering transform improves note onset detection and instrument recognition in music transcription.
problem Note onset detection and instrument recognition in music transcription.
method Multiscale scattering operators applied to MIDI-driven datasets and real musical pieces.
result Scattering transform outperforms other sound representations for note onset detection and instrument recognition.
Study of SK-N-AS cells' response to methamidophos using transcriptomics.
problem Understanding the transcriptional response of SK-N-AS cells to methamidophos exposure.
method Combination of statistical and machine learning methods for anomaly detection and causal network inference.
result Identification of key processes and transcripts involved in the response to methamidophos.
With different genomes available, unsupervised learning algorithms are essential in learning genome-wide biological insights. Especially, the functional characterization of different genomes is essential for us to understand lives. In this book chapter, we review the state-of-the-art unsupervised learning algorithms fo…
Generative AI predicts economic activity from corporate transcripts.
problem Predicting economic activity using existing measures like surveys.
method Extracted managerial expectations from transcripts using generative AI.
result AI Economy Score predicts economic activity up to 10 quarters ahead.
Bayesian model fuses diverse microbiome data types.
problem Challenges in fusing different types of microbiome data.
method Flexible multinomial-Gaussian generative model with variational EM algorithm.
result Inferred latent variables provide common dimensionality reduction and predictive posterior distribution.
Study detects SLI in children from spontaneous narrative transcripts.
problem Detecting Specific Language Impairment (SLI) in children.
method Three-stage pipeline: feature extraction, dimensionality reduction, and classification.
result 97.13% accuracy in identifying SLI from transcripts.
This paper explores how combining quantitative factors and news from LLMs improves stock return prediction.
problem Improving stock return prediction using quantitative factors and news.
method Introduces a fusion learning framework to learn unified representations from factors and LLM-generated newsflow, comparing combination, summation, and attentive methods. Explores mixture models and decoupled training approaches.
result Effective multimodal modeling of factors and news improves stock return prediction and selection.
Model earnings call transcripts for better stock price prediction.
problem Predicting future stock price movements using earnings call transcripts.
method Deep learning framework with an attention mechanism to encode text data into vectors for predicting stock price movements.
result The proposed model outperforms traditional machine learning methods in stock price prediction.
Improves seq2seq speech recognition by addressing overconfidence and incomplete transcriptions.
problem Overconfidence and incomplete transcriptions in seq2seq speech recognition.
method Proposed practical solutions to address overconfidence and incomplete transcriptions using a trigram language model.
result Achieved competitive speaker independent word error rates (6.7%) with a trigram language model.
Components of biological systems interact with each other in order to carry out vital cell functions. Such information can be used to improve estimation and inference, and to obtain better insights into the underlying cellular mechanisms. Discovering regulatory interactions among genes is therefore an important problem…
Predicts student dropout using transcript data.
problem Student attrition in higher education.
method Machine learning model using transcript data.
result Dropout can be accurately predicted from a single term of transcript data.
Study analyzes deep learning models for financial sentiment in earnings calls.
problem Leveraging NLP for sentiment analysis in financial transcripts.
method Comparative analysis of BERT, FinBERT, and ULMFiT models.
result Models' strengths and limitations in financial sentiment analysis.
New method improves music transcription by treating frequency distributions holistically.
problem Small frequency shifts and variations in sound timbre harm traditional fit measures.
method Optimal transportation and new holistic frequency distribution measure.
result Simplified note templates lead to faster, state-of-the-art performance.
We investigate the problem of modeling symbolic sequences of polyphonic music in a completely general piano-roll representation. We introduce a probabilistic model based on distribution estimators conditioned on a recurrent neural network that is able to discover temporal dependencies in high-dimensional sequences. Our…
Neural networks predict TED Talk ratings from transcripts, removing bias.
problem Predicting public speaking performance from speech transcripts.
method Causal diagram modeling, word sequence and dependency tree based neural networks.
result Average F-score of 0.77, significantly outperforming baseline methods.
New task aligns molecular structure with gene expression changes.
problem Modeling the relationship between chemical structure and gene expression changes.
method Developed a cross-modal small molecule retrieval task and a coordinated deep learning approach to align chemical structure and gene expression profiles.
result Demonstrated the feasibility of the new task and highlighted the limitations of current data and systems.
The paper improves ASR accuracy using semi-supervised learning and dropout.
problem Improving ASR accuracy with limited labeled data.
method Training a seed model on limited labeled data, using dropout for uncertainty, and data selection for diversity.
result The approach significantly reduces ASR errors compared to baseline.
Dilated convolutions model long-distance genomic dependencies effectively.
problem Detecting regulatory elements from raw DNA with long-distance dependencies.
method Developed and used a novel dataset for dilated convolutional neural networks.
result Dilated convolutions are effective at modeling regulatory elements in the human genome.
Predicts customer call intent for auto dealerships using CNN.
problem Understanding customer intent from phone calls for better service.
method Developed a CNN-based supervised learning model for multi-class classification.
result CNN model performs well on customer call intent classification.
Neural network generates music scores directly from polyphonic audio.
problem Transcribing music scores directly from polyphonic audio.
method Convolutional Recurrent Neural Network (CRNN) with CTC loss function.
result Model can learn to transcribe scores directly from audio signals.
The paper uses attention networks for character-based handwritten text transcription.
problem Handwritten text recognition with improved character-level alignment.
method Attentional encoder-decoder networks trained on character sequences, comparing different activation functions.
result Softmax attention provides more precise character alignment than sigmoid attention.
Study combines speaker verification and voice trigger detection in a single network.
problem Separate training for speaker verification and voice trigger detection.
method Multi-task learning with a single network trained on both tasks.
result Single network achieves comparable accuracy to independent models for each task.
Model predicts active and passive cosponsorship motivations in Congress.
problem Identifying motivations behind cosponsorship in U.S. Congress.
method Encoder+RGCN model learning from bill texts and speeches.
result F1-score of 0.88 for predicting active and passive cosponsorship.
iRF detects stable high-order interactions in genomics data.
problem Understanding high-order interactions in genomics data.
method Iterative Random Forest algorithm (iRF) for stable high-order interaction detection.
result iRF identifies stable high-order interactions with computational cost similar to Random Forest.
Anonymization reduces economic signal extraction from financial texts.
problem Reducing meaningful economic signals from financial texts due to anonymization.
method Analyzed the impact of anonymization on textual understanding and economic signal extraction.
result Information loss due to anonymization is severe and pervasive, outweighing its benefits in certain financial applications.
Black-box adversarial examples improve ASR system accuracy.
problem Improving ASR system accuracy through targeted adversarial examples.
method Combining genetic algorithms and gradient estimation for black-box attacks.
result Achieved 89.25% targeted attack similarity with 94.6% audio file similarity.
Framework ranks sectors influenced by Indian Union Budgets.
problem Real-time analysis of budgetary impacts on sector-specific equity performance.
method Fine-tuned embeddings and language models for sector identification and performance ranking.
result 0.997 NDCG score in predicting sector ranks based on post-budget performances.