This study identifies sentence relationships in legal transcripts.
problem Improving understanding of legal case proceedings through sentence relationships.
method Combining machine learning and rule-based approach to classify sentence relationships.
result First study to use discourse relationships for legal court case transcripts.
We present the Bayesian Echo Chamber, a new Bayesian generative model for social interaction data. By modeling the evolution of people's language usage over time, this model discovers latent influence relationships between them. Unlike previous work on inferring influence, which has primarily focused on simple temporal…
Study finds managers' tenure and education influence their choice between in-court and out-of-court restructuring.
problem Exploring managers' characteristics and their impact on restructuring decisions.
method Empirical investigation using upper echelons theory and data from 342 managers of French firms.
result Managers with longer tenure and higher education levels prefer private restructuring over court involvement.
Bi-LSTM classifies legal documents for Brazil's supreme court.
problem Clogging of Brazil's supreme court due to high volume of lawsuit cases.
method Used a Bidirectional Long Short-Term Memory (Bi-LSTM) network.
result Successfully classified legal documents for efficient case allocation.
Successful attempts to predict judges' votes shed light into how legal decisions are made and, ultimately, into the behavior and evolution of the judiciary. Here, we investigate to what extent it is possible to make predictions of a justice's vote based on the other justices' votes in the same case. For our predictions…
Paper presents datasets from European Court of Human Rights judgments for classification studies.
problem Lack of accessible and reproducible datasets for legal judgments.
method Automated open-source scripts for data collection and feature transformation; experimental campaign on machine learning algorithms.
result Consistently good accuracy (75.86% - 98.32%) across binary datasets, with an average accuracy of 96.45%.
What happens when the Supreme Court of the United States decides a case impacting one or more publicly-traded firms? While many have observed anecdotal evidence linking decisions or oral arguments to abnormal stock returns, few have rigorously or systematically investigated the behavior of equities around Supreme Court…
Recent advances in high-throughput cDNA sequencing (RNA-Seq) technology have revolutionized transcriptome studies. A major motivation for RNA-Seq is to map the structure of expressed transcripts at nucleotide resolution. With accurate computational tools for transcript reconstruction, this technology may also become us…
End-to-end ASR error detection using audio-transcript entailment.
problem Detecting transcription errors in ASR systems to prevent error propagation.
method Proposes a novel end-to-end approach using audio-transcript entailment, with acoustic and linguistic encoders.
result Achieves CER of 26.2% on all transcription errors and 23% on medical errors specifically, improving by 12% and 15.4% respectively over a strong baseline.
TF-MoDISco finds transcription factor motifs from genomic data.
problem Identifying transcription factor motifs from genomic sequence data.
method Algorithm for motif discovery from basepair-level importance scores.
result Improved version v0.5.6.5 of TF-MoDISco.
Adversarial learning improves music transcription accuracy.
problem Conditional independence of labels in deep learning models limits transcription performance.
method Adversarial training scheme operating on time-frequency representations to reduce inter-label dependencies.
result Adversarial learning reduces error rate and increases model confidence.
siRF identifies transcription factor binding near enhancers in flies.
problem Identifying functional transcription factor binding near enhancers.
method Signed iterative random forests (siRF) for machine learning.
result Infers regulatory interactions among transcription factors and enhancers.
Computational approaches to transcription factor binding site identification have been actively researched for the past decade. Negative examples have long been utilized in de novo motif discovery and have been shown useful in transcription factor binding site search as well. However, understanding of the roles of nega…
Inverse Drum Machine separates drum mixes using transcription and synthesis.
problem Separating individual drum tracks from mixed recordings.
method Analysis-by-synthesis framework combining deep learning and automatic transcription.
result Separation quality comparable to supervised methods requiring isolated stems.
Researchers predict NBA player salaries using machine learning, avoiding overfitting.
problem Predicting NBA player salaries based on performance statistics.
method Selected important determinants, used Random Forest machine learning, avoided overfitting.
result Very satisfactory salary predictions identified for important factors.
Physically-inspired Gaussian process models study post-transcriptional regulation in Drosophila.
problem Understanding spatiotemporal interactions between mRNAs and gap proteins in post-transcriptional regulation.
method Two physically-inspired Gaussian process models based on reaction-diffusion equations, tested with mRNA expression data.
result Novel GP model requires only kernel function differentiation, simplifying spatial discretisation.
Motivation: Cell-biological processes are regulated through a complex network of interactions between genes and their products. The processes, their activating conditions, and the associated transcriptional responses are often unknown. Organism-wide modeling of network activation can reveal unique and shared mechanisms…
Paper proposes M2H-GAN to improve speech theme identification.
problem Limited ASR transcripts for speech theme identification.
method Uses M2H-GAN, a GAN-based approach, to generate TRS-like ASR transcripts.
result Improves speech theme identification performance close to human levels.
Study on supply chain networks using wire transfers in Brazil.
problem Understanding economic integration and specialization in Brazilian cities.
method Constructed a directed and weighted network of wire transfers between cities, analyzed centrality measures, and used econometric analysis.
result Disassortative mixing pattern in trade network, stronger after recession, and impact of court efficiency on economic transactions.
Machine learning predicts ECHR judgments on human rights violations.
problem Predicting the outcome of ECHR judgments on human rights violations.
method Auto-sklearn for model selection, N-grams, word embeddings, doc2vec, echr2vec for feature extraction, cross-validation for accuracy assessment.
result Features from echr2vec embedding provided the highest cross-validation accuracy for 5 Articles, overall test accuracy was 68.83%.
Study of SK-N-AS cells' response to methamidophos using transcriptomics.
problem Understanding the transcriptional response of SK-N-AS cells to methamidophos exposure.
method Combination of statistical and machine learning methods for anomaly detection and causal network inference.
result Identification of key processes and transcripts involved in the response to methamidophos.
A new ranking model uses nonnegative matrix factorization for tennis players.
problem Modeling latent variables influencing tennis player performance.
method Combines Bradley-Terry-Luce model with nonnegative matrix factorization.
result Model identifies surface type as key determinant of male player performance.
Generative AI predicts economic activity from corporate transcripts.
problem Predicting economic activity using existing measures like surveys.
method Extracted managerial expectations from transcripts using generative AI.
result AI Economy Score predicts economic activity up to 10 quarters ahead.
Study detects SLI in children from spontaneous narrative transcripts.
problem Detecting Specific Language Impairment (SLI) in children.
method Three-stage pipeline: feature extraction, dimensionality reduction, and classification.
result 97.13% accuracy in identifying SLI from transcripts.
In training a deep learning system to perform audio transcription, two practical problems may arise. Firstly, most datasets are weakly labelled, having only a list of events present in each recording without any temporal information for training. Secondly, deep neural networks need a very large amount of labelled train…
Model earnings call transcripts for better stock price prediction.
problem Predicting future stock price movements using earnings call transcripts.
method Deep learning framework with an attention mechanism to encode text data into vectors for predicting stock price movements.
result The proposed model outperforms traditional machine learning methods in stock price prediction.
We advance the state of the art in polyphonic piano music transcription by using a deep convolutional and recurrent neural network which is trained to jointly predict onsets and frames. Our model predicts pitch onset events and then uses those predictions to condition framewise pitch predictions. During inference, we r…
To survive environmental conditions, cells transcribe their response activities into encoded mRNA sequences in order to produce certain amounts of protein concentrations. The external conditions are mapped into the cell through the activation of special proteins called transcription factors (TFs). Due to the difficult …
Many of the recent approaches to polyphonic piano note onset transcription require training a machine learning model on a large piano database. However, such approaches are limited by dataset availability; additional training data is difficult to produce, and proposed systems often perform poorly on novel recording con…
Study analyzes deep learning models for financial sentiment in earnings calls.
problem Leveraging NLP for sentiment analysis in financial transcripts.
method Comparative analysis of BERT, FinBERT, and ULMFiT models.
result Models' strengths and limitations in financial sentiment analysis.
We investigate the problem of modeling symbolic sequences of polyphonic music in a completely general piano-roll representation. We introduce a probabilistic model based on distribution estimators conditioned on a recurrent neural network that is able to discover temporal dependencies in high-dimensional sequences. Our…
Neural networks predict TED Talk ratings from transcripts, removing bias.
problem Predicting public speaking performance from speech transcripts.
method Causal diagram modeling, word sequence and dependency tree based neural networks.
result Average F-score of 0.77, significantly outperforming baseline methods.
Formulates LGFO to measure fair ML systems using legal signals.
problem Formally incompatible measures of unfairness in ML systems.
method Uses legal signals to measure social cost of unfairness.
result LGFO aligns with societal view of unfairness.
Research shows higher damages may encourage more disclosure in corporate disputes.
problem How to resolve disputes over undisclosed material events in a way that encourages voluntary disclosure.
method Dynamic continuous-time model of management's equilibrium disclosure decision.
result Increased damages may lead to an endogenous increase in voluntary disclosure.
New task aligns molecular structure with gene expression changes.
problem Modeling the relationship between chemical structure and gene expression changes.
method Developed a cross-modal small molecule retrieval task and a coordinated deep learning approach to align chemical structure and gene expression profiles.
result Demonstrated the feasibility of the new task and highlighted the limitations of current data and systems.
Hidden Markov Models (HMMs) are a ubiquitous tool to model time series data, and have been widely used in two main tasks of Automatic Music Transcription (AMT): note segmentation, i.e. identifying the played notes after a multi-pitch estimation, and sequential post-processing, i.e. correcting note segmentation using tr…
The study constructs and shows isotopy of high-dimensional Legendrian spheres.
problem Understanding Legendrian spheres in contact manifolds of any dimension.
method Three Legendrian sphere constructions using open books and a doubling procedure.
result These constructions are isotopic to the Legendrian unknot.
The paper improves ASR accuracy using semi-supervised learning and dropout.
problem Improving ASR accuracy with limited labeled data.
method Training a seed model on limited labeled data, using dropout for uncertainty, and data selection for diversity.
result The approach significantly reduces ASR errors compared to baseline.
Predicts customer call intent for auto dealerships using CNN.
problem Understanding customer intent from phone calls for better service.
method Developed a CNN-based supervised learning model for multi-class classification.
result CNN model performs well on customer call intent classification.
Neural network generates music scores directly from polyphonic audio.
problem Transcribing music scores directly from polyphonic audio.
method Convolutional Recurrent Neural Network (CRNN) with CTC loss function.
result Model can learn to transcribe scores directly from audio signals.
Model predicts active and passive cosponsorship motivations in Congress.
problem Identifying motivations behind cosponsorship in U.S. Congress.
method Encoder+RGCN model learning from bill texts and speeches.
result F1-score of 0.88 for predicting active and passive cosponsorship.
This paper explores a variety of models for frame-based music transcription, with an emphasis on the methods needed to reach state-of-the-art on human recordings. The translation-invariant network discussed in this paper, which combines a traditional filterbank with a convolutional neural network, was the top-performin…
Study combines speaker verification and voice trigger detection in a single network.
problem Separate training for speaker verification and voice trigger detection.
method Multi-task learning with a single network trained on both tasks.
result Single network achieves comparable accuracy to independent models for each task.
Automatic Music Transcription (AMT) consists in automatically estimating the notes in an audio recording, through three attributes: onset time, duration and pitch. Probabilistic Latent Component Analysis (PLCA) has become very popular for this task. PLCA is a spectrogram factorization method, able to model a magnitude …
Anonymization reduces economic signal extraction from financial texts.
problem Reducing meaningful economic signals from financial texts due to anonymization.
method Analyzed the impact of anonymization on textual understanding and economic signal extraction.
result Information loss due to anonymization is severe and pervasive, outweighing its benefits in certain financial applications.
We use automatic speech recognition to assess spoken English learner pronunciation based on the authentic intelligibility of the learners' spoken responses determined from support vector machine (SVM) classifier or deep learning neural network model predictions of transcription correctness. Using numeric features produ…
Framework ranks sectors influenced by Indian Union Budgets.
problem Real-time analysis of budgetary impacts on sector-specific equity performance.
method Fine-tuned embeddings and language models for sector identification and performance ranking.
result 0.997 NDCG score in predicting sector ranks based on post-budget performances.
Each year, roughly 30% of first-year students at US baccalaureate institutions do not return for their second year and over $9 billion is spent educating these students. Yet, little quantitative research has analyzed the causes and possible remedies for student attrition. Here, we describe initial efforts to model stud…