SCROLLS benchmarks long text NLP tasks, improving existing models.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Learning latent representations from long text sequences is an important first step in many natural language processing applications. Recurrent Neural Networks (RNNs) have become a cornerstone for this challenging task. However, the quality of sentences during RNN-based decoding (reconstruction) decreases with the leng…
ZeroSCROLLS benchmarks zero-shot natural language understanding over long texts.
Improved Mamba model for long-range sequence tasks.
Model criticism tool evaluates text coherence and structure in generated long-form text.
System identifies language of transliterated text.
New method predicts nonfactuality in LLM responses using semantic isotropy.
Many challenges in natural language processing require generating text, including language translation, dialogue generation, and speech recognition. For all of these problems, text generation becomes more difficult as the text becomes longer. Current language models often struggle to keep track of coherence for long pi…
We apply text analysis approaches for a specialized search engine for 3D CAD models and associated products. The main goals are to distinguish between actual product descriptions and other text on a website, as well as to decide whether a given text is or contains a product name. For this we use paragraph vectors for t…
A latent-variable model is introduced for text matching, inferring sentence representations by jointly optimizing generative and discriminative objectives. To alleviate typical optimization challenges in latent-variable models for text, we employ deconvolutional networks as the sequence decoder (generator), providing l…
Linguistic calibration improves long-form text confidence.
CNNs, RNNs, GCNs, and CapsNets have shown significant insights in representation learning and are widely used in various text mining tasks such as large-scale multi-label text classification. However, most existing deep models for multi-label text classification consider either the non-consecutive and long-distance sem…
Sparse and short news headlines can be arbitrary, noisy, and ambiguous, making it difficult for classic topic model LDA (latent Dirichlet allocation) designed for accommodating long text to discover knowledge from them. Nonetheless, some of the existing research about text-based crude oil forecasting employs LDA to exp…
Proposes a flexible neural recommendation framework for better prediction performance.
Shorter adversarial prompts help protect LLMs from jailbreak attacks.
A novel method extracts topological features from word embeddings for text classification.
Proves Payne conjecture for buckling and membrane eigenvalues.
Traditional sequence-to-sequence (seq2seq) models and other variations of the attention-mechanism such as hierarchical attention have been applied to the text summarization problem. Though there is a hierarchy in the way humans use language by forming paragraphs from sentences and sentences from words, hierarchical mod…
Learning representations that accurately capture long-range dependencies in sequential inputs -- including text, audio, and genomic data -- is a key problem in deep learning. Feed-forward convolutional models capture only feature interactions within finite receptive fields while recurrent architectures can be slow and …
Paper fine-tunes a language model to predict long-term stock buy signals.
The Generative Adversarial Network (GAN) has achieved great success in generating realistic (real-valued) synthetic data. However, convergence issues and difficulties dealing with discrete data hinder the applicability of GAN to text. We propose a framework for generating realistic text via adversarial training. We emp…
Paper calculates topological complexity of robot movement in narrow aisles.
Unified Long-Moody and Katz methods for constructing local systems.
This thesis evaluates text-based vs audio-based classification of mental health interviews.
This paper fine-tunes LLMs for stock return prediction using financial news.
One-hot CNN (convolutional neural network) has been shown to be effective for text categorization (Johnson & Zhang, 2015). We view it as a special case of a general framework which jointly trains a linear model with a non-linear feature generator consisting of `text region embedding + pooling'. Under this framework, we…
DreamPropeller accelerates text-to-3D generation by 4.7x with minimal loss in quality.
TRM improves long-horizon LLM RL by masking divergent sequences.
TRM improves long-horizon reinforcement learning for LLMs by masking divergent sequences.
Recent work in learning ontologies (hierarchical and partially-ordered structures) has leveraged the intrinsic geometry of spaces of learned representations to make predictions that automatically obey complex structural constraints. We explore two extensions of one such model, the order-embedding model for hierarchical…
Study enhances cryptocurrency sentiment analysis using TikTok and Twitter data.
Summarization of long sequences into a concise statement is a core problem in natural language processing, requiring non-trivial understanding of the input. Based on the promising results of graph neural networks on highly structured data, we develop a framework to extend existing sequence encoders with a graph compone…
AMI framework improves text generation by optimizing mutual information between source and target.
Study proposes a multimodal model for cardiovascular risk prediction using EHRs.
NoLBERT avoids lookback and lookahead biases for better econometric inference.
FINCH dataset enables financial Text-to-SQL tasks, improving model evaluation.
This text discusses several popular explanatory methods that go beyond the error measurements and plots traditionally used to assess machine learning models. Some of the explanatory methods are accepted tools of the trade while others are rigorously derived and backed by long-standing theory. The methods, decision tree…
Many real-world problems, including multi-speaker text-to-speech synthesis, can greatly benefit from the ability to meta-learn large models with only a few task-specific components. Updating only these task-specific modules then allows the model to be adapted to low-data tasks for as many steps as necessary without ris…
Markov chain Monte Carlo (MCMC) algorithms are ubiquitous in probability theory in general and in machine learning in particular. A Markov chain is devised so that its stationary distribution is some probability distribution of interest. Then one samples from the given distribution by running the Markov chain for a "lo…
Graph convolutional networks (GCNs) have shown the powerful ability in text structure representation and effectively facilitate the task of text classification. However, challenges still exist in adapting GCN on learning discriminative features from texts due to the main issue of graph variants incurred by the textual …
The paper predicts financial markets using news text and semantic network analysis.
The paper proposes a new framework to generate synthetic data with human-like imperfections to prevent model collapse.
New algorithm learns optimal policy for average reward MDPs with sample complexity matching lower bound.
ICP improves text infilling and POS tagging with valid confidence sets.
Study integrates climate and text data to improve credit default prediction.
Bayesian Neural Networks detect gravitational wave events with high accuracy and real-time potential.
In this paper we deal with the offline handwriting text recognition (HTR) problem with reduced training datasets. Recent HTR solutions based on artificial neural networks exhibit remarkable solutions in referenced databases. These deep learning neural networks are composed of both convolutional (CNN) and long short-ter…
Medical applications challenge today's text categorization techniques by demanding both high accuracy and ease-of-interpretation. Although deep learning has provided a leap ahead in accuracy, this leap comes at the sacrifice of interpretability. To address this accuracy-interpretability challenge, we here introduce, fo…