A new text representation model combines CNN and VAE for better semantic extraction.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A method for disentangling text representations without supervision.
New text-to-image diffusion models improve scene understanding for AI agents.
Most of the information is stored as text, so text mining is regarded as having high commercial potential. Aiming at the semantic constraint problem of classification methods based on sparse representation, we propose a weighted recurrent neural network (W-RNN), which can fully extract text serialization semantic infor…
We establish a gluing construction for Higgs bundles over a connected sum of Riemann surfaces in terms of solutions to the -Hitchin equations using the linearization of a relevant elliptic operator. The construction can be used to provide model Higgs bundles in all the exce…
Let be a surface bundle over a circle with monodromy . We study deformations of certain reducible representations of into , obtained by composing a reducible representation into with the irreducible representation $\text{SL}(2,\mathb…
Learning text representation is crucial for text classification and other language related tasks. There are a diverse set of text representation networks in the literature, and how to find the optimal one is a non-trivial problem. Recently, the emerging Neural Architecture Search (NAS) techniques have demonstrated good…
Strict plurisubharmonicity proven for Teichmüller energy on Hitchin representations.
New text as data techniques offer a great promise: the ability to inductively discover measures that are useful for testing social science theories of interest from large collections of text. We introduce a conceptual framework for making causal inferences with discovered measures as a treatment or outcome. Our framewo…
A new framework uses text descriptions to improve protein design.
Visual question answering is a recently proposed artificial intelligence task that requires a deep understanding of both images and texts. In deep learning, images are typically modeled through convolutional neural networks, and texts are typically modeled through recurrent neural networks. While the requirement for mo…
LFD method improves text classification by making features clearer and less label-leaking.
Enhances GNNs with text features for better fake news detection.
The paper studies mapping class group actions on character varieties of surfaces.
Paper develops heavy-tailed embeddings for better text classification and augmentation.
Recent progress in AutoML has lead to state-of-the-art methods (e.g., AutoSKLearn) that can be readily used by non-experts to approach any supervised learning problem. Whereas these methods are quite effective, they are still limited in the sense that they work for tabular (matrix formatted) data only. This paper descr…
Autoencoders have been successful in learning meaningful representations from image datasets. However, their performance on text datasets has not been widely studied. Traditional autoencoders tend to learn possibly trivial representations of text documents due to their confounding properties such as high-dimensionality…
Generative autoencoders offer a promising approach for controllable text generation by leveraging their latent sentence representations. However, current models struggle to maintain coherent latent spaces required to perform meaningful text manipulations via latent vector operations. Specifically, we demonstrate by exa…
CLIP learns joint image-text representations for zero-shot learning.
In this continuation of \cite{BM}, we prove the following: Let be a cocompact lattice, and let be an irreducible representation. Then the holomorphic vector bundle associated to is polystab…
Given the fundamental group of a finite-volume complete hyperbolic -manifold , it is possible to associate to any representation a numerical invariant called volume. This invariant is bounded by the hyperbolic volume of and satisfies a rigidity condition: if the …
The paper defines and calculates Reidemeister torsion for a specific class of representations.
Manually labelling large collections of text data is a time-consuming, expensive, and laborious task, but one that is necessary to support machine learning based on text datasets. Active learning has been shown to be an effective way to alleviate some of the effort required in utilising large collections of unlabelled …
We present a comprehensive study on the use of autoencoders for modelling text data, in which (differently from previous studies) we focus our attention on the following issues: i) we explore the suitability of two different models bDA and rsDA for constructing deep autoencoders for text data at the sentence level; ii)…
Bi-directional LSTMs are a powerful tool for text representation. On the other hand, they have been shown to suffer various limitations due to their sequential nature. We investigate an alternative LSTM structure for encoding text, which consists of a parallel state for each word. Recurrent steps are used to perform lo…
Characterizes flag geometries for Hitchin representations in SL3(R).
POTA improves short text clustering by generating reliable pseudo-labels.
DPNR preserves privacy of text representations using differential privacy.
Let be a closed surface of genus . In this paper, we investigate the relationship between hyperbolic cone-structure on and representations of the fundamental group into . We consider surfaces of genus greater than and we show that, under suitable conditions, every representation $ρ:π_…
Automated sentiment analysis and opinion mining is a complex process concerning the extraction of useful subjective information from text. The explosion of user generated content on the Web, especially the fact that millions of users, on a daily basis, express their opinions on products and services to blogs, wikis, so…
The paper formalizes how concepts are encoded in text-guided generative models and provides a method to manipulate them.
Paper tackles supervision bottleneck in machine learning.
The paper addresses causal estimation for text data with apparent overlap violations.
We generalize arc coordinates for maximal representations on a pair of pants.
It has been known since the time of Nielsen that the mapping class group of a surface of genus and one puncture acts faithfully by homeomorphisms on the circle. In this note, we show that this standard representation of the mapping class group is not rigid, precisely, if is a…
CNNs, RNNs, GCNs, and CapsNets have shown significant insights in representation learning and are widely used in various text mining tasks such as large-scale multi-label text classification. However, most existing deep models for multi-label text classification consider either the non-consecutive and long-distance sem…
Let be a surface of genus at least . A representation is said to be purely hyperbolic if its image consists only of hyperbolic elements other than the identity. We may wonder under which conditions such representations arise as holonomy of a hyperbolic cone-structur…
Trains word embeddings from music and text data to link music contexts.
A novel method extracts topological features from word embeddings for text classification.
This paper shows how to approximate CAT(-1) representations by Fuchsian ones.
This paper fine-tunes LLMs for stock return prediction using financial news.
Learning latent representations from long text sequences is an important first step in many natural language processing applications. Recurrent Neural Networks (RNNs) have become a cornerstone for this challenging task. However, the quality of sentences during RNN-based decoding (reconstruction) decreases with the leng…
Recent work in learning ontologies (hierarchical and partially-ordered structures) has leveraged the intrinsic geometry of spaces of learned representations to make predictions that automatically obey complex structural constraints. We explore two extensions of one such model, the order-embedding model for hierarchical…
Interpretable text-response modelling for structured outcomes
A commonly used evaluation metric for text-to-image synthesis is the Inception score (IS) \cite{inceptionscore}, which has been shown to be a quality metric that correlates well with human judgment. However, IS does not reveal properties of the generated images indicating the ability of a text-to-image synthesis method…
Over the last few years, machine learning over graph structures has manifested a significant enhancement in text mining applications such as event detection, opinion mining, and news recommendation. One of the primary challenges in this regard is structuring a graph that encodes and encompasses the features of textual …
Model uses LLM features to predict stock returns effectively.
This study reviews text-based stock market analysis methods.