A new text representation model combines CNN and VAE for better semantic extraction.
problem Difficult to effectively extract semantic features and distinguish polysemy in text data.
method Integrates CNN for feature extraction and VAE for consistent Gaussian distribution.
result The model outperforms traditional classification algorithms in text classification tasks.
TextNAS finds optimal text representation networks using neural architecture search.
problem Finding the optimal text representation networks is challenging.
method Proposes a novel search space for text representation and uses automatic neural architecture search.
result Automatic search discovers network architectures that outperform state-of-the-art models on text classification and natural language inference tasks.
W-RNN improves text classification by extracting serialized text semantics.
problem Semantic constraint in sparse representation classification methods.
method Weighted RNN using word vectors and recurrent neural networks.
result W-RNN outperforms other methods in precision, recall, F1, and loss values.
A method for disentangling text representations without supervision.
problem Challenges in learning disentangled representations of natural language.
method Information-theoretic guidance to induce independent style and content embeddings.
result High quality disentangled representations in terms of content and style preservation.
New text-to-image diffusion models improve scene understanding for AI agents.
problem Fine-grained scene understanding for AI agents from text and images.
method Pre-trained text-to-image diffusion models optimized for generating images from text prompts.
result Policies learned with Stable Control Representations outperform state-of-the-art approaches on various control tasks.
We establish a gluing construction for Higgs bundles over a connected sum of Riemann surfaces in terms of solutions to the Sp(4,R)-Hitchin equations using the linearization of a relevant elliptic operator. The construction can be used to provide model Higgs bundles in all the 2g−3 exce…
Let Mφ be a surface bundle over a circle with monodromy φ:S→S. We study deformations of certain reducible representations of π1(Mφ) into SL(n,C), obtained by composing a reducible representation into SL(2,C) with the irreducible representation $\text{SL}(2,\mathb…
Strict plurisubharmonicity proven for Teichmüller energy on Hitchin representations.
problem Proving strict plurisubharmonicity of Teichmüller energy for Hitchin representations.
method Analyzing energy functional E on Teichmüller space associated to Hitchin representations. result Strict plurisubharmonicity of energy functional E proven. New text as data techniques offer a great promise: the ability to inductively discover measures that are useful for testing social science theories of interest from large collections of text. We introduce a conceptual framework for making causal inferences with discovered measures as a treatment or outcome. Our framewo…
A new framework uses text descriptions to improve protein design.
problem Lack of effective methods to incorporate textual descriptions in protein design.
method ProteinDT framework that combines text and protein structural information.
result ProteinDT significantly improves protein design accuracy and performance.
BERT-based word embeddings improve active learning for text datasets.
problem Efficiently labelling large text datasets for machine learning.
method Evaluation of text representation mechanisms (BERT vs. bag of words) in active learning.
result BERT-based word embeddings significantly improve active learning performance.
Visual question answering is a recently proposed artificial intelligence task that requires a deep understanding of both images and texts. In deep learning, images are typically modeled through convolutional neural networks, and texts are typically modeled through recurrent neural networks. While the requirement for mo…
LFD method improves text classification by making features clearer and less label-leaking.
problem Creating interpretable text representations that are both predictive and understandable.
method LFD method: proposes lexical and semantic features from contrastive text pairs, screens candidates using κ, and selects features by residual gain. result LFD features achieve higher human-human and human-LLM agreement than baseline concepts and are less label-leaking.
Enhances GNNs with text features for better fake news detection.
problem Detecting disinformation on social media using GNNs.
method Integrates Transformer-based textual features into GNNs.
result Contextual text representations improve GNN performance by 33.8% in Macro F1.
The paper studies mapping class group actions on character varieties of surfaces.
problem Understanding the dynamics of mapping class group actions on relative extPSL(2,R)-character varieties. method Definition and proof of simple-stability and primitive-stability of representations.
result Holonomies of hyperbolic cone surfaces are simple-stable and primitive-stable.
Paper develops heavy-tailed embeddings for better text classification and augmentation.
problem Improving text classification, especially for extreme values.
method Develops heavy-tailed embeddings using multivariate extreme value theory and introduces a scale-invariant classifier.
result The classifier outperforms baselines and generates meaningful augmented text.
Recent progress in AutoML has lead to state-of-the-art methods (e.g., AutoSKLearn) that can be readily used by non-experts to approach any supervised learning problem. Whereas these methods are quite effective, they are still limited in the sense that they work for tabular (matrix formatted) data only. This paper descr…
Autoencoders have been successful in learning meaningful representations from image datasets. However, their performance on text datasets has not been widely studied. Traditional autoencoders tend to learn possibly trivial representations of text documents due to their confounding properties such as high-dimensionality…
Generative autoencoders offer a promising approach for controllable text generation by leveraging their latent sentence representations. However, current models struggle to maintain coherent latent spaces required to perform meaningful text manipulations via latent vector operations. Specifically, we demonstrate by exa…
CLIP learns joint image-text representations for zero-shot learning.
problem Understanding and improving zero-shot transfer performance in CLIP.
method Formal study of transferrable representation learning and analysis of zero-shot transfer performance.
result Proposes a new CLIP-type approach that outperforms existing methods.
New method evaluates text-to-image synthesis for realism, variety, and semantic accuracy.
problem Lack of metrics revealing semantic accuracy in text-to-image synthesis.
method Uses Inception network representations and t-SNE visualization for semantic evaluation.
result Classification accuracy of generated images to real images' visual concepts correlates with semantic accuracy.
In this continuation of \cite{BM}, we prove the following: Let Γ⊂SL(2,C) be a cocompact lattice, and let ρ:Γ→GL(r,C) be an irreducible representation. Then the holomorphic vector bundle Eρ⟶SL(2,C)/Γ associated to ρ is polystab…
Given the fundamental group Γ of a finite-volume complete hyperbolic 3-manifold M, it is possible to associate to any representation ρ:Γ→Isom(H3) a numerical invariant called volume. This invariant is bounded by the hyperbolic volume of M and satisfies a rigidity condition: if the …
The paper defines and calculates Reidemeister torsion for a specific class of representations.
problem Defining and calculating Reidemeister torsion for G-Anosov representations.
method Symplectic chain complex method to establish a novel formula for R-torsion.
result Reidemeister torsion is well-defined and calculated for G-Anosov representations.
We present a comprehensive study on the use of autoencoders for modelling text data, in which (differently from previous studies) we focus our attention on the following issues: i) we explore the suitability of two different models bDA and rsDA for constructing deep autoencoders for text data at the sentence level; ii)…
Bi-directional LSTMs are a powerful tool for text representation. On the other hand, they have been shown to suffer various limitations due to their sequential nature. We investigate an alternative LSTM structure for encoding text, which consists of a parallel state for each word. Recurrent steps are used to perform lo…
Characterizes flag geometries for Hitchin representations in SL3(R).
problem Understanding flag geometries associated with Hitchin representations in SL3(R).
method Geometric characterization based on invariant foliations and refraction flows.
result Constructs refraction flows for positive roots in general sl_n(R), with highest root flows being C^1+α.
POTA improves short text clustering by generating reliable pseudo-labels.
problem Limited discriminative representations in short texts.
method POTA uses instance-level attention and optimal transport for semantic consistency and cluster structure.
result POTA outperforms state-of-the-art methods in short text clustering.
DPNR preserves privacy of text representations using differential privacy.
problem Privacy leakage in deep learning text representations.
method DPNR uses Differential Privacy to provide formal privacy guarantees and dropout masking for enhanced privacy.
result DPNR reduces privacy leakage without significantly sacrificing main task performance.
Let S be a closed surface of genus g. In this paper, we investigate the relationship between hyperbolic cone-structure on S and representations of the fundamental group into PSL2R. We consider surfaces of genus greater than g and we show that, under suitable conditions, every representation $ρ:π_…
Automated sentiment analysis and opinion mining is a complex process concerning the extraction of useful subjective information from text. The explosion of user generated content on the Web, especially the fact that millions of users, on a daily basis, express their opinions on products and services to blogs, wikis, so…
The paper formalizes how concepts are encoded in text-guided generative models and provides a method to manipulate them.
problem Encoding and manipulating concepts in text-guided generative models.
method Formalizing concepts as subspaces of a representation space, developing algebraic manipulation methods.
result The ability to manipulate concepts in generative models through algebraic operations on the representation.
Paper tackles supervision bottleneck in machine learning.
problem Difficulty in generating supervision signals for learning models.
method Describes several learning paradigms to alleviate the supervision bottleneck.
result Illustrates the benefit of these paradigms in inducing semantic representations from text.
The paper addresses causal estimation for text data with apparent overlap violations.
problem Estimating causal effects from text data with unknown confounders and apparent overlap.
method Uses supervised representation learning to create a representation that preserves confounding information while eliminating predictive information, satisfying overlap assumptions.
result Shows how to obtain robust causal estimation in the presence of apparent overlap violations.
We generalize arc coordinates for maximal representations on a pair of pants.
problem Maximal representations of reflection groups on hyperbolic surfaces.
method Introducing geometric parameters and reflections in Siegel space.
result Natural parametrization of maximal representations into PSp(4, R).
Model identifies urgent radiology reports with high accuracy.
problem Lack of annotated training data for text analysis.
method Self-supervised contextual language representation using BERT.
result Model achieved 97.0% precision, 93.3% recall, and 95.1% F-measure.
It has been known since the time of Nielsen that the mapping class group Modg,1 of a surface of genus g and one puncture acts faithfully by homeomorphisms on the circle. In this note, we show that this standard representation of the mapping class group is not rigid, precisely, if G<Modg,1 is a…
CNNs, RNNs, GCNs, and CapsNets have shown significant insights in representation learning and are widely used in various text mining tasks such as large-scale multi-label text classification. However, most existing deep models for multi-label text classification consider either the non-consecutive and long-distance sem…
Let S be a surface of genus g at least 2. A representation ρ:π1S⟶PSL2R is said to be purely hyperbolic if its image consists only of hyperbolic elements other than the identity. We may wonder under which conditions such representations arise as holonomy of a hyperbolic cone-structur…
Trains word embeddings from music and text data to link music contexts.
problem Varying vocabulary size and musical relevance in word embeddings.
method Combines general text and music-specific data to train word embeddings.
result Trained embeddings better associate music contexts with compositions.
A novel method extracts topological features from word embeddings for text classification.
problem High dimensional and noisy text representations in natural language processing.
method Persistent homology for topological data analysis on word embeddings.
result Topological features outperform conventional text mining features on long textual documents.
Model captures author language diffusion over time.
problem Lack of author identity and temporal context in language models.
method Temporal language model conditioning on author and temporal vectors.
result Beat temporal and non-temporal baselines, learns time-varying author representations.
This paper fine-tunes LLMs for stock return prediction using financial news.
problem Improving stock return forecasting accuracy using LLMs.
method Fine-tuning LLMs with text and forecasting modules, comparing encoder-only and decoder-only models, and integrating token-level representations.
result LLMs' aggregated token-level embeddings enhance return predictions for long-only and long-short portfolios.
This paper shows how to approximate CAT(-1) representations by Fuchsian ones.
problem Approximating representations of surface groups in CAT(-1) spaces.
method Constructing equivariant maps from the hyperbolic plane to CAT(-1) spaces.
result Every CAT(-1) representation can be approximated by a Fuchsian one.
Learning latent representations from long text sequences is an important first step in many natural language processing applications. Recurrent Neural Networks (RNNs) have become a cornerstone for this challenging task. However, the quality of sentences during RNN-based decoding (reconstruction) decreases with the leng…
Recent work in learning ontologies (hierarchical and partially-ordered structures) has leveraged the intrinsic geometry of spaces of learned representations to make predictions that automatically obey complex structural constraints. We explore two extensions of one such model, the order-embedding model for hierarchical…
Interpretable text-response modelling for structured outcomes
problem Predicting structured responses alongside textual data
method Joint non-negative matrix factorisation and binomial regression
result Recovering stable response-relevant textual signals
Over the last few years, machine learning over graph structures has manifested a significant enhancement in text mining applications such as event detection, opinion mining, and news recommendation. One of the primary challenges in this regard is structuring a graph that encodes and encompasses the features of textual …