Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

316394125 · Oct 201919922001200920182026
48 results for text visualization

Proposes CNN Inception + Gate model for visual question answering.

problem Deep understanding of images and texts for visual question answering.
method Proposes a CNN Inception + Gate model for learning textual representations.
result Improves question representations and overall accuracy in visual question answering.

New method evaluates text-to-image synthesis for realism, variety, and semantic accuracy.

problem Lack of metrics revealing semantic accuracy in text-to-image synthesis.
method Uses Inception network representations and t-SNE visualization for semantic evaluation.
result Classification accuracy of generated images to real images' visual concepts correlates with semantic accuracy.

DBNet improves natural language image localization and detection.

problem Natural language-based visual entity localization with limited accuracy.
method Discriminative bimodal neural network (DBNet) trained with extensive negative samples.
result Significantly outperforms previous methods on Visual Genome dataset.

Diestel-Leader graphs are neither hyperbolic nor CAT(0), so their visual boundaries may be pathological. Indeed, we show that for d>2d>2, DLd(q)\partial\text{DL}_d(q) carries the indiscrete topology. On the other hand, DL2(q)\partial\text{DL}_2(q), while not Hausdorff, is T1T_1, totally disconnected, and compact. Since $\text{D…

2013-07-08abs ↗pdf ↗

Proposes MR-SNE for multimodal data visualization.

problem Visualizing data from multiple domains with relations across them.
method Extends t-SNE to compute augmented relations and jointly embed them in a low-dimensional space.
result Demonstrates promising performance in visualizing Flickr and Animal with Attributes 2 datasets.

Survey on GANs for generating visual arts, music, and literature.

problem Tackles the challenge of generating art using GANs.
method Uses generative adversarial networks (GANs) to generate visual arts, music, and literary text.
result Performance comparison and description of various GAN architectures presented.

NegToMe uses images to guide text-based models away from unwanted visual elements.

problem Insufficient text-based adversarial guidance for complex visual concepts.
method Negative token merging (NegToMe) using visual features from reference images.
result Significantly enhances output diversity and reduces visual similarity to copyrighted content.

ZegOT uses optimal transport to zero-shot segment images with text prompts.

problem Zero-shot semantic segmentation with limited image-text alignment knowledge.
method ZegOT uses optimal transport to match multiple text prompts with frozen image embeddings.
result ZegOT achieves state-of-the-art performance in zero-shot semantic segmentation.

Visuals in scientific papers are used to express complex ideas; this study uses them to identify knowledge domains.

problem Scientific figures are underutilized in literature analysis.
method Encoded scientific figures into visual signatures and used distances between signatures to compare communities of practice.
result Figures can differentiate knowledge domains as effectively as text or citation patterns.

Tri-modal model predicts personality traits from audio, visual, and text data.

problem Predicting personality traits from multimodal data.
method Stacked Convolutional Neural Networks for each modality, multimodal fusion on decision-level and by concatenation.
result Multimodal fusion outperforms individual modalities, improving personality trait prediction by 9.4%.

We investigate ways in which to improve the interpretability of LDA topic models by better analyzing and visualizing their outputs. We focus on examining what we refer to as topic similarity networks: graphs in which nodes represent latent topics in text collections and links represent similarity among topics. We descr…

2014-09-26abs ↗pdf ↗

In the wake of the still ongoing global financial crisis, bank interdependencies have come into focus in trying to assess linkages among banks and systemic risk. To date, such analysis has largely been based on numerical data. By contrast, this study attempts to gain further insight into bank interconnections by tappin…

2014-06-30abs ↗pdf ↗

Visual spoofing bypasses spam filters and plagiarism detection.

problem Vulnerability in spam filters that can be exploited by visually similar but differently encoded characters.
method Replaces characters with visually similar but differently encoded characters from a different alphabet.
result Spammers can create messages that bypass existing spam filters.

Paper tackles zero-shot activity recognition using video features and text embeddings.

problem Zero-shot activity recognition with videos.
method Auto-encoder model for multimodal joint embedding, 3D convolutional action recognition for visual features, GloVe word embeddings for textual features.
result Improved zero-shot recognition results with top-n accuracy and mean Nearest Neighbor Overlap.

Paper proposes ICCN to learn correlations between text, audio, and video for multimodal sentiment analysis.

problem Improving multimodal sentiment analysis by learning hidden correlations between text and audio/video features.
method Interaction Canonical Correlation Network (ICCN) using deep canonical correlation analysis (DCCA).
result Empirical results confirm the effectiveness of ICCN in capturing useful information from all three views.

Method retrieves similar fashion items from images and text, enabling style refinement.

problem Lack of intuitive, interactive refinement in search engines for fashion items.
method Joint visual-textual embedding training, Mini-Batch Match Retrieval, attribute extraction.
result Improved performance in multimodal style search, demonstrated through benchmark.

Improved real-time visualizations of conversation turns using dynamic attention weights.

problem Uniform attention weights in sequential analysis tasks prevent meaningful visualization.
method Developed a method to track changes in turn importance over time.
result More informative real-time visuals confirmed by human reviewers.

Generative model generates realistic text without reinforcement learning.

problem Training GANs for natural language processing is challenging due to discrete sequences.
method Autoencoder learns low-dimensional sentence representation, GAN generates vectors in this space.
result Model generates realistic text with competitive BLEU scores and human ratings.

Reduces memory costs for developing countries by replacing images with text.

problem High memory costs associated with multimodal webpages in developing countries.
method Canonical Correlation Analysis (CCA) to replace high-cost modality (images) with low-cost modality (text).
result Reduces memory costs by at least 83.35% through eye-tracking experiments.

Twitmo analyzes geo-tagged Twitter data for topic modeling and visualization.

problem Analyzing public discourse on Twitter for various topics, parties, or individuals.
method Collects and preprocesses geo-tagged Tweets, applies LDA, CTM, STM, and visualizes results.
result Automatic pooling of Tweets into pseudo-documents improves topic coherence.

Adversarial training method improved word embeddings in text classification.

problem Improving semi-supervised text classification performance.
method Adapted adversarial and virtual adversarial training to word embeddings in RNNs.
result Achieved state-of-the-art results on semi-supervised and supervised tasks.

We describe a new method for visualizing topics, the distributions over terms that are automatically extracted from large text corpora using latent variable models. Our method finds significant nn-grams related to a topic, which are then used to help understand and interpret the underlying distribution. Compared with …

2009-07-06abs ↗pdf ↗

The paper introduces multimodal generative models to improve data marginal likelihood.

problem Improving data marginal likelihood in multimodal settings.
method Derives variational bounds on the evidence for multimodal deep generative models, generalizes objectives for different model types, and benchmarks across various datasets.
result Multimodal VAEs excel in image, label, and text datasets with and without weak supervision.

SeqAttnGAN generates interactive images based on multi-turn text descriptions.

problem Interactive image editing with multi-turn textual commands.
method SeqAttnGAN uses a neural state tracker and GAN framework for sequential image generation and refinement.
result SeqAttnGAN outperforms state-of-the-art models on interactive image editing tasks.

We introduce SARR for symmetric object pose estimation, improving CNN performance.

problem Ambiguities in symmetric object orientations hinder deep learning pose estimation.
method Numeric rotation representation using symmetry-derived trigonometric identities.
result SARR enables standard CNNs to achieve state-of-the-art performance.

Paper develops a framework for generating coherent image captions using visual features and hierarchical topics.

problem Generating semantically coherent paragraphs to describe image content.
method Plug-and-play hierarchical-topic-guided image paragraph generation framework integrating visual extractor and deep topic model.
result Proposed models can distill interpretable multi-layer semantic topics and generate diverse and coherent captions.