Visualizes text models with in-text and word-as-pixel highlighting.
problem Diagnosing and understanding text models, especially for slang and historical texts.
method In-text annotations and word-as-pixel graphics.
result Interconnected methods help diagnose and understand text models.
Proposes CNN Inception + Gate model for visual question answering.
problem Deep understanding of images and texts for visual question answering.
method Proposes a CNN Inception + Gate model for learning textual representations.
result Improves question representations and overall accuracy in visual question answering.
PySS3 simplifies access to SS3's text classification and visualization.
problem Limited availability of an open-source SS3 implementation.
method Developed PySS3, an open-source Python package implementing SS3.
result PySS3 enables robust, explainable, and trusty text classification.
New method evaluates text-to-image synthesis for realism, variety, and semantic accuracy.
problem Lack of metrics revealing semantic accuracy in text-to-image synthesis.
method Uses Inception network representations and t-SNE visualization for semantic evaluation.
result Classification accuracy of generated images to real images' visual concepts correlates with semantic accuracy.
DBNet improves natural language image localization and detection.
problem Natural language-based visual entity localization with limited accuracy.
method Discriminative bimodal neural network (DBNet) trained with extensive negative samples.
result Significantly outperforms previous methods on Visual Genome dataset.
Diestel-Leader graphs are neither hyperbolic nor CAT(0), so their visual boundaries may be pathological. Indeed, we show that for d>2, ∂DLd(q) carries the indiscrete topology. On the other hand, ∂DL2(q), while not Hausdorff, is T1, totally disconnected, and compact. Since $\text{D…
Proposes MR-SNE for multimodal data visualization.
problem Visualizing data from multiple domains with relations across them.
method Extends t-SNE to compute augmented relations and jointly embed them in a low-dimensional space.
result Demonstrates promising performance in visualizing Flickr and Animal with Attributes 2 datasets.
Scene text magnifier enhances readability for visually impaired.
problem Helps visually impaired read natural scene text.
method Four CNN-based networks: character erasing, extraction, magnify, synthesis.
result Effective text magnification without background alteration.
Survey on GANs for generating visual arts, music, and literature.
problem Tackles the challenge of generating art using GANs.
method Uses generative adversarial networks (GANs) to generate visual arts, music, and literary text.
result Performance comparison and description of various GAN architectures presented.
NegToMe uses images to guide text-based models away from unwanted visual elements.
problem Insufficient text-based adversarial guidance for complex visual concepts.
method Negative token merging (NegToMe) using visual features from reference images.
result Significantly enhances output diversity and reduces visual similarity to copyrighted content.
Improves video search by balancing text and visual modalities.
problem Modality imbalance in video search models, focusing mainly on text matching.
method Proposes MBVR with MS samples and DM to balance modalities.
result Empirically shows significant improvement in modality balance and search effectiveness.
ZegOT uses optimal transport to zero-shot segment images with text prompts.
problem Zero-shot semantic segmentation with limited image-text alignment knowledge.
method ZegOT uses optimal transport to match multiple text prompts with frozen image embeddings.
result ZegOT achieves state-of-the-art performance in zero-shot semantic segmentation.
regvis.net offers a visual survey of regulatory visualization.
problem Lack of a comprehensive resource for regulatory visualization.
method Collection and manual tagging of 80+ publications, creation of a searchable webpage.
result First publication set tailored for regulatory visualization.
SAFE detects fake news by analyzing text and image similarities.
problem Detecting fake news with less focus on text-image similarity.
method SAFE uses neural networks to extract text and visual features, then learns their relationship to predict fake news.
result SAFE effectively recognizes fake news based on text, images, or mismatches.
A new method extracts linguistic objects from text using CNNs.
problem Lack of interpretability in deep learning models for text.
method Weighted extension of Text Deconvolution Saliency (wTDS) measure.
result Extracts interpretable linguistic objects from text.
Visuals in scientific papers are used to express complex ideas; this study uses them to identify knowledge domains.
problem Scientific figures are underutilized in literature analysis.
method Encoded scientific figures into visual signatures and used distances between signatures to compare communities of practice.
result Figures can differentiate knowledge domains as effectively as text or citation patterns.
Tri-modal model predicts personality traits from audio, visual, and text data.
problem Predicting personality traits from multimodal data.
method Stacked Convolutional Neural Networks for each modality, multimodal fusion on decision-level and by concatenation.
result Multimodal fusion outperforms individual modalities, improving personality trait prediction by 9.4%.
We investigate ways in which to improve the interpretability of LDA topic models by better analyzing and visualizing their outputs. We focus on examining what we refer to as topic similarity networks: graphs in which nodes represent latent topics in text collections and links represent similarity among topics. We descr…
LAMVI-2 visualizes word embedding model tuning for developers.
problem Tuning deep learning models is complex and time-consuming.
method Introduces LAMVI-2, a visual analytics system for comparing hyperparameter settings.
result LAMVI-2 helps developers quickly and accurately choose effective models.
Taxicab correspondence analysis visualizes sparse text data sets.
problem Visualization of extremely sparse contingency tables.
method Robust variant of correspondence analysis for sparse data.
result Visualized an 8265-dimensional textual data set.
In the wake of the still ongoing global financial crisis, bank interdependencies have come into focus in trying to assess linkages among banks and systemic risk. To date, such analysis has largely been based on numerical data. By contrast, this study attempts to gain further insight into bank interconnections by tappin…
Improves text-to-image generation with bidirectional capabilities.
problem Generating realistic images from text descriptions.
method Integrates text and image modalities using MMVR architecture with n-gram cost function and multiple sentences.
result Significant improvement in image quality over existing methods (over 20%).
Visual spoofing bypasses spam filters and plagiarism detection.
problem Vulnerability in spam filters that can be exploited by visually similar but differently encoded characters.
method Replaces characters with visually similar but differently encoded characters from a different alphabet.
result Spammers can create messages that bypass existing spam filters.
AttViz offers online visualizations of neural language model attention mechanisms.
problem Limited interpretability of neural language models.
method Online toolkit for exploring self-attention mechanisms.
result Visualizations help understand model decision-making.
Paper tackles zero-shot activity recognition using video features and text embeddings.
problem Zero-shot activity recognition with videos.
method Auto-encoder model for multimodal joint embedding, 3D convolutional action recognition for visual features, GloVe word embeddings for textual features.
result Improved zero-shot recognition results with top-n accuracy and mean Nearest Neighbor Overlap.
Generative model controls text attributes for realistic sentences.
problem Challenges in generating natural language sentences with desired attributes.
method Combines variational auto-encoders and holistic attribute discriminators for semantic structure imposition.
result Effective generation of realistic sentences with desired attributes.
Adaptive cross-modal few-shot learning improves performance in image classification.
problem Few-shot classification with limited data.
method Adaptive combination of visual and semantic features.
result Model outperforms uni-modality and modality-alignment methods.
Enhances speech emotion recognition by integrating visual data with attention mechanisms.
problem Improving emotion detection accuracy by combining multiple modalities.
method Introducing an attention mechanism to combine audio, text, and video modalities.
result Significant improvement of 3.65% in weighted accuracy.
Visualizes Bangladeshi laws for quicker searching.
problem Difficulty in finding relevant Bangladeshi laws.
method Doc2Vec for node layout, link mining for citation networks, named entity recognition for quick section finding.
result Users find the tool faster and more intuitive for legal research.
Paper tackles supervision bottleneck in machine learning.
problem Difficulty in generating supervision signals for learning models.
method Describes several learning paradigms to alleviate the supervision bottleneck.
result Illustrates the benefit of these paradigms in inducing semantic representations from text.
Paper proposes ICCN to learn correlations between text, audio, and video for multimodal sentiment analysis.
problem Improving multimodal sentiment analysis by learning hidden correlations between text and audio/video features.
method Interaction Canonical Correlation Network (ICCN) using deep canonical correlation analysis (DCCA).
result Empirical results confirm the effectiveness of ICCN in capturing useful information from all three views.
Method retrieves similar fashion items from images and text, enabling style refinement.
problem Lack of intuitive, interactive refinement in search engines for fashion items.
method Joint visual-textual embedding training, Mini-Batch Match Retrieval, attribute extraction.
result Improved performance in multimodal style search, demonstrated through benchmark.
Improved real-time visualizations of conversation turns using dynamic attention weights.
problem Uniform attention weights in sequential analysis tasks prevent meaningful visualization.
method Developed a method to track changes in turn importance over time.
result More informative real-time visuals confirmed by human reviewers.
Generative model generates realistic text without reinforcement learning.
problem Training GANs for natural language processing is challenging due to discrete sequences.
method Autoencoder learns low-dimensional sentence representation, GAN generates vectors in this space.
result Model generates realistic text with competitive BLEU scores and human ratings.
Reduces memory costs for developing countries by replacing images with text.
problem High memory costs associated with multimodal webpages in developing countries.
method Canonical Correlation Analysis (CCA) to replace high-cost modality (images) with low-cost modality (text).
result Reduces memory costs by at least 83.35% through eye-tracking experiments.
Develops a deep model for joint image-text learning.
problem Bidirectional joint image-text modeling.
method Variational hetero-encoder randomized GAN (VHE-GAN).
result Achieves state-of-the-art performance in image-text learning and generation.
Twitmo analyzes geo-tagged Twitter data for topic modeling and visualization.
problem Analyzing public discourse on Twitter for various topics, parties, or individuals.
method Collects and preprocesses geo-tagged Tweets, applies LDA, CTM, STM, and visualizes results.
result Automatic pooling of Tweets into pseudo-documents improves topic coherence.
Adversarial training method improved word embeddings in text classification.
problem Improving semi-supervised text classification performance.
method Adapted adversarial and virtual adversarial training to word embeddings in RNNs.
result Achieved state-of-the-art results on semi-supervised and supervised tasks.
We describe a new method for visualizing topics, the distributions over terms that are automatically extracted from large text corpora using latent variable models. Our method finds significant n-grams related to a topic, which are then used to help understand and interpret the underlying distribution. Compared with …
The paper introduces multimodal generative models to improve data marginal likelihood.
problem Improving data marginal likelihood in multimodal settings.
method Derives variational bounds on the evidence for multimodal deep generative models, generalizes objectives for different model types, and benchmarks across various datasets.
result Multimodal VAEs excel in image, label, and text datasets with and without weak supervision.
Visual analytics tool detects and corrects concept drift in data streams.
problem Concept drift causes inaccurate predictions in evolving data.
method DriftVis combines drift detection and visualization.
result Visual analytics supports detection, examination, and correction of concept drift.
SeqAttnGAN generates interactive images based on multi-turn text descriptions.
problem Interactive image editing with multi-turn textual commands.
method SeqAttnGAN uses a neural state tracker and GAN framework for sequential image generation and refinement.
result SeqAttnGAN outperforms state-of-the-art models on interactive image editing tasks.
New visual quality index for fuzzy clustering.
problem No accurate quality index for fuzzy clustering across datasets.
method Proposes a new visual quality index and graph-based solution.
result Validated through extensive experiments on various datasets.
We introduce SARR for symmetric object pose estimation, improving CNN performance.
problem Ambiguities in symmetric object orientations hinder deep learning pose estimation.
method Numeric rotation representation using symmetry-derived trigonometric identities.
result SARR enables standard CNNs to achieve state-of-the-art performance.
Voice-controlled e-commerce app enhances user experience for visually impaired.
problem Limited user experience for visually impaired in e-commerce applications.
method Proposes a voice-controlled e-commerce application using IBM Watson speech-to-text.
result Demonstrates enhanced usability for visually impaired users.
Researchers aim to understand TextCNN's learning on NLP datasets.
problem Interpreting TextCNN is challenging due to its black box nature.
method Deep visualization tools are used to understand TextCNN's functions and correlations.
result Functions of different convolutional kernels and correlations between them are explored.
Improves text-to-image diffusion models using GFlowNets.
problem Aligning diffusion models with text descriptions.
method Post-training diffusion models with GFlowNets to generate high-reward images.
result Effective alignment of large-scale text-to-image diffusion models with reward information.
Paper develops a framework for generating coherent image captions using visual features and hierarchical topics.
problem Generating semantically coherent paragraphs to describe image content.
method Plug-and-play hierarchical-topic-guided image paragraph generation framework integrating visual extractor and deep topic model.
result Proposed models can distill interpretable multi-layer semantic topics and generate diverse and coherent captions.