Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

16324864 · Sep 201919922001200920172026
48 results for text

Scene text magnifier aims to magnify text in natural scene images without recognition. It could help the special groups, who have myopia or dyslexia to better understand the scene. In this paper, we design the scene text magnifier through interacted four CNN-based networks: character erasing, character extraction, char…

2019-06-17abs ↗pdf ↗

The paper classifies actions of a specific group on certain manifolds.

problem Classifying analytic actions of a specific semi-orthogonal group on manifolds.
method Adapting Uchida's construction, the paper explicitly constructs actions on specific manifolds and demonstrates that any action is covered by these.
result Any analytic action of the semi-orthogonal group on a closed, connected manifold is covered by the constructed actions.

The paper solves the Nielsen realization problem for high degree del Pezzo surfaces.

problem Which finite subgroups of the mapping class group of a del Pezzo surface lift to the diffeomorphism group?
method Classification and partial answers for d7d \geq 7, equivariant connected sum for d=6d = 6.
result Complete classification for d7d \geq 7, partial answer for d=6d = 6.

A new text representation model combines CNN and VAE for better semantic extraction.

problem Difficult to effectively extract semantic features and distinguish polysemy in text data.
method Integrates CNN for feature extraction and VAE for consistent Gaussian distribution.
result The model outperforms traditional classification algorithms in text classification tasks.

We propose two algorithms that can find local minima faster than the state-of-the-art algorithms in both finite-sum and general stochastic nonconvex optimization. At the core of the proposed algorithms is One-epoch-SNVRG+\text{One-epoch-SNVRG}^+ using stochastic nested variance reduction (Zhou et al., 2018a), which outperforms the s…

2018-06-22abs ↗pdf ↗

Improved text summarization using belief propagation on weighted bipartite graphs.

problem Text summarization from a graph theory perspective.
method Generalized belief propagation algorithm for weighted bipartite graphs.
result Our algorithm outperforms greedy methods in text summarization tasks.

Let Man\text{Man}_{*} denote the category of closed, connected, oriented and based 33-manifolds, with basepoint preserving diffeomorphisms between them. Juhász, Thurston and Zemke showed that the Heegaard Floer invariants are natural with respect to diffeomorphisms, in the sense that there are functors $HF^{\circ}: \te…

2019-08-17abs ↗pdf ↗

Paper generates diverse, readable adversarial texts from scratch.

problem Text classification models are easily fooled by adversarial examples.
method Trained a conditional variational autoencoder (VAE) with adversarial loss and utilized GANs to generate consistent adversarial texts.
result Successfully generates adversarial texts with higher success rate and acceptable quality.

The paper studies symplectic structures on character varieties of Sasakian threefolds.

problem Character varieties of Sasakian threefolds and their symplectic structures.
method Constructing a natural algebraic 2-form and showing its properties.
result The restriction of the 2-form to the space of irreducible SU(r) homomorphisms is symplectic.

Most of the information is stored as text, so text mining is regarded as having high commercial potential. Aiming at the semantic constraint problem of classification methods based on sparse representation, we propose a weighted recurrent neural network (W-RNN), which can fully extract text serialization semantic infor…

2019-09-28abs ↗pdf ↗

Recent years have seen remarkable progress of text generation in different contexts, such as the most common setting of generating text from scratch, and the emerging paradigm of retrieval-and-rewriting. Text infilling, which fills missing text portions of a sentence or paragraph, is also of numerous use in real life, …

2019-01-01abs ↗pdf ↗

Joyce showed that for a classical knot KK, the involutory medial quandle IMQ(K)\text{IMQ}(K) is isomorphic to the core quandle of the homology group H1(X2)H_1(X_2), where X2X_2 is the cyclic double cover of S3\mathbb S ^3, branched over KK. It follows that IMQ(K)=detK|\text{IMQ}(K)| = | \det K |. In the present paper, the extension o…

2019-02-27abs ↗pdf ↗

Convolutional neural network improves assertion detection in multi-label clinical text.

problem Detecting assertions in multi-label clinical text with rich descriptions.
method Developed a CNN architecture for multi-label scope detection.
result At least 12% improvement over state-of-the-art on multi-label clinical text.

Proposes a flexible neural recommendation framework for better prediction performance.

problem Data sparsity, cold start problem, and long-tail distribution in recommendations.
method A modular neural recommendation framework that includes a neural collaborative filtering part and a text processing part as a regularizer.
result Achieves better prediction performance than state-of-the-art text-aware methods using a simple text processing approach.

Let G(n)\text{G}(n) be equal either to PO(n,1),PU(n,1)\text{PO}(n,1),\text{PU}(n,1) or PSp(n,1)\text{PSp}(n,1) and let ΓG(n)Γ\leq \text{G}(n) be a uniform lattice. Denote by HKn\mathbb{H}^n_K the hyperbolic space associated to G(n)\text{G}(n), where KK is a division algebra over the reals of dimension d=dimRKd=\dim_{\mathbb{R}} K. Assume $d(n-1) \ge…

2019-09-17abs ↗pdf ↗

Traditionally, text generation models take in a sequence of text as input, and iteratively generate the next most probable word using pre-trained parameters. In this work, we propose the architecture to use images instead of text as the input of the text generation model, called StoryGen. In the architecture, we design…

2020-01-16abs ↗pdf ↗

MoleculeSTM learns from molecule structures and texts for better drug design.

problem Lack of integration between chemical structures and textual knowledge in AI drug discovery.
method Jointly learns chemical structures and texts via contrastive learning, using a large dataset.
result MoleculeSTM achieves state-of-the-art performance in zero-shot tasks like structure-text retrieval and molecule editing.

Image captioning has demonstrated models that are capable of generating plausible text given input images or videos. Further, recent work in image generation has shown significant improvements in image quality when text is used as a prior. Our work ties these concepts together by creating an architecture that can enabl…

2018-09-27abs ↗pdf ↗

Estimates watermarked content proportions in mixed-source texts.

problem Optimally estimating the proportion of watermarked content in texts with mixed sources.
method Casting the problem as estimating a proportion parameter in a mixture model based on pivotal statistics.
result Proposes efficient estimators for watermark proportion and shows their accuracy through evaluations.

Bayesian Topic Regression models causal inference with text and numerical data.

problem Causal inference using observational text data with both text and numerical confounders.
method Combines supervised Bayesian topic model with Bayesian regression framework, respecting the Frisch-Waugh-Lovell theorem.
result Joint approach recovers ground truth with lower bias than benchmarks, superior prediction results compared to separate approaches.

SpeakerStew verifies 46 languages with reduced training and inference costs.

problem Speaker verification for 46 languages with smart speaker interactions.
method Pooling multilingual data, triage between text-dependent and text-independent models.
result Training on multiple languages generalizes well and reduces computational requirements.

In this work, we study abstractive text summarization by exploring different models such as LSTM-encoder-decoder with attention, pointer-generator networks, coverage mechanisms, and transformers. Upon extensive and careful hyperparameter tuning we compare the proposed architectures against each other for the abstractiv…

2019-03-24abs ↗pdf ↗

Many challenges in natural language processing require generating text, including language translation, dialogue generation, and speech recognition. For all of these problems, text generation becomes more difficult as the text becomes longer. Current language models often struggle to keep track of coherence for long pi…

2018-10-20abs ↗pdf ↗

We simplify and prove a splitting theorem for mapping class groups of connect sums of S2imesS1S^2 imes S^1.

problem Understanding the structure of mapping class groups of specific 3-manifolds.
method We prove a splitting theorem for the mapping class group of connect sums of S2imesS1S^2 imes S^1, simplifying Laudenbach's original proof.
result The mapping class group of connect sums of S2imesS1S^2 imes S^1 is the semidirect product of Out(Fn)(F_n) by (Z/2)n(\mathbb{Z}/2)^n.

Let XX be a compact connected Riemann surface of genus gg, with g2g \geq 2. For each d<η(X)d <η(X), where η(X)η(X) is the gonality of XX, the symmetric product Symd(X)\text{Sym}^d(X) embeds into Picd(X)\text{Pic}^d(X) by sending an effective divisor of degree dd to the corresponding holomorphic line bundle. Therefore, the restrict…

2016-08-07abs ↗pdf ↗

Let MM and NN be two closed CC^{\infty} manifolds and let Diffc(M)\text{Diff}_c(M) denote the group of CC^{\infty} diffeomorphisms isotopic to the identity. We prove that any (discrete) group homomorphism between Diffc(M)\text{Diff}_c(M) and Diffc(N)\text{Diff}_c(N) is continuous. We also show that a non-trivial group homomorphism $…

2013-07-16abs ↗pdf ↗

Unsupervised text embedding has shown great power in a wide range of NLP tasks. While text embeddings are typically learned in the Euclidean space, directional similarity is often more effective in tasks such as word similarity and document clustering, which creates a gap between the training stage and usage stage of t…

2019-11-04abs ↗pdf ↗

The paper studies extPin± ext{Pin}^{\pm}-structures on non-oriented 4-manifolds via Lefschetz fibrations.

problem Understanding extPin± ext{Pin}^{\pm}-structures on non-oriented 4-manifolds and Lefschetz fibrations.
method Extending work on orientable settings, the paper uses Lefschetz fibrations to analyze extPin± ext{Pin}^{\pm}-structures on non-orientable 4-manifolds and vector bundles.
result The paper provides existence results of extPin+ ext{Pin}^{+} and extPin ext{Pin}^--structures on closed non-orientable 4-manifolds and Lefschetz fibrations over the 2-sphere.

Deep learning models outperform classical methods in text classification.

problem Improving text classification accuracy using deep learning.
method Comprehensive review of deep learning models and datasets for text classification.
result Deep learning models outperform classical methods on various text classification tasks.

Study stabilizers of isotropic classes in rational 4-manifolds, finding diffeomorphisms that almost preserve Lefschetz fibrations.

problem Stabilizers of isotropic classes in rational 4-manifolds.
method Analysis of stabilizers under mapping class group action, finding diffeomorphisms almost preserving Lefschetz fibrations.
result Diffeomorphisms can almost preserve genus-0 Lefschetz fibrations, answering Nielsen realization problem.