Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,932 papers · 148 categories

Trend · papers per month

12.5%25.0%37.5%50.0% · Jul 199319922001200920172026
48 results for word order preservation

Let (W,S) be a finite rank Coxeter system with W infinite. We prove that the limit weak order on the blocks of infinite reduced words of W is encoded by the topology of the Tits boundary of the Davis complex X of W. We consider many special cases, including W word hyperbolic, and X with isolated flats. We establish tha…

2013-01-05abs ↗pdf ↗

Proposes a novel framework for multi-label text classification.

problem Lack of coherent consideration of non-consecutive and long-distance semantics and hierarchical relations among labels.
method Hierarchical taxonomy-aware and attentional graph capsule recurrent CNNs framework.
result Significantly improves multi-label text classification performance.

The paper presents a method to preserve privacy in text analysis using calibrated noise.

problem Accurately learning from user data while maintaining privacy.
method Calibrated multivariate perturbations applied to word embeddings to achieve geo-indistinguishability.
result The method provides better privacy guarantees than baseline models with minimal utility loss.

The goal of homomorphic encryption is to encrypt data such that another party can operate on it without being explicitly exposed to the content of the original data. We introduce an idea for a privacy-preserving transformation on natural language data, inspired by homomorphic encryption. Our primary tool is {\em obfusc…

2019-04-21abs ↗pdf ↗

Sequence-to-Sequence (seq2seq) modeling has rapidly become an important general-purpose NLP tool that has proven effective for many text-generation and sequence-labeling tasks. Seq2seq builds on deep neural language modeling and inherits its remarkable accuracy in estimating local, next-word distributions. In this work…

2016-06-09abs ↗pdf ↗

Word embedding models have become a fundamental component in a wide range of Natural Language Processing (NLP) applications. However, embeddings trained on human-generated corpora have been demonstrated to inherit strong gender stereotypes that reflect social constructs. To address this concern, in this paper, we propo…

2018-08-29abs ↗pdf ↗

New method neutralizes gender bias in word embeddings without losing semantic information.

problem Gender biases in word embeddings trained on human-generated corpora.
method Latent Disentanglement and Counterfactual Generation with siamese auto-encoder and gradient reversal layer.
result Our method outperforms existing debiasing methods in preserving semantic information and neutralizing gender biases.

Standard sequential generation methods assume a pre-specified generation order, such as text generation methods which generate words from left to right. In this work, we propose a framework for training models of text generation that operate in non-monotonic orders; the model directly learns good orders, without any ad…

2019-02-05abs ↗pdf ↗

LPL optimizes embeddings to align local neighborhoods, improving cross-lingual word alignment.

problem Aligning embeddings across different datasets and languages.
method Locality Preserving Loss (LPL) optimizes model to project embeddings while maintaining local neighborhoods and aligning them.
result LPL-based alignment leads to better and consistent accuracy, especially in small training set settings.

We show that the Goldman flows preserve the holomorphic structure on the moduli space of homomorphisms of the fundamental group of a Riemann surface into U(1), in other words the Jacobian.

2008-02-24abs ↗pdf ↗

For text analysis, one often resorts to a lossy representation that either completely ignores word order or embeds each word as a low-dimensional dense feature vector. In this paper, we propose convolutional Poisson factor analysis (CPFA) that directly operates on a lossless representation that processes the words in e…

2019-05-14abs ↗pdf ↗

Top2Vec finds topic vectors from documents and words without needing stop words or custom settings.

problem Topic modeling weaknesses, including needing known topics, stop words, and custom settings.
method Joint document and word semantic embedding to find topic vectors automatically.
result Top2Vec finds more informative and representative topics than probabilistic models.

Most popular word embedding techniques involve implicit or explicit factorization of a word co-occurrence based matrix into low rank factors. In this paper, we aim to generalize this trend by using numerical methods to factor higher-order word co-occurrence based arrays, or \textit{tensors}. We present four word embedd…

2017-04-10abs ↗pdf ↗

A quantum model classifies financial sentiment by mapping text chunks to quantum circuits.

problem Classifying financial texts with high accuracy and preserving semantic information.
method Chunked diagrams are mapped to quantum circuits, with a Transformer encoder and type embeddings added for context.
result The hybrid model improves sentiment classification over a simple averaging baseline.

We propose an unsupervised object matching method for relational data, which finds matchings between objects in different relational datasets without correspondence information. For example, the proposed method matches documents in different languages in multi-lingual document-word networks without dictionaries nor ali…

2018-10-09abs ↗pdf ↗

By representing words with probability densities rather than point vectors, probabilistic word embeddings can capture rich and interpretable semantic information and uncertainty. The uncertainty information can be particularly meaningful in capturing entailment relationships -- whereby general words such as "entity" co…

2018-04-26abs ↗pdf ↗

Improves dialogue response model interpretability using attention and regularization.

problem Improving interpretability of dual encoder models for dialogue response suggestions.
method Integrates attention mechanism and novel regularization loss to emphasize important words.
result Improves model accuracy and interpretability compared to existing methods.

Paper reproduces and enhances a method for cross-lingual word embeddings.

problem Creating robust cross-lingual mappings of word embeddings without supervision.
method Reproduces and enhances a self-learning method with grid search for hyperparameters.
result Model's robustness is demonstrated across four new languages.

A method for authorship attribution based on function word adjacency networks (WANs) is introduced. Function words are parts of speech that express grammatical relationships between other words but do not carry lexical meaning on their own. In the WANs in this paper, nodes are function words and directed edges stand in…

2014-06-17abs ↗pdf ↗

The paper proposes a privacy-preserving method for text data using Hyperbolic space.

problem Preserving user privacy in text data while maintaining utility for machine learning.
method Word representations in Hyperbolic space to provide privacy, sampling from a probability distribution.
result Demonstrates significant privacy guarantees (20x greater) compared to Euclidean space.

In topic modeling, many algorithms that guarantee identifiability of the topics have been developed under the premise that there exist anchor words -- i.e., words that only appear (with positive probability) in one topic. Follow-up work has resorted to three or higher-order statistics of the data corpus to relax the an…

2016-11-15abs ↗pdf ↗

In the field of Natural Language Processing (NLP), we revisit the well-known word embedding algorithm word2vec. Word embeddings identify words by vectors such that the words' distributional similarity is captured. Unexpectedly, besides semantic similarity even relational similarity has been shown to be captured in word…

2018-06-20abs ↗pdf ↗

Paper proposes AXE loss for non-autoregressive machine translation, improving performance.

problem Challenges in training non-autoregressive models due to lack of autoregressive factors and cross entropy loss penalties.
method Proposes aligned cross entropy (AXE) loss function using a differentiable dynamic program for better word order alignment.
result AXE-based training improves performance on major WMT benchmarks and sets a new state of the art for non-autoregressive models.

In this article we introduce order preserving representations of fundamental groups of surfaces into Lie groups with bi-invariant orders. By relating order preserving representations to weakly maximal representations, introduced in arXiv:1305.2620, we show that order preserving representations into Lie groups of Hermit…

2016-01-10abs ↗pdf ↗

We study stable commutator length (scl) in free products via surface maps into a wedge of spaces. We prove that scl is piecewise rational linear if it vanishes on each factor of the free product, generalizing the main result in Danny Calegari's paper "Scl, sails and surgery". We further prove that the property of isome…

2016-11-22abs ↗pdf ↗

In this paper we prove that an isometry between orbit spaces of two proper isometric actions is smooth if it preserves the codimension of the orbits or if the orbit spaces have no boundary. In other words, we generalize Myers-Steenrod's theorem for orbit spaces. These results are proved in the more general context of s…

2011-11-26abs ↗pdf ↗

This paper tackles rare word problem in low-resource language pairs using NMT.

problem Rare word problem in neural machine translation, especially for low-resource languages.
method Three solutions: enhanced source context, morphology learning, and wordnet synonyms.
result Significant improvements in BLEU scores (+1.0 points) on English-Vietnamese and Japanese-Vietnamese.

Many loss functions in representation learning are invariant under a continuous symmetry transformation. For example, the loss function of word embeddings (Mikolov et al., 2013) remains unchanged if we simultaneously rotate all word and context embedding vectors. We show that representation learning models for time ser…

2018-03-08abs ↗pdf ↗

A scalable framework preserves personalized higher-order network proximities.

problem Lack of expressive methods to preserve personalized higher-order network proximities.
method Incorporates random walk into a sound objective to preserve arbitrary higher-order proximities and introduces random walk with restart for personalized-weighted preservation.
result Consistently and substantially outperforms state-of-the-art methods on real-world networks.

The problem of topic modeling can be seen as a generalization of the clustering problem, in that it posits that observations are generated due to multiple latent factors (e.g., the words in each document are generated as a mixture of several active topics, as opposed to just one). This increased representational power …

2012-04-30abs ↗pdf ↗

This research explores using kernels in the softmax layer for better contextual word classification.

problem Improving contextual word classification accuracy.
method Replacing the inner product in the softmax layer with various kernel functions and comparing their performance.
result Different kernel settings yield varying performance in contextual word classification tasks.

Attention layers are sensitive to single words, improving generalization over random features.

problem Understanding why attention layers are effective in NLP tasks.
method Study of word sensitivity in random features using BERT-Base word embeddings.
result Attention layers have high word sensitivity, improving generalization over random features.

We prove that any regular Casimir in 3D magnetohydrodynamics is a function of the magnetic helicity and cross-helicity. In other words, these two helicities are the only independent regular integral invariants of the coadjoint action of the MHD group SDiff(M)X(M)\text{SDiff}(M)\ltimes\mathfrak X^*(M), which is the semidirect pro…

2019-01-14abs ↗pdf ↗