BERT model improves cross-lingual document retrieval.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Many information retrieval algorithms rely on the notion of a good distance that allows to efficiently compare objects of different nature. Recently, a new promising metric called Word Mover's Distance was proposed to measure the divergence between text passages. In this paper, we demonstrate that this metric can be ex…
\emph{Sentiment Quantification} (i.e., the task of estimating the relative frequency of sentiment-related classes -- such as \textsf{Positive} and \textsf{Negative} -- in a set of unlabelled documents) is an important topic in sentiment analysis, as the study of sentiment-related quantities and trends across a populati…
The authors advocate for more rigorous unsupervised cross-lingual learning methods.
Structural correspondence learning (SCL) is an effective method for cross-lingual sentiment classification. This approach uses unlabeled documents along with a word translation oracle to automatically induce task specific, cross-lingual correspondences. It transfers knowledge through identifying important features, i.e…
Paper develops a new unsupervised scoring function for cross-lingual document alignment.
CCAligned creates a massive web document dataset for cross-lingual research.
Develops a cross-lingual hate speech detection model using pre-trained Transformers.
Study uncovers financial trends from cross-lingual news data.
Paper reproduces and enhances a method for cross-lingual word embeddings.
Paper refines cross-lingual word embeddings using Manhattan norm.
Because it is not feasible to collect training data for every language, there is a growing interest in cross-lingual transfer learning. In this paper, we systematically explore zero-shot cross-lingual transfer learning on reading comprehension tasks with a language representation model pre-trained on multi-lingual corp…
LPL optimizes embeddings to align local neighborhoods, improving cross-lingual word alignment.
Parallel sentences are a relatively scarce but extremely useful resource for many applications including cross-lingual retrieval and statistical machine translation. This research explores our methodology for mining such data from previously obtained comparable corpora. The task is highly practical since non-parallel m…
NMIXX fine-tunes embeddings for finance, outperforming general models in Korean.
Parallel sentences are a relatively scarce but extremely useful resource for many applications including cross-lingual retrieval and statistical machine translation. This research explores our new methodologies for mining such data from previously obtained comparable corpora. The task is highly practical since non-para…
We introduce BilBOWA (Bilingual Bag-of-Words without Alignments), a simple and computationally-efficient model for learning bilingual distributed representations of words which can scale to large monolingual datasets and does not require word-aligned parallel training data. Instead it trains directly on monolingual dat…
Although over 100 languages are supported by strong off-the-shelf machine translation systems, only a subset of them possess large annotated corpora for named entity recognition. Motivated by this fact, we leverage machine translation to improve annotation-projection approaches to cross-lingual named entity recognition…
We solve a key problem in cross-lingual learning using a novel approach.
Simple framework decouples word alignment and multilingual embedding mapping.
A novel OT-based method for aligning hyperbolic representations.
Introduces PELP for graph-enhanced word embeddings.
This paper introduces PyDCI, a new implementation of Distributional Correspondence Indexing (DCI) written in Python. DCI is a transfer learning method for cross-domain and cross-lingual text classification for which we had provided an implementation (here called JaDCI) built on top of JaTeCS, a Java framework for text …
Study quantifies gender bias in language models across 7 languages.
Cross-lingual Text Classification (CLC) consists of automatically classifying, according to a common set C of classes, documents each written in one of a set of languages L, and doing so more accurately than when naively classifying each document via its corresponding language-specific classifier. In order to obtain an…
A framework for multi-label sentiment analysis in 100 languages with dynamic weighting.
Improves retrieval accuracy for hierarchical documents, especially for distant matches.
UPR hybrid model improves phase retrieval performance.
This paper improves image retrieval accuracy through novel relevance feedback methods.
A new deep learning model improves phase retrieval performance.
Signal retrieval from a series of indirect measurements is a common task in many imaging, metrology and characterization platforms in science and engineering. Because most of the indirect measurement processes are well-described by physical models, signal retrieval can be solved with an iterative optimization that enfo…
Transformer models improve query-document retrieval efficiency and accuracy.
On most sponsored search platforms, advertisers bid on some keywords for their advertisements (ads). Given a search request, ad retrieval module rewrites the query into bidding keywords, and uses these keywords as keys to select Top N ads through inverted indexes. In this way, an ad will not be retrieved even if querie…
Introduces MPR to measure and optimize representation across intersectional groups in retrieval.
GMC benchmark isolates retrieval in Transformers, revealing max-margin alignment.
Combines deep learning and iterative methods for robust phase retrieval.
New benchmark for non-rigid 3D human shape retrieval.
Motivation: Public and private repositories of experimental data are growing to sizes that require dedicated methods for finding relevant data. To improve on the state of the art of keyword searches from annotations, methods for content-based retrieval have been proposed. In the context of gene expression experiments, …
Sharp asymptotics derived for phase retrieval and compressed sensing with random generative priors.
A new method, InfoGuide, improves automatic clustering analysis.
Most content-based image retrieval systems consider either one single query, or multiple queries that include the same object or represent the same semantic information. In this paper we consider the content-based image retrieval problem for multiple query images corresponding to different image semantics. We propose a…
For the task of generating complex outputs such as source code, editing existing outputs can be easier than generating complex outputs from scratch. With this motivation, we propose an approach that first retrieves a training example based on the input (e.g., natural language description) and then edits it to the desir…
Research on interference has provided evidence that the formation of dependencies between non-adjacent words relies on a cue-based retrieval mechanism. Two different models can account for one of the main predictions of interference, i.e., a slowdown at a retrieval site, when several items share a feature associated wi…
With the wide development of black-box machine learning algorithms, particularly deep neural network (DNN), the practical demand for the reliability assessment is rapidly rising. On the basis of the concept that `Bayesian deep learning knows what it does not know,' the uncertainty of DNN outputs has been investigated a…
The paper introduces subgraph nomination for finding similar subgraphs in networks.
HybridRAG combines KGs and vector retrieval for financial document Q&A.
This paper improves retrieval for LLMs in financial document Q&A.
Extends phase retrieval methods to handle sensing vector errors.