Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

18365371 · May 202619922001200920172026
48 results for lexical ratio

Paper introduces lexical ratio to measure portfolio diversification.

problem Traditional diversification metrics overlook non-numerical relationships.
method Uses textual data to capture diversification dimensions through entropy-based insights.
result Lexical ratio (LR) outperforms traditional metrics in optimizing portfolio returns.

Measuring the distance between concepts is an important field of study of Natural Language Processing, as it can be used to improve tasks related to the interpretation of those same concepts. WordNet, which includes a wide variety of concepts associated with words (i.e., synsets), is often used as a source for computin…

2018-04-24abs ↗pdf ↗

Proposes a framework for compositional generalization in language models.

problem Lack of compositional generalization in neural networks compared to humans.
method Introduces Generalized Grammar Rules (GGRs) for transduction tasks, formalizing symmetry-based constraints.
result Framework enables models to generalize compositionally, similar to human learning.

This paper maps the large-scale variation of the Spanish language by employing a corpus based on geographically tagged Twitter messages. Lexical dialects are extracted from an analysis of variants of tens of concepts. The resulting maps show linguistic variation on an unprecedented scale across the globe. We discuss th…

2015-11-16abs ↗pdf ↗

In this work we approach the task of learning multilingual word representations in an offline manner by fitting a generative latent variable model to a multilingual dictionary. We model equivalent words in different languages as different views of the same word generated by a common latent variable representing their l…

2019-05-14abs ↗pdf ↗

This paper investigates the influence of different acoustic features, audio-events based features and automatic speech translation based lexical features in complex emotion recognition such as curiosity. Pretrained networks, namely, AudioSet Net, VoxCeleb Net and Deep Speech Net trained extensively for different speech…

2018-10-31abs ↗pdf ↗

Generating paraphrases that are lexically similar but semantically different is a challenging task. Paraphrases of this form can be used to augment data sets for various NLP tasks such as machine reading comprehension and question answering with non-trivial negative examples. In this article, we propose a deep variatio…

2019-11-27abs ↗pdf ↗

We perform a large-scale analysis of language diatopic variation using geotagged microblogging datasets. By collecting all Twitter messages written in Spanish over more than two years, we build a corpus from which a carefully selected list of concepts allows us to characterize Spanish varieties on a global scale. A clu…

2014-07-26abs ↗pdf ↗

Study explains Zipf's law using geometric mechanisms from a finite alphabet.

problem Explains Zipf's law in language without relying on linguistic elements.
method Uses the Full Combinatorial Word Model (FCWM) to generate geometric distributions of word lengths.
result Supports predictions of power-law rank-frequency curves, matching various languages.

Part-of-speech (POS) tagging is a fundamental component for performing natural language tasks such as parsing, information extraction, and question answering. When POS taggers are trained in one domain and applied in significantly different domains, their performance can degrade dramatically. We present a methodology f…

2014-10-31abs ↗pdf ↗

Proposes a method to ensure low losses across all subpopulations in large datasets.

problem Standard practice of minimizing average loss fails to guarantee low losses across all subpopulations in heterogeneous datasets.
method Convex procedure that controls worst-case performance over all subpopulations of a given size with finite-sample convergence guarantees.
result Empirically, the worst-case procedure learns models that do well against unseen subpopulations.

The study distills news sources to analyze stock reactions, finding sentiment has asymmetric and sector-specific effects.

problem Analyzing the influence of financial text sources on stock reactions.
method Mixed text sources from professional platforms, blogs, and message boards were distilled using different lexica to analyze sentiment variables.
result Sentiment has an asymmetric and sector-specific effect on stock reactions.

The process of translation is ambiguous, in that there are typically many valid trans- lations for a given sentence. This gives rise to significant variation in parallel cor- pora, however, most current models of machine translation do not account for this variation, instead treating the prob- lem as a deterministic pr…

2018-05-28abs ↗pdf ↗

By representing words with probability densities rather than point vectors, probabilistic word embeddings can capture rich and interpretable semantic information and uncertainty. The uncertainty information can be particularly meaningful in capturing entailment relationships -- whereby general words such as "entity" co…

2018-04-26abs ↗pdf ↗

Annotating temporal relations (TempRel) between events described in natural language is known to be labor intensive, partly because the total number of TempRels is quadratic in the number of events. As a result, only a small number of documents are typically annotated, limiting the coverage of various lexical/semantic …

2018-04-18abs ↗pdf ↗

Text style transfer aims to modify the style of a sentence while keeping its content unchanged. Recent style transfer systems often fail to faithfully preserve the content after changing the style. This paper proposes a structured content preserving model that leverages linguistic information in the structured fine-gra…

2018-10-15abs ↗pdf ↗

Pre-trained word embeddings encode general word semantics and lexical regularities of natural language, and have proven useful across many NLP tasks, including word sense disambiguation, machine translation, and sentiment analysis, to name a few. In supervised tasks such as multiclass text classification (the focus of …

2019-11-26abs ↗pdf ↗

Word vectors are at the core of many natural language processing tasks. Recently, there has been interest in post-processing word vectors to enrich their semantic information. In this paper, we introduce a novel word vector post-processing technique based on matrix conceptors (Jaeger2014), a family of regularized ident…

2018-11-17abs ↗pdf ↗

Recently, sentiment analysis has received a lot of attention due to the interest in mining opinions of social media users. Sentiment analysis consists in determining the polarity of a given text, i.e., its degree of positiveness or negativeness. Traditionally, Sentiment Analysis algorithms have been tailored to a speci…

2016-12-15abs ↗pdf ↗

We discuss - in what is intended to be a pedagogical fashion - generalized "mean-to-risk" ratios for portfolio optimization. The Sharpe ratio is only one example of such generalized "mean-to-risk" ratios. Another example is what we term the Fano ratio (which, unlike the Sharpe ratio, is independent of the time horizon)…

2017-11-29abs ↗pdf ↗

Optimal option portfolios under Sharpe Ratio maximization with skew-elliptical t-distributed returns

problem Optimal option portfolios under Sharpe Ratio maximization
method Formulation for explicit portfolio weights
result Different optimal portfolios for Sharpe Ratio and return-to-Value-at-Risk (VaR) ratio

Omega ratio, defined as the probability-weighted ratio of gains over losses at a given level of expected return, has been advocated as a better performance indicator compared to Sharpe and Sortino ratio as it depends on the full return distribution and hence encapsulates all information about risk and return. We comput…

2019-10-15abs ↗pdf ↗

We present a new methodology of computing incremental contribution for performance ratios for portfolio like Sharpe, Treynor, Calmar or Sterling ratios. Using Euler's homogeneous function theorem, we are able to decompose these performance ratios as a linear combination of individual modified performance ratios. This a…

2018-07-13abs ↗pdf ↗

The paper proposes an asset allocation strategy using the Sortino ratio for better performance.

problem Traditional asset allocation methods like the Sharpe ratio do not penalize negative returns adequately.
method The Sortino ratio is used to maximize asset allocation, penalizing only negative return variances.
result The Sortino ratio-based strategy outperforms traditional methods like the Kelly criterion.