Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

3567111,0671,422 · Jun 202019922001200920172026
48 results for Spanish language models

SpanishTinyRoBERTa distills large Spanish models into efficient question-answering models.

problem Efficient Spanish question-answering models for resource-constrained environments.
method Knowledge distillation from large Spanish language models onto a smaller model.
result SpanishTinyRoBERTa achieves comparable performance to large models with faster inference.

We perform a large-scale analysis of language diatopic variation using geotagged microblogging datasets. By collecting all Twitter messages written in Spanish over more than two years, we build a corpus from which a carefully selected list of concepts allows us to characterize Spanish varieties on a global scale. A clu…

2014-07-26abs ↗pdf ↗

Information extraction is an important task in NLP, enabling the automatic extraction of data for relational database filling. Historically, research and data was produced for English text, followed in subsequent years by datasets in Arabic, Chinese (ACE/OntoNotes), Dutch, Spanish, German (CoNLL evaluations), and many …

2019-11-19abs ↗pdf ↗

This paper maps the large-scale variation of the Spanish language by employing a corpus based on geographically tagged Twitter messages. Lexical dialects are extracted from an analysis of variants of tens of concepts. The resulting maps show linguistic variation on an unprecedented scale across the globe. We discuss th…

2015-11-16abs ↗pdf ↗

We describe a simple neural language model that relies only on character-level inputs. Predictions are still made at the word-level. Our model employs a convolutional neural network (CNN) and a highway network over characters, whose output is given to a long short-term memory (LSTM) recurrent neural network language mo…

2015-08-26abs ↗pdf ↗

Develops a cross-lingual hate speech detection model using pre-trained Transformers.

problem Detecting hate speech in low-resource languages.
method Utilizes frozen Transformer language models and AXEL attention-based classification block for zero-shot and few-shot learning.
result Demonstrates highly competitive results on English and Spanish subsets of the HatEval challenge.

This work tackles the problem of learning a set of language specific acoustic units from unlabeled speech recordings given a set of labeled recordings from other languages. Our approach may be described by the following two steps procedure: first the model learns the notion of acoustic units from the labelled data and …

2019-04-08abs ↗pdf ↗

We present a recurrent encoder-decoder deep neural network architecture that directly translates speech in one language into text in another. The model does not explicitly transcribe the speech into text in the source language, nor does it require supervision from the ground truth source language transcription during t…

2017-03-24abs ↗pdf ↗

Study on Spanish households' investment choices in housing, deposits, and stocks.

problem Investment decisions of Spanish households in housing, deposits, and stocks.
method Theoretical model considering indivisible and illiquid housing assets, financial constraints, and actual choices compared.
result Households underinvest in stocks and deposits compared to optimal choices, but mortgage investments are efficient.

Recently, sentiment analysis has received a lot of attention due to the interest in mining opinions of social media users. Sentiment analysis consists in determining the polarity of a given text, i.e., its degree of positiveness or negativeness. Traditionally, Sentiment Analysis algorithms have been tailored to a speci…

2016-12-15abs ↗pdf ↗

Sentiment analysis (SA) is a task related to understanding people's feelings in written text; the starting point would be to identify the polarity level (positive, neutral or negative) of a given text, moving on to identify emotions or whether a text is humorous or not. This task has been the subject of several researc…

2018-11-29abs ↗pdf ↗

Case study shows impact of co-optimizing energy and reserve for wind energy.

problem Impact of lack of co-optimization of energy and reserve in high wind penetration scenarios.
method Developed two models with and without co-optimization, calibrated with Spanish market parameters.
result Models show significant differences in energy and reserve management.

In this paper, we identify an interesting kind of error in the output of Unsupervised Neural Machine Translation (UNMT) systems like \textit{Undreamt}(footnote). We refer to this error type as \textit{Scrambled Translation problem}. We observe that UNMT models which use \textit{word shuffle} noise (as in case of Undrea…

2019-10-30abs ↗pdf ↗

Study examines barriers to grid-connected battery systems in Spain, finding high cycle cost remains main obstacle.

problem Barriers to grid-connected battery systems in Spain's deregulated electricity market.
method Utilization analysis and concept of 'potentially profitable utilization time' introduced.
result High cycle cost remains the main barrier for grid-connected battery systems in Spain.

The study compares VaR and ES models for tail risk of electricity futures, finding AR(1)-GARCH(1,1) with Student-t distribution best.

problem Modeling tail risk of electricity futures contracts in various markets.
method Comparison of VaR and ES models using AR(1)-GARCH(1,1) with Student-t distribution, historical simulation, and quantile regression.
result AR(1)-GARCH(1,1) with Student-t distribution is the best-performing model for tail risk estimation.

We propose a novel deep learning architecture suitable for the prediction of investor interest for a given asset in a given time frame. This architecture performs both investor clustering and modelling at the same time. We first verify its superior performance on a synthetic scenario inspired by real data and then appl…

2019-09-11abs ↗pdf ↗

This paper introduces compositional data analysis for financial ratios, improving industry-level analysis.

problem Statistical issues with standard financial ratios at industry level.
method Compositional data analysis techniques for financial ratios.
result Improved analysis of financial ratios using compositional data methods.

Study on time-zero efficiency of European power derivatives markets using statistical tests and trading rules.

problem Assessing time-zero efficiency in European power derivatives markets.
method Statistical tests based on the law of one price and trading rules based on price differentials and no-arbitrage violations applied to daily data of three European power markets.
result Definite conclusions on time-zero efficiency are not possible for French and Spanish markets due to liquidity and representativeness challenges.

We present an analysis of the price impact associated with trades effected by different financial firms. Using data from the Spanish Stock Market, we find a high degree of heterogeneity across different market members, both in the instantaneous impact functions and in the time-dependent market response to trades by ind…

2011-09-01abs ↗pdf ↗

THieF improves day-ahead electricity price prediction accuracy by reconciling hourly and block forecasts.

problem Improving accuracy in predicting day-ahead electricity prices.
method Temporal hierarchy forecasting (THieF) reconciling hourly and block forecasts.
result THieF significantly improves accuracy (up to 13%) at all levels of prediction.

We analyze an exhaustive data-set of new-cars monthly sales. The set refers to 10 years of Spanish sales of more than 6500 different car model configurations and a total of 10M sold cars, from January 2007 to January 2017. We find that for those model configurations with a monthly market-share higher than 0.1% the sale…

2017-05-09abs ↗pdf ↗

We present a novel method for solving Canonical Correlation Analysis (CCA) in a sparse convex framework using a least squares approach. The presented method focuses on the scenario when one is interested in (or limited to) a primal representation for the first view while having a dual representation for the second view…

2009-08-19abs ↗pdf ↗

New vine copula method forecasts portfolio risk measures robust to market downturns.

problem Inaccurate risk measure estimation for financial portfolios due to lack of cross-dependency capture.
method Combines vine copulas with ARMA-GARCH models for marginal risk estimation.
result Portfolio is robust to American market downturns but not European market.

These are Lecture Notes of a course given by the author at the French-Spanish School "Tresses in Pau", held in Pau (France) in October 2009. It is basically an introduction to distinct approaches and techniques that can be used to show results in braid groups. Using these techniques we provide several proofs of well kn…

2010-10-02abs ↗pdf ↗

New ASR system handles multiple languages without needing language-specific encoding.

problem Joint training of data-rich and data-scarce languages in a single model.
method Transforms all languages to a single writing system through transliteration, separating modeling and rendering.
result Language-agnostic multilingual ASR system reduces WER up to 10% over language-dependent models.

The paper applies math and physics to language models, introducing entropy and geometric concepts.

problem Understanding and improving language models to approximate intelligent language.
method Formal definitions, functional analysis, topology, thermodynamics, and set theory.
result Entropy function reveals key obstacles for LLMs and offers insights into language models.

Paper reviews neurolinguistics and language technologies, emphasizing mutual enrichment.

problem Understanding brain activity during language processing.
method Brain imaging studies and natural language representations.
result Development of brain-aware natural language representations.