Deep learning predicts readmissions from less structured data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Many methods have been used to recognize author personality traits from text, typically combining linguistic feature engineering with shallow learning models, e.g. linear regression or Support Vector Machines. This work uses deep-learning-based models and atomic features of text, the characters, to build hierarchical, …
SpanishTinyRoBERTa distills large Spanish models into efficient question-answering models.
We perform a large-scale analysis of language diatopic variation using geotagged microblogging datasets. By collecting all Twitter messages written in Spanish over more than two years, we build a corpus from which a carefully selected list of concepts allows us to characterize Spanish varieties on a global scale. A clu…
Sentiment analysis (SA) is a task related to understanding people's feelings in written text; the starting point would be to identify the polarity level (positive, neutral or negative) of a given text, moving on to identify emotions or whether a text is humorous or not. This task has been the subject of several researc…
Study on Spanish households' investment choices in housing, deposits, and stocks.
Unified BERT model improves NER across multiple languages.
We present a recurrent encoder-decoder deep neural network architecture that directly translates speech in one language into text in another. The model does not explicitly transcribe the speech into text in the source language, nor does it require supervision from the ground truth source language transcription during t…
This paper maps the large-scale variation of the Spanish language by employing a corpus based on geographically tagged Twitter messages. Lexical dialects are extracted from an analysis of variants of tens of concepts. The resulting maps show linguistic variation on an unprecedented scale across the globe. We discuss th…
This paper describes the participation of Amobee in the shared sentiment analysis task at SemEval 2018. We participated in all the English sub-tasks and the Spanish valence tasks. Our system consists of three parts: training task-specific word embeddings, training a model consisting of gated-recurrent-units (GRU) with …
Recently, sentiment analysis has received a lot of attention due to the interest in mining opinions of social media users. Sentiment analysis consists in determining the polarity of a given text, i.e., its degree of positiveness or negativeness. Traditionally, Sentiment Analysis algorithms have been tailored to a speci…
Author profiling is the characterization of an author through some key attributes such as gender, age, and language. In this paper, a RNN model with Attention (RNNwA) is proposed to predict the gender of a twitter user using their tweets. Both word level and tweet level attentions are utilized to learn 'where to look'.…
In this paper we extend the concept of Competitivity Graph to compare series of rankings with ties ({\em partial rankings}). We extend the usual method used to compute Kendall's coefficient for two partial rankings to the concept of evolutive Kendall's coefficient for a series of partial rankings. The theoretical frame…
Case study shows impact of co-optimizing energy and reserve for wind energy.
End-to-end Sanskrit TTS developed with limited data, achieving good quality.
Study examines barriers to grid-connected battery systems in Spain, finding high cycle cost remains main obstacle.
Intensive development of urban systems creates a number of challenges for urban planners and policy makers in order to maintain sustainable growth. Running efficient urban policies requires meaningful urban metrics, which could quantify important urban characteristics including various aspects of an actual human behavi…
Study on time-zero efficiency of European power derivatives markets using statistical tests and trading rules.
We propose a novel deep learning architecture suitable for the prediction of investor interest for a given asset in a given time frame. This architecture performs both investor clustering and modelling at the same time. We first verify its superior performance on a synthetic scenario inspired by real data and then appl…
New L0 norm added to TDA for market analysis.
Survey on non-positively curved cube complexes and geometric group theory.
The understanding of complex social or economic systems is an important scientific challenge. Here we present a comprehensive study of the Spanish Stock Exchange showing that most financial firms trading in that market are characterized by a resulting strategy and can be classified in groups of firms with different spe…
The study compares VaR and ES models for tail risk of electricity futures, finding AR(1)-GARCH(1,1) with Student-t distribution best.
This paper introduces compositional data analysis for financial ratios, improving industry-level analysis.
We present an analysis of the price impact associated with trades effected by different financial firms. Using data from the Spanish Stock Market, we find a high degree of heterogeneity across different market members, both in the instantaneous impact functions and in the time-dependent market response to trades by ind…
Develops a cross-lingual hate speech detection model using pre-trained Transformers.
We present a novel method for solving Canonical Correlation Analysis (CCA) in a sparse convex framework using a least squares approach. The presented method focuses on the scenario when one is interested in (or limited to) a primal representation for the first view while having a dual representation for the second view…
These are Lecture Notes of a course given by the author at the French-Spanish School "Tresses in Pau", held in Pau (France) in October 2009. It is basically an introduction to distinct approaches and techniques that can be used to show results in braid groups. Using these techniques we provide several proofs of well kn…
Proposes a multilingual email segmentation benchmark and model.
Morphology in unbalanced languages remains a big challenge in the context of machine translation. In this paper, we propose to de-couple machine translation from morphology generation in order to better deal with the problem. We investigate the morphology simplification with a reasonable trade-off between expected gain…
We describe a simple neural language model that relies only on character-level inputs. Predictions are still made at the word-level. Our model employs a convolutional neural network (CNN) and a highway network over characters, whose output is given to a long short-term memory (LSTM) recurrent neural network language mo…
We empirically study the market impact of trading orders. We are specifically interested in large trading orders that are executed incrementally, which we call hidden orders. These are reconstructed based on information about market member codes using data from the Spanish Stock Market and the London Stock Exchange. We…
Financial markets are systems with the complex behavior, that can be hardly analyzed by means of linear methods. Recurrence Quantification Analysis (RQA) is a nonlinear methodology, which is able to work with the nonstationary and short data series. Thus, we apply RQA for the studying of the critical events on financia…
The objective of this paper is to fill a gap in the literature on internationalization, in relation to the absence of objective and measurable performance indicators on the process of how firms sequentially enter external markets. To that end, this research develops a quantitative tool that can be used as a performance…
THieF improves day-ahead electricity price prediction accuracy by reconciling hourly and block forecasts.
Study reveals clusters of resilient and vulnerable Spanish agri-food firms post-Ukraine-Russia war.
Scene text magnifier aims to magnify text in natural scene images without recognition. It could help the special groups, who have myopia or dyslexia to better understand the scene. In this paper, we design the scene text magnifier through interacted four CNN-based networks: character erasing, character extraction, char…
The paper classifies actions of a specific group on certain manifolds.
This work tackles the problem of learning a set of language specific acoustic units from unlabeled speech recordings given a set of labeled recordings from other languages. Our approach may be described by the following two steps procedure: first the model learns the notion of acoustic units from the labelled data and …
Bayesian-Deep Learning model predicts Covid-19 evolution in Spain.
W-RNN improves text classification by extracting serialized text semantics.
The paper solves the Nielsen realization problem for high degree del Pezzo surfaces.
We establish a gluing construction for Higgs bundles over a connected sum of Riemann surfaces in terms of solutions to the -Hitchin equations using the linearization of a relevant elliptic operator. The construction can be used to provide model Higgs bundles in all the exce…
A new text representation model combines CNN and VAE for better semantic extraction.
We propose two algorithms that can find local minima faster than the state-of-the-art algorithms in both finite-sum and general stochastic nonconvex optimization. At the core of the proposed algorithms is using stochastic nested variance reduction (Zhou et al., 2018a), which outperforms the s…
We analyze an exhaustive data-set of new-cars monthly sales. The set refers to 10 years of Spanish sales of more than 6500 different car model configurations and a total of 10M sold cars, from January 2007 to January 2017. We find that for those model configurations with a monthly market-share higher than 0.1% the sale…
New families of Lie groups with special foliations discovered.
Improved text summarization using belief propagation on weighted bipartite graphs.