Study maps Spanish dialects on Twitter, revealing urban vs regional variations.
problem Understanding linguistic diversity across Spanish-speaking regions.
method Geographically tagged Twitter corpus, machine learning for dialect analysis.
result Urban dialects show international characteristics, rural dialects regional uniformity.
We perform a large-scale analysis of language diatopic variation using geotagged microblogging datasets. By collecting all Twitter messages written in Spanish over more than two years, we build a corpus from which a carefully selected list of concepts allows us to characterize Spanish varieties on a global scale. A clu…
Paper tackles Arabic question similarity, outperforming state-of-the-art.
problem Detecting semantically similar questions in Arabic is challenging.
method Utilizes contextualized word representations (ELMo embeddings) trained on MSA and dialectic sentences, combined with a pairwise similarity layer.
result Achieves 93% F1-score on Modern Standard Arabic benchmark and 82% on dialectical benchmark.
It is argued that arguments for strict prohibition of interests must be based on the use of arguments from authority. This is carried out by first making a survey of so-called dialectical roots for interest prohibition and then demonstrating that for at least one important positive interest bearing financial product, t…
SpanishTinyRoBERTa distills large Spanish models into efficient question-answering models.
problem Efficient Spanish question-answering models for resource-constrained environments.
method Knowledge distillation from large Spanish language models onto a smaller model.
result SpanishTinyRoBERTa achieves comparable performance to large models with faster inference.
Paper surveys and introduces Acoustic Dialect Decoder for voice translation.
problem Machine understanding of natural language in speech translation.
method Recognition, Translation, and Synthesis units using HMMs, RNNs, and HTS.
result Initial successful translation of English to Tamil.
Graphemes outperform phonemes in end-to-end models for English Voice-search and multi-dialect tasks.
problem Comparing phoneme-based and grapheme-based sub-word units in end-to-end models.
method Detailed experiments comparing phoneme-based and grapheme-based end-to-end models on large vocabulary English Voice-search and multi-dialect tasks.
result Graphemes outperform phonemes in end-to-end models for English Voice-search and multi-dialect tasks.
ArSentD-LEV dataset improves sentiment analysis in Levantine Arabic tweets.
problem Challenges in sentiment analysis of Arabic tweets, especially Levantine dialect.
method Created a dataset of 4,000 Levantine Arabic tweets with detailed sentiment and topic annotations.
result Improved performance of sentiment classifiers with detailed annotations.
MPSA-DenseNet improves accent classification accuracy.
problem Accurate English accent identification.
method Combines multi-task learning and PSA attention mechanism with DenseNet.
result MPSA-DenseNet outperforms other models in accent classification.
Study on Spanish households' investment choices in housing, deposits, and stocks.
problem Investment decisions of Spanish households in housing, deposits, and stocks.
method Theoretical model considering indivisible and illiquid housing assets, financial constraints, and actual choices compared.
result Households underinvest in stocks and deposits compared to optimal choices, but mortgage investments are efficient.
Paper tackles morphology simplification for Chinese-Spanish machine translation.
problem Challenges in morphology generation for unbalanced languages in machine translation.
method Proposes a new neural architecture for morphological simplification, combining embedding, convolutional, and recurrent neural network layers.
result Obtains over 98% accuracy in gender classification, over 93% in number classification, and an overall translation improvement of 0.7 METEOR.
Amobee won 3rd and 1st place in SemEval 2018 sentiment classification tasks.
problem Sentiment classification in multiple languages.
method Training GRU-CNN model with word embeddings and stacking ensembles.
result 3rd and 1st place in valence ordinal classification sub-tasks in English and Spanish.
In this paper we extend the concept of Competitivity Graph to compare series of rankings with ties ({\em partial rankings}). We extend the usual method used to compute Kendall's coefficient for two partial rankings to the concept of evolutive Kendall's coefficient for a series of partial rankings. The theoretical frame…
Improved neural model predicts gender from tweets.
problem Predicting gender from Twitter text.
method RNN model with attention, LSA-reduced n-gram features.
result Improved model achieves state-of-the-art performance on English tweets.
Case study shows impact of co-optimizing energy and reserve for wind energy.
problem Impact of lack of co-optimization of energy and reserve in high wind penetration scenarios.
method Developed two models with and without co-optimization, calibrated with Spanish market parameters.
result Models show significant differences in energy and reserve management.
We explain the meaning of local symmetries in physics.
problem Understanding the meaning of local symmetries in physics.
method We argue that general covariance and gauge principles are principles of epistemic access to physical laws, leading to ontological insights.
result Relationality is a core notion in gauge field theory, encoded by local symmetries.
Develops a tool to measure gradual internationalization performance.
problem Lack of objective performance indicators for gradual internationalization.
method Quantitative tool based on export data, tested in Spanish wine sector.
result Creation of an international priority index for analyzing geographically differentiated strategies.
The starting point of this article is the question "How to retrieve fingerprints of rhythm in written texts?" We address this problem in the case of Brazilian and European Portuguese. These two dialects of Modern Portuguese share the same lexicon and most of the sentences they produce are superficially identical. Yet t…
Study examines barriers to grid-connected battery systems in Spain, finding high cycle cost remains main obstacle.
problem Barriers to grid-connected battery systems in Spain's deregulated electricity market.
method Utilization analysis and concept of 'potentially profitable utilization time' introduced.
result High cycle cost remains the main barrier for grid-connected battery systems in Spain.
Neural network framework for language recognition considers sequence information and improves accuracy.
problem Challenging task of automatic language identification in noisy conditions.
method Proposes a neural network framework with bidirectional LSTM and attention modeling for relevance weighting.
result Significant improvements over conventional methods in noisy conditions and multi-speaker speech.
Intensive development of urban systems creates a number of challenges for urban planners and policy makers in order to maintain sustainable growth. Running efficient urban policies requires meaningful urban metrics, which could quantify important urban characteristics including various aspects of an actual human behavi…
Deep learning predicts readmissions from less structured data.
problem Predicting readmissions from non-standard, unstructured medical records.
method Proposes a deep learning architecture that handles less structured data, including Spanish text.
result Achieves AUROC of 0.76 on a Chilean medical dataset, comparable to US results.
Study on time-zero efficiency of European power derivatives markets using statistical tests and trading rules.
problem Assessing time-zero efficiency in European power derivatives markets.
method Statistical tests based on the law of one price and trading rules based on price differentials and no-arbitrage violations applied to daily data of three European power markets.
result Definite conclusions on time-zero efficiency are not possible for French and Spanish markets due to liquidity and representativeness challenges.
In the era of deep learning several unsupervised models have been developed to capture the key features in unlabeled handwritten data. Popular among them is the Restricted Boltzmann Machines RBM. However, due to the novelty in handwritten multidialect data, the RBM may fail to generate an efficient representation. In t…
New L0 norm added to TDA for market analysis.
problem Improving TDA tools for market prediction.
method Defined and applied L0 norm in TDA for four markets.
result Enhanced TDA tools for market analysis.
Unified BERT model improves NER across multiple languages.
problem Language-specific NER models limit data extraction.
method Jointly trained multilingual BERT with regularization.
result Unified model outperforms monolingual models on various datasets.
Survey on non-positively curved cube complexes and geometric group theory.
problem None explicitly stated, focuses on introduction.
method Lecture notes and mini-course teaching.
result Introduction to non-positively curved cube complexes and geometric group theory.
The understanding of complex social or economic systems is an important scientific challenge. Here we present a comprehensive study of the Spanish Stock Exchange showing that most financial firms trading in that market are characterized by a resulting strategy and can be classified in groups of firms with different spe…
The study compares VaR and ES models for tail risk of electricity futures, finding AR(1)-GARCH(1,1) with Student-t distribution best.
problem Modeling tail risk of electricity futures contracts in various markets.
method Comparison of VaR and ES models using AR(1)-GARCH(1,1) with Student-t distribution, historical simulation, and quantile regression.
result AR(1)-GARCH(1,1) with Student-t distribution is the best-performing model for tail risk estimation.
This paper introduces compositional data analysis for financial ratios, improving industry-level analysis.
problem Statistical issues with standard financial ratios at industry level.
method Compositional data analysis techniques for financial ratios.
result Improved analysis of financial ratios using compositional data methods.
Deep learning predicts investor interest in stocks.
problem Predicting investor interest in stocks.
method Supervised clustering deep learning architecture.
result Superior performance on synthetic and real-world data.
Estimates labels from bagged data with label proportion constraints.
problem Estimating individual labels from aggregated data with label proportion constraints.
method Relaxed proportion constraints, bag differences constraints, intuitive formulation and algorithm.
result Achieves high accuracy in various domains (income level, sentiment analysis, geographical differences in dialect).
Improved LID for multilingual speakers using context-aware models.
problem Low accuracy for languages spoken by multilingual speakers, especially with accented speech.
method Coarser-grained acoustic model and integration with interaction context signals.
result Average 97% accuracy across all language combinations, 60% improvement in worst-case accuracy.
New GE2E loss improves speaker verification efficiency and accuracy.
problem Improving speaker verification accuracy and efficiency.
method Proposed a new loss function (GE2E) that updates the network to emphasize difficult examples.
result Decreased EER by more than 10% and reduced training time by 60%.
Model learns accent patterns from small datasets.
problem Automatic speech recognition struggles with rare accents.
method Automatically retrieves phonological generalizations from a small dataset.
result Model generates a million phonological variations of words.
We present an analysis of the price impact associated with trades effected by different financial firms. Using data from the Spanish Stock Market, we find a high degree of heterogeneity across different market members, both in the instantaneous impact functions and in the time-dependent market response to trades by ind…
Develops a cross-lingual hate speech detection model using pre-trained Transformers.
problem Detecting hate speech in low-resource languages.
method Utilizes frozen Transformer language models and AXEL attention-based classification block for zero-shot and few-shot learning.
result Demonstrates highly competitive results on English and Spanish subsets of the HatEval challenge.
CL methods improve monolingual ASR models across new tasks without forgetting past data.
problem Catastrophic Forgetting in monolingual ASR models when adapting to new domains or accents.
method Implement and compare various Continual Learning methods for monolingual ASR.
result Best CL method reduces performance gap by over 40% with minimal past data.
We present a novel method for solving Canonical Correlation Analysis (CCA) in a sparse convex framework using a least squares approach. The presented method focuses on the scenario when one is interested in (or limited to) a primal representation for the first view while having a dual representation for the second view…
These are Lecture Notes of a course given by the author at the French-Spanish School "Tresses in Pau", held in Pau (France) in October 2009. It is basically an introduction to distinct approaches and techniques that can be used to show results in braid groups. Using these techniques we provide several proofs of well kn…
UQE uses LLMs to analyze unstructured data efficiently.
problem Efficient analytics on unstructured data.
method Proposes UQE, a query engine that uses LLMs to interpret UQL queries.
result Demonstrates efficient analytics on various unstructured data types.
Deep-learning model detects personality traits from short texts across multiple languages.
problem Recognizing author personality traits from short texts.
method Uses deep-learning-based models and atomic features of text to build hierarchical, vectorial word and sentence representations.
result Shows state-of-the-art performance across five traits and three languages (English, Spanish, and Italian).
Proposes a multilingual email segmentation benchmark and model.
problem Lack of multilingual email zoning corpora and models.
method Analysis of existing corpora, development of multilingual benchmark, introduction of OKAPI model.
result OKAPI model achieves state-of-the-art performance in English and generalizes well to unseen languages.
We describe a simple neural language model that relies only on character-level inputs. Predictions are still made at the word-level. Our model employs a convolutional neural network (CNN) and a highway network over characters, whose output is given to a long short-term memory (LSTM) recurrent neural network language mo…
We empirically study the market impact of trading orders. We are specifically interested in large trading orders that are executed incrementally, which we call hidden orders. These are reconstructed based on information about market member codes using data from the Spanish Stock Market and the London Stock Exchange. We…
Financial markets are systems with the complex behavior, that can be hardly analyzed by means of linear methods. Recurrence Quantification Analysis (RQA) is a nonlinear methodology, which is able to work with the nonstationary and short data series. Thus, we apply RQA for the studying of the critical events on financia…
THieF improves day-ahead electricity price prediction accuracy by reconciling hourly and block forecasts.
problem Improving accuracy in predicting day-ahead electricity prices.
method Temporal hierarchy forecasting (THieF) reconciling hourly and block forecasts.
result THieF significantly improves accuracy (up to 13%) at all levels of prediction.
Study reveals clusters of resilient and vulnerable Spanish agri-food firms post-Ukraine-Russia war.
problem Financial resilience of agri-food companies in Spain during the Ukraine-Russia conflict.
method Cluster analysis using centred log-ratios for compositional data of financial ratios.
result Increase in resilient firms by 2023, highlighting sectoral adaptation to economic challenges.