Computational model uncovers linguistic universals.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
In this paper we present a review of the existing typologies of Internet service users. We zoom in on social networking services including blogs and crowdsourcing websites. Based on the results of the analysis of the considered typologies obtained by means of FCA we developed a new user typology of a certain class of I…
ATD measures language distance using neural models, recovering linguistic groupings.
Performance metrics (error measures) are vital components of the evaluation frameworks in various fields. The intention of this study was to overview of a variety of performance metrics and approaches to their classification. The main goal of the study was to develop a typology that will help to improve our knowledge a…
We investigated the impact of noisy linguistic features on the performance of a Japanese speech synthesis system based on neural network that uses WaveNet vocoder. We compared an ideal system that uses manually corrected linguistic features including phoneme and prosodic information in training and test sets against a …
A new method extracts linguistic objects from text using CNNs.
GTI network learns linguistic features for multi-task sequence tagging.
Model shows how confidence feedback can lead to different crisis outcomes.
The starting point of this article is the question "How to retrieve fingerprints of rhythm in written texts?" We address this problem in the case of Brazilian and European Portuguese. These two dialects of Modern Portuguese share the same lexicon and most of the sentences they produce are superficially identical. Yet t…
Study investigates how simple speech sounds can form abstract categories.
Linguistic calibration improves long-form text confidence.
The abstract discusses parallels between Galois theory and Stone-Weierstrass theorem in various fields.
BERT captures linguistic features in separate semantic and syntactic subspaces.
Our usage of language is not solely reliant on cognition but is arguably determined by myriad external factors leading to a global variability of linguistic patterns. This issue, which lies at the core of sociolinguistics and is backed by many small-scale studies on face-to-face communication, is addressed here by cons…
ContextBench benchmarks methods for generating linguistically fluent inputs that activate specific latent features in language models.
We study the problem of interpreting trained classification models in the setting of linguistic data sets. Leveraging a parse tree, we propose to assign least-squares based importance scores to each word of an instance by exploiting syntactic constituency structure. We establish an axiomatic characterization of these i…
Convolutional neural networks have been successfully applied to various NLP tasks. However, it is not obvious whether they model different linguistic patterns such as negation, intensification, and clause compositionality to help the decision-making process. In this paper, we apply visualization techniques to observe h…
Recent research in psycholinguistics has provided increasing evidence that humans predict upcoming content. Prediction also affects perception and might be a key to robustness in human language processing. In this paper, we investigate the factors that affect human prediction by building a computational model that can …
One of the most prevalent symptoms among the elderly population, dementia, can be detected by classifiers trained on linguistic features extracted from narrative transcripts. However, these linguistic features are impacted in a similar but different fashion by the normal aging process. Aging is therefore a confounding …
We define general linguistic intelligence as the ability to reuse previously acquired knowledge about a language's lexicon, syntax, semantics, and pragmatic conventions to adapt to new tasks quickly. Using this definition, we analyze state-of-the-art natural language understanding models and conduct an extensive empiri…
The insufficient understanding of the credit network structure was recognized as a key factor for regulators' underestimation of the destructive systematic risk during the financial crisis that started in 2007. The existing credit network research either took a macro perspective to clarify the topological properties of…
Investigates neural TTS systems for Japanese and English.
Enhances speech quality in noisy environments using symbolic sequential modeling.
Qwant Research improves clinical case matching and information retrieval.
This paper studies users' perception regarding a controversial product, namely self-driving (autonomous) cars. To find people's opinion regarding this new technology, we used an annotated Twitter dataset, and extracted the topics in positive and negative tweets using an unsupervised, probabilistic model known as topic …
Paper proposes MMD-Sense-Analysis for detecting word sense shifts.
We introduce Markov substitute processes, a new model at the crossroad of statistics and formal grammars, and prove its main property : Markov substitute processes with a given support form an exponential family.
Interpersonal relations are fickle, with close friendships often dissolving into enmity. In this work, we explore linguistic cues that presage such transitions by studying dyadic interactions in an online strategy game where players form alliances and break those alliances through betrayal. We characterize friendships …
Analysts use vague language in reports to convey useful information about future payoffs.
Study shows mutual information can reward structure learning agents without expert systems.
This is the first of two articles in which we provide detailed and self-contained account of the construction of a system of Kuranishi structures on the moduli spaces of pseudo holomorphic disks, using the exponential decay estimate given in [FOOO7]. This article completes the construction of a Kuranishi structure of a…
Study uses social media to analyze COVID-19 impact.
We present the Bayesian Echo Chamber, a new Bayesian generative model for social interaction data. By modeling the evolution of people's language usage over time, this model discovers latent influence relationships between them. Unlike previous work on inferring influence, which has primarily focused on simple temporal…
OT domain adaptation improves aphasia detection across languages.
EigenNoise provides a competitive word vector initialization scheme without pre-training data.
New RL framework learns task completion without prior knowledge.
Paper predicts Indian stocks using news psycholinguistic features.
New distress dictionary improves bankruptcy prediction from disclosure text.
This paper analyzes text in financial disclosures to improve financial analysis.
This paper proposes a new method to connect language and physical actions in reinforcement learning.
In this paper, we propose a new kernel-based co-occurrence measure that can be applied to sparse linguistic expressions (e.g., sentences) with a very short learning time, as an alternative to pointwise mutual information (PMI). As well as deriving PMI from mutual information, we derive this new measure from the Hilbert…
This paper presents an agent-based artificial cryptocurrency market in which heterogeneous agents buy or sell cryptocurrencies, in particular Bitcoins. In this market, there are two typologies of agents, Random Traders and Chartists, which interact with each other by trading Bitcoins. Each agent is initially endowed wi…
We study how to leverage off-the-shelf visual and linguistic data to cope with out-of-vocabulary answers in visual question answering task. Existing large-scale visual datasets with annotations such as image class labels, bounding boxes and region descriptions are good sources for learning rich and diverse visual conce…
Group discussions are essential for organizing every aspect of modern life, from faculty meetings to senate debates, from grant review panels to papal conclaves. While costly in terms of time and organization effort, group discussions are commonly seen as a way of reaching better decisions compared to solutions that do…
This paper presents sampling-based speech parameter generation using moment-matching networks for Deep Neural Network (DNN)-based speech synthesis. Although people never produce exactly the same speech even if we try to express the same linguistic and para-linguistic information, typical statistical speech synthesis pr…
Study reduces redundant information in multi-modal datasets.
Paper introduces methods for more reliable probabilistic predictions with confidence intervals.
Word meaning changes over time, depending on linguistic and extra-linguistic factors. Associating a word's correct meaning in its historical context is a central challenge in diachronic research, and is relevant to a range of NLP tasks, including information retrieval and semantic search in historical texts. Bayesian m…