Citizen science projects are successful at gathering rich datasets for various applications. However, the data collected by citizen scientists are often biased --- in particular, aligned more with the citizens' preferences than with scientific objectives. We propose the Shift Compensation Network (SCN), an end-to-end l…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New model estimates species population trends from citizen science data.
The paper uses a graph autoencoder to learn unbiased plant-pollinator interaction embeddings.
EcoCast predicts biodiversity risks using satellite data and citizen science records.
This paper improves prediction accuracy for multi-input classification tasks using p-value aggregation.
This paper explores a real-world fundamental theme under a data science perspective. It specifically discusses whether fraud or manipulation can be observed in and from municipality income tax size distributions, through their aggregation from citizen fiscal reports. The study case pertains to official data obtained fr…
The Prescriptive Canvas improves business outcomes by directly prescribing actions based on predictions.
In this paper we seek methods to effectively detect urban micro-events. Urban micro-events are events which occur in cities, have limited geographical coverage and typically affect only a small group of citizens. Because of their scale these are difficult to identify in most data sources. However, by using citizen sens…
Bird sound data collected with unattended microphones for automatic surveys, or mobile devices for citizen science, typically contain multiple simultaneously vocalizing birds of different species. However, few works have considered the multi-label structure in birdsong. We propose to use an ensemble of classifier chain…
Understanding how species are distributed across landscapes over time is a fundamental question in biodiversity research. Unfortunately, most species distribution models only target a single species at a time, despite strong ecological evidence that species are not independently distributed. We propose Deep Multi-Speci…
We suggest an analytical approach for Pareto-Zipf law, where we assume random multiplicative noise and fragmentation processes for the growth of the number of citizens of each city and the number of the cities, respectively.
Developed a neural topic model for classifying COVID-19 disinformation.
Machine Learning and Artificial Intelligence are considered an integral part of the Fourth Industrial Revolution. Their impact, and far-reaching consequences, while acknowledged, are yet to be comprehended. These technologies are very specialized, and few organizations and select highly trained professionals have the w…
This paper identifies the salient factors that characterize the inequality income distribution for Romania. Data analysis is rigorously carried out using sophisticated techniques borrowed from classical statistics (Theil). Decomposition of the inequalities measured by the Theil index is also performed. This study relie…
Media seems to have become more partisan, often providing a biased coverage of news catering to the interest of specific groups. It is therefore essential to identify credible information content that provides an objective narrative of an event. News communities such as digg, reddit, or newstrust offer recommendations,…
Mosquitoes are a major vector for malaria, causing hundreds of thousands of deaths in the developing world each year. Not only is the prevention of mosquito bites of paramount importance to the reduction of malaria transmission cases, but understanding in more forensic detail the interplay between malaria, mosquito vec…
The scale of ongoing and future electromagnetic surveys pose formidable challenges to classify astronomical objects. Pioneering efforts on this front include citizen science campaigns adopted by the Sloan Digital Sky Survey (SDSS). SDSS datasets have been recently used to train neural network models to classify galaxie…
Machine learning services pose privacy risks if data or model parameters are compromised.
Ordinal data is omnipresent in almost all multiuser-generated feedback - questionnaires, preferences etc. This paper investigates modelling of ordinal data with Gaussian restricted Boltzmann machines (RBMs). In particular, we present the model architecture, learning and inference procedures for both vector-variate and …
TIMME detects Twitter users' ideology from sparse, heterogeneous data.
DeepMaxent uses neural networks to improve species distribution models.
This study simulates the evolution of artificial economies in order to understand the tax relevance of administrative boundaries in the quality of life of its citizens. The modeling involves the construction of a computational algorithm, which includes citizens, bounded into families; firms and governments; all of them…
Isobenefit Lines can offer a certain range of applicability in Location Theory and Gravitational Models for Urban and Geography Economics, in positional decision processes made by citizens, and, last but not least, in land value and property market theories and analysis. The value of a land, or a property, in a generic…
This paper presents a simple agent-based model of an economic system, populated by agents playing different games according to their different view about social cohesion and tax payment. After a first set of simulations, correctly replicating results of existing literature, a wider analysis is presented in order to stu…
In a densely populated city like Dhaka (Bangladesh), a growing number of high-rise buildings is an inevitable reality. However, they pose mental health risks for citizens in terms of detachment from natural light, sky view, greenery, and environmental landscapes. The housing economy and rent structure in different area…
The world of cryptocurrency is not transparent enough though it was established for innate transparent tracking of capital flows. The most contributing factor is the violation of securities laws and scam in Initial Coin Offering (ICO) which is used to raise capital through crowdfunding. There is a lack of proper regula…
Data Science is currently a popular field of science attracting expertise from very diverse backgrounds. Current learning practices need to acknowledge this and adapt to it. This paper summarises some experiences relating to such learning approaches from teaching a postgraduate Data Science module, and draws some learn…
This paper is the first attempt to formalize a new field of economics; studding the Intangibles Goods available on the Internet. We are taking advantage of the digital world's specific rules, in particular the zero marginal cost, to propose a theory of trading & sharing unified. A function based money is created as a w…
ML methods improve planetary science data analysis.
Machine learning improves wildfire science and management, but requires expert knowledge.
Causal inference from observational data is the goal of many data analyses in the health and social sciences. However, academic statistics has often frowned upon data analyses with a causal objective. The introduction of the term "data science" provides a historic opportunity to redefine data analysis in such a way tha…
The increasing complexity of mobility plus the growing population in cities, together with the importance of privacy when sharing data from vehicles or any device, makes traffic forecasting that uses data from infrastructure and citizens an open and challenging task. In this paper, we introduce a novel approach to deal…
Today, the prominence of data science within organizations has given rise to teams of data science workers collaborating on extracting insights from data, as opposed to individual data scientists working alone. However, we still lack a deep understanding of how data science workers collaborate in practice. In this work…
We propose Edward, a Turing-complete probabilistic programming language. Edward defines two compositional representations---random variables and inference. By treating inference as a first class citizen, on a par with modeling, we show that probabilistic programming can be as flexible and computationally efficient as t…
Foundation models alter medical data science workflow, challenging veridical data science principles.
Back cover text: Megaprojects and Risk provides the first detailed examination of the phenomenon of megaprojects. It is a fascinating account of how the promoters of multibillion-dollar megaprojects systematically and self-servingly misinform parliaments, the public and the media in order to get projects approved and b…
Defines data science as a natural ecosystem with challenges and missions.
This paper presents thirteen datasets for binary, multiclass and multilabel classification based on the European Court of Human Rights judgments since its creation. The interest of such datasets is explained through the prism of the researcher, the data scientist, the citizen and the legal practitioner. Contrarily to m…
Browsing and finding relevant information for Bangladeshi laws is a challenge faced by all law students and researchers in Bangladesh, and by citizens who want to learn about any legal procedure. Some law archives in Bangladesh are digitized, but lack proper tools to organize the data meaningfully. We present a text vi…
I studied what role the US stock markets and money markets have possibly played in the Gross Private Domestic Investment (GPDI) of the United States from the year 1959 to the year 2001, Gross Private Domestic Investment refers to the total amount of investment spending by businesses and firms located within the borders…
Data mining revealed a cluster of economic, psychological, social and cultural indicators that in combination predicted corruption and wealth of European nations. This prosperity syndrome of self-reliant citizens, efficient division of labor, a sophisticated scientific community, and respect for the law, was clearly di…
Data science enhances knot theory by analyzing invariant relations.
Donoho's JCGS (in press) paper is a spirited call to action for statisticians, who he points out are losing ground in the field of data science by refusing to accept that data science is its own domain. (Or, at least, a domain that is becoming distinctly defined.) He calls on writings by John Tukey, Bill Cleveland, and…
Data science models, although successful in a number of commercial domains, have had limited applicability in scientific problems involving complex physical phenomena. Theory-guided data science (TGDS) is an emerging paradigm that aims to leverage the wealth of scientific knowledge for improving the effectiveness of da…
The goal of this article is to inspire data scientists to participate in the debate on the impact that their professional work has on society, and to become active in public debates on the digital world as data science professionals. How do ethical principles (e.g., fairness, justice, beneficence, and non-maleficence) …
OOD-trained Bayesian neural networks perform similarly to frequentist methods in uncertainty quantification.
Heart disease is one of the most common diseases in middle-aged citizens. Among the vast number of heart diseases, the coronary artery disease (CAD) is considered as a common cardiovascular disease with a high death rate. The most popular tool for diagnosing CAD is the use of medical imaging, e.g., angiography. However…
Politicians world-wide frequently promise a better life for their citizens. We find that the probability that a country will increase its {\it per capita} GDP ({\it gdp}) rank within a decade follows an exponential distribution with decay constant . We use the Corruption Perceptions Index (CPI) and the Global …