With this work we try to analyse the agglomeration process in the Portuguese regions, using the New Economic Geography models. In these models the base idea is that where has increasing returns to scale in the manufactured industry and low transport costs, there is agglomeration. Of referring, as summary conclusion, th…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Newly available data on the spatial distribution of retail activities in cities makes it possible to build models formalized at the level of the single retailer. Current models tackle consumer location choices at an aggregate level and the opportunity new data offers for modeling at the retail unit level lacks a theore…
The aim of this paper is to analyze the processes of polarization and agglomeration, to explain the mechanisms and causes of these phenomena in order to identify similarities and differences. As the main implication of this study should be noted that both process pretend to explain the concentration of economic activit…
DEFRAG accelerates extreme classification by reducing feature dimensions.
ANN clusters multi-view data by agglomerating subviews and avoiding postprocessing.
The aim of this paper is to analyze the relationship between inter-industry, intra-industry and inter-regional clustering and demand for labor by companies in Portugal. Is expected at the outset that there is more demand for work where the agglomeration is greater. It should be noted, as a summary conclusion, the resul…
Efficiently clusters data with weak assumptions, robust to contamination.
This work aims to study the Portuguese regional agglomeration process, using the linear form the New Economic Geography models that emphasize the importance of spatial factors (distance, costs of transport and communication) in explaining of the concentration of economic activity in certain locations. In a theoretical …
A new text clustering method using NMF and LSA improves stability and performance.
In this work, we revisit fast dimension reduction approaches, as with random projections and random sampling. Our goal is to summarize the data to decrease computational costs and memory footprint of subsequent analysis. Such dimension reduction can be very efficient when the signals of interest have a strong structure…
The paper explores various forms of calibration scores and their implications for fairness.
New bounds improve linkage methods for clustering, distinguishing complete-link from single-link.
Using open source data, we observe the fascinating dynamics of nighttime light. Following a global economic regime shift, the planetary center of light can be seen moving eastwards at a pace of about 60 km per year. Introducing spatial light Gini coefficients, we find a universal pattern of human settlements across dif…
Unified method for simultaneous denoising and clustering.
Infinite mixture models are commonly used for clustering. One can sample from the posterior of mixture assignments by Monte Carlo methods or find its maximum a posteriori solution by optimization. However, in some problems the posterior is diffuse and it is hard to interpret the sampled partitionings. In this paper, we…
Study a simplified model of multiverse structure with synchronized timelines.
Sum-of-norms clustering is a method for assigning points in to clusters, , using convex optimization. Recently, Panahi et al.\ proved that sum-of-norms clustering is guaranteed to recover a mixture of Gaussians under the restriction that the number of samples is not too large. The pu…
Spatial orderness metric improves CNN performance for non-spatial data.
Inverse inference, or "brain reading", is a recent paradigm for analyzing functional magnetic resonance imaging (fMRI) data, based on pattern recognition and statistical learning. By predicting some cognitive variables related to brain activation maps, this approach aims at decoding brain activity. Inverse inference ta…
It has always been a great challenge for clustering algorithms to automatically determine the cluster numbers according to the distribution of datasets. Several approaches have been proposed to address this issue, including the recent promising work which incorporate Bayesian Nonparametrics into the -means clusterin…
Developing a scientific understanding of cities in a fast urbanizing world is essential for planning sustainable urban systems. Recently, it was shown that income and wealth creation follow increasing returns, scaling superlinearly with city size. We study scaling of per capita incomes for separate census defined incom…
With the large-scale penetration of the internet, for the first time, humanity has become linked by a single, open, communications platform. Harnessing this fact, we report insights arising from a unified internet activity and location dataset of an unparalleled scope and accuracy drawn from over a trillion (1.5$\times…
With this study we want to test the validity of the well known "Verdoorn's Law" which considers the relationship between the growth of productivity and output in the case of the Portuguese economy at a regional and sectoral levels (NUTs II) for the period 1995-1999. The importance of some additional variables in the or…
This study predicts stock prices using various machine and deep learning models.
CAF-HFCM automatically forms a cluster hierarchy and optimizes the number of clusters without trial-and-validation.
Clarifies contributions of model architectures and resources in SOTA language models.
Data preprocessing improves data quality for robust data mining.
Big data sets must be carefully partitioned into statistically similar data subsets that can be used as representative samples for big data analysis tasks. In this paper, we propose the random sample partition (RSP) data model to represent a big data set as a set of non-overlapping data subsets, called RSP data blocks,…
A new method for handling imbalanced big data using ensembles and smart data.
Prevents sensitive data generation in diffusion models using labeled and unlabeled data.
Study reveals Data Shapley's inconsistent performance in data selection tasks.
PRRO generates synthetic tabular data that improves SL performance and class distribution.
Defines data science as a natural ecosystem with challenges and missions.
Synthetic data enhances analytics but requires careful volume management.
This paper introduces C-DSL to improve data mining outcomes by considering context.
Proposes using probabilistic models for privacy-preserving synthetic data.
New test ensures quality of shared data in machine learning.
Paper creates fair synthetic data ensuring equal predictions across sensitive attributes.
A new method classifies multiple correlated data streams simultaneously.
DPA preserves data distribution in reduced dimensions.
Efficient synthetic data generation improves model performance on tabular data.
For most problems in science and engineering we can obtain data sets that describe the observed system from various perspectives and record the behavior of its individual components. Heterogeneous data sets can be collectively mined by data fusion. Fusion can focus on a specific target relation and exploit directly ass…
DAERNN models censored data using neural networks with data augmentation.
Data preprocessing techniques are devoted to correct or alleviate errors in data. Discretization and feature selection are two of the most extended data preprocessing techniques. Although we can find many proposals for static Big Data preprocessing, there is little research devoted to the continuous Big Data problem. A…
Data collection is a major bottleneck in machine learning and an active research topic in multiple communities. There are largely two reasons data collection has recently become a critical issue. First, as machine learning is becoming more widely-used, we are seeing new applications that do not necessarily have enough …
This paper quantifies uncertainty in Data Shapley using statistical inference.
DCoM uses deep neural networks to detect semantic data types from raw column values.
Causal inference from observational data is the goal of many data analyses in the health and social sciences. However, academic statistics has often frowned upon data analyses with a causal objective. The introduction of the term "data science" provides a historic opportunity to redefine data analysis in such a way tha…