Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,932 papers · 148 categories

Trend · papers per month

10.7%21.3%32.0%42.6% · Jun 202019922001200920172026
48 results for Data Agglomeration

Newly available data on the spatial distribution of retail activities in cities makes it possible to build models formalized at the level of the single retailer. Current models tackle consumer location choices at an aggregate level and the opportunity new data offers for modeling at the retail unit level lacks a theore…

2016-12-16abs ↗pdf ↗

The aim of this paper is to analyze the processes of polarization and agglomeration, to explain the mechanisms and causes of these phenomena in order to identify similarities and differences. As the main implication of this study should be noted that both process pretend to explain the concentration of economic activit…

2011-10-25abs ↗pdf ↗

DEFRAG accelerates extreme classification by reducing feature dimensions.

problem High precision and scalability in assigning labels from a vast label space.
method Adaptive feature agglomeration to reduce feature dimensions.
result Significant reduction in training and prediction times (up to 40%) for extreme classification algorithms.

The aim of this paper is to analyze the relationship between inter-industry, intra-industry and inter-regional clustering and demand for labor by companies in Portugal. Is expected at the outset that there is more demand for work where the agglomeration is greater. It should be noted, as a summary conclusion, the resul…

2011-10-25abs ↗pdf ↗

Efficiently clusters data with weak assumptions, robust to contamination.

problem General-shaped clustering under weak parametric assumptions with data contamination.
method Two-step hybrid robust clustering algorithm combining trimmed k-means and hierarchical agglomeration.
result Outperforms state-of-the-art methods in various applications.

This work aims to study the Portuguese regional agglomeration process, using the linear form the New Economic Geography models that emphasize the importance of spatial factors (distance, costs of transport and communication) in explaining of the concentration of economic activity in certain locations. In a theoretical …

2011-10-25abs ↗pdf ↗

A new text clustering method using NMF and LSA improves stability and performance.

problem Text data's large, sparse term-document matrix makes clustering difficult.
method Proposes a new feature agglomeration method based on NMF and deterministic K-Means initialization.
result Significantly improves clustering performance and stability.

The paper explores various forms of calibration scores and their implications for fairness.

problem The evaluation of probabilistic predictions through calibration.
method The authors organize three grouping choices and one agglomeration of group errors, providing a framework for comparing and creating new calibration scores.
result The study demonstrates that appropriate choices of grouping can provide notions of (sub-)group or individual fairness.

New bounds improve linkage methods for clustering, distinguishing complete-link from single-link.

problem Improving bounds on linkage methods for clustering quality.
method Developed new bounds for complete-link and average-link methods in agglomeration clustering.
result Separated complete-link from single-link in terms of approximation for diameter.

Using open source data, we observe the fascinating dynamics of nighttime light. Following a global economic regime shift, the planetary center of light can be seen moving eastwards at a pace of about 60 km per year. Introducing spatial light Gini coefficients, we find a universal pattern of human settlements across dif…

2013-03-12abs ↗pdf ↗

Infinite mixture models are commonly used for clustering. One can sample from the posterior of mixture assignments by Monte Carlo methods or find its maximum a posteriori solution by optimization. However, in some problems the posterior is diffuse and it is hard to interpret the sampled partitionings. In this paper, we…

2013-10-01abs ↗pdf ↗

Study a simplified model of multiverse structure with synchronized timelines.

problem Understanding the complex structure of the local multiverse with multiple universes.
method Time-amalgamated globally hyperbolic model of multiverse as a collection of parallel universes.
result Elementary particles are transcosmic strings with multiple endpoints on parallel universes.

Sum-of-norms clustering is a method for assigning nn points in Rd\mathbb{R}^d to KK clusters, 1Kn1\le K\le n, using convex optimization. Recently, Panahi et al.\ proved that sum-of-norms clustering is guaranteed to recover a mixture of Gaussians under the restriction that the number of samples is not too large. The pu…

2019-02-19abs ↗pdf ↗

Inverse inference, or "brain reading", is a recent paradigm for analyzing functional magnetic resonance imaging (fMRI) data, based on pattern recognition and statistical learning. By predicting some cognitive variables related to brain activation maps, this approach aims at decoding brain activity. Inverse inference ta…

2011-05-02abs ↗pdf ↗

It has always been a great challenge for clustering algorithms to automatically determine the cluster numbers according to the distribution of datasets. Several approaches have been proposed to address this issue, including the recent promising work which incorporate Bayesian Nonparametrics into the kk-means clusterin…

2013-06-13abs ↗pdf ↗

Developing a scientific understanding of cities in a fast urbanizing world is essential for planning sustainable urban systems. Recently, it was shown that income and wealth creation follow increasing returns, scaling superlinearly with city size. We study scaling of per capita incomes for separate census defined incom…

2015-09-03abs ↗pdf ↗

With this study we want to test the validity of the well known "Verdoorn's Law" which considers the relationship between the growth of productivity and output in the case of the Portuguese economy at a regional and sectoral levels (NUTs II) for the period 1995-1999. The importance of some additional variables in the or…

2011-10-25abs ↗pdf ↗

This study predicts stock prices using various machine and deep learning models.

problem Predicting stock price movements is challenging but possible.
method Agglomerative approach combining statistical, machine learning, and deep learning models.
result Deep learning models outperform traditional methods in stock price prediction.

CAF-HFCM automatically forms a cluster hierarchy and optimizes the number of clusters without trial-and-validation.

problem Challenges in determining the optimal number of clusters in fuzzy c-means.
method CAF-HFCM, an auto-fused hierarchical fuzzy c-means method.
result Automatic agglomeration and optimal number of clusters without validity indices.

Clarifies contributions of model architectures and resources in SOTA language models.

problem Difficulty in disentangling contributions of model architectures and resources.
method Overview of large pre-trained language models, focusing on their use of new architectures and resources.
result Identifies potential starting points for benchmark comparisons and areas for improvement.

Big data sets must be carefully partitioned into statistically similar data subsets that can be used as representative samples for big data analysis tasks. In this paper, we propose the random sample partition (RSP) data model to represent a big data set as a set of non-overlapping data subsets, called RSP data blocks,…

2017-12-12abs ↗pdf ↗

Prevents sensitive data generation in diffusion models using labeled and unlabeled data.

problem Generating sensitive data in diffusion models using unlabeled data.
method Positive-Unlabeled Diffusion Models, approximating ELBO with labeled and unlabeled data.
result Prevents the generation of sensitive data without compromising image quality.

Study reveals Data Shapley's inconsistent performance in data selection tasks.

problem Inconsistency of Data Shapley's performance in data selection across different settings.
method Hypothesis testing framework and identification of utility functions.
result Data Shapley's performance is no better than random selection without specific constraints.

PRRO generates synthetic tabular data that improves SL performance and class distribution.

problem Low SL utility of synthetic data due to class imbalance and overlooked data relationships.
method Data pruning and column reordering to optimize SL utility.
result Synthetic data generated with PRRO enhances predictive performance and class distribution.

Defines data science as a natural ecosystem with challenges and missions.

problem Challenges and missions in data science due to 5D complexities and data life cycle phases.
method Systemic and data-centric view of data science as a fusion of data universe and its challenges, formalizing a general-purpose architecture.
result Essential data science as a natural ecosystem integrating specific disciplines and high-impact applications.

This paper introduces C-DSL to improve data mining outcomes by considering context.

problem Data collection ambiguities, data imbalance, hidden biases, lack of domain info, and data incompleteness.
method Developed Context-Driven Data Science Lifecycle (C-DSL) to address data quality issues.
result Tangible improvements to data mining outcomes were achieved through C-DSL.

Proposes using probabilistic models for privacy-preserving synthetic data.

problem Designing high-quality synthetic data for privacy preservation.
method Formulate the problem through probabilistic modelling, choosing a model for the data.
result Statistical discoveries can be reliably reproduced from synthetic data.

Paper creates fair synthetic data ensuring equal predictions across sensitive attributes.

problem Ensuring fair predictions across sensitive attributes in synthetic data.
method Equalizing target probability distributions across sensitive attributes in synthetic data generation.
result Synthetic data provides strong fair predictions, equal across all thresholds.

A new method classifies multiple correlated data streams simultaneously.

problem Classifying multiple correlated data streams in practical scenarios.
method Double-Coupling Support Vector Machines (DC-SVM) considers both internal and external correlations.
result The proposed method outperforms traditional methods on artificial and real-world data streams.

Efficient synthetic data generation improves model performance on tabular data.

problem Improving model robustness and performance with scarce or low-quality data.
method Hardness characterization to identify high-value training points, generating synthetic data only from these points.
result Synthetic data generated from hardest points outperforms non-targeted methods on tabular datasets.

For most problems in science and engineering we can obtain data sets that describe the observed system from various perspectives and record the behavior of its individual components. Heterogeneous data sets can be collectively mined by data fusion. Fusion can focus on a specific target relation and exploit directly ass…

2013-07-02abs ↗pdf ↗

DAERNN models censored data using neural networks with data augmentation.

problem Handling censored data in expectile regression.
method Data augmentation based Expectile Regression Neural Networks (ERNNs).
result DAERNN outperforms existing censored ERNNs methods and achieves comparable predictive performance to fully observed data.

Data preprocessing techniques are devoted to correct or alleviate errors in data. Discretization and feature selection are two of the most extended data preprocessing techniques. Although we can find many proposals for static Big Data preprocessing, there is little research devoted to the continuous Big Data problem. A…

2018-10-14abs ↗pdf ↗

This paper quantifies uncertainty in Data Shapley using statistical inference.

problem Uncertainty in data valuation due to dynamic data distribution.
method Established relationship with U-statistics and quantified uncertainty using statistical inference.
result Confidence intervals for Data Shapley estimations are provided.

DCoM uses deep neural networks to detect semantic data types from raw column values.

problem Detecting semantic data types from dirty and unseen data.
method DCoM employs multi-input NLP-based deep neural networks trained on 686,765 data columns.
result DCoM outperforms existing methods significantly on 78 different semantic data types.