Extreme classification seeks to assign each data point, the most relevant labels from a universe of a million or more labels. This task is faced with the dual challenge of high precision and scalability, with millisecond level prediction times being a benchmark. We propose DEFRAG, an adaptive feature agglomeration tech…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A new text clustering method using NMF and LSA improves stability and performance.
With this work we try to analyse the agglomeration process in the Portuguese regions, using the New Economic Geography models. In these models the base idea is that where has increasing returns to scale in the manufactured industry and low transport costs, there is agglomeration. Of referring, as summary conclusion, th…
The aim of this paper is to analyze the processes of polarization and agglomeration, to explain the mechanisms and causes of these phenomena in order to identify similarities and differences. As the main implication of this study should be noted that both process pretend to explain the concentration of economic activit…
The aim of this paper is to analyze the relationship between inter-industry, intra-industry and inter-regional clustering and demand for labor by companies in Portugal. Is expected at the outset that there is more demand for work where the agglomeration is greater. It should be noted, as a summary conclusion, the resul…
The paper explores various forms of calibration scores and their implications for fairness.
In this work, we revisit fast dimension reduction approaches, as with random projections and random sampling. Our goal is to summarize the data to decrease computational costs and memory footprint of subsequent analysis. Such dimension reduction can be very efficient when the signals of interest have a strong structure…
This work aims to study the Portuguese regional agglomeration process, using the linear form the New Economic Geography models that emphasize the importance of spatial factors (distance, costs of transport and communication) in explaining of the concentration of economic activity in certain locations. In a theoretical …
Newly available data on the spatial distribution of retail activities in cities makes it possible to build models formalized at the level of the single retailer. Current models tackle consumer location choices at an aggregate level and the opportunity new data offers for modeling at the retail unit level lacks a theore…
Infinite mixture models are commonly used for clustering. One can sample from the posterior of mixture assignments by Monte Carlo methods or find its maximum a posteriori solution by optimization. However, in some problems the posterior is diffuse and it is hard to interpret the sampled partitionings. In this paper, we…
ANN clusters multi-view data by agglomerating subviews and avoiding postprocessing.
Efficiently clusters data with weak assumptions, robust to contamination.
New bounds improve linkage methods for clustering, distinguishing complete-link from single-link.
Inverse inference, or "brain reading", is a recent paradigm for analyzing functional magnetic resonance imaging (fMRI) data, based on pattern recognition and statistical learning. By predicting some cognitive variables related to brain activation maps, this approach aims at decoding brain activity. Inverse inference ta…
Using open source data, we observe the fascinating dynamics of nighttime light. Following a global economic regime shift, the planetary center of light can be seen moving eastwards at a pace of about 60 km per year. Introducing spatial light Gini coefficients, we find a universal pattern of human settlements across dif…
Study a simplified model of multiverse structure with synchronized timelines.
Sum-of-norms clustering is a method for assigning points in to clusters, , using convex optimization. Recently, Panahi et al.\ proved that sum-of-norms clustering is guaranteed to recover a mixture of Gaussians under the restriction that the number of samples is not too large. The pu…
Unified method for simultaneous denoising and clustering.
With the large-scale penetration of the internet, for the first time, humanity has become linked by a single, open, communications platform. Harnessing this fact, we report insights arising from a unified internet activity and location dataset of an unparalleled scope and accuracy drawn from over a trillion (1.5$\times…
Developing a scientific understanding of cities in a fast urbanizing world is essential for planning sustainable urban systems. Recently, it was shown that income and wealth creation follow increasing returns, scaling superlinearly with city size. We study scaling of per capita incomes for separate census defined incom…
With this study we want to test the validity of the well known "Verdoorn's Law" which considers the relationship between the growth of productivity and output in the case of the Portuguese economy at a regional and sectoral levels (NUTs II) for the period 1995-1999. The importance of some additional variables in the or…
It has always been a great challenge for clustering algorithms to automatically determine the cluster numbers according to the distribution of datasets. Several approaches have been proposed to address this issue, including the recent promising work which incorporate Bayesian Nonparametrics into the -means clusterin…
CAF-HFCM automatically forms a cluster hierarchy and optimizes the number of clusters without trial-and-validation.
This study predicts stock prices using various machine and deep learning models.
Clarifies contributions of model architectures and resources in SOTA language models.
Identifies features most relevant to concept drift in data.
New method evaluates feature interactions using orthogonal variance decomposition.
Paper predicts EEG features from acoustic features using RNN and GAN.
Introduces RFI for assessing feature importance relative to any subset of features.
A single pre-trained agent guides feature selection using knockoffs.
Feature networks link ML features via graph structure for enhanced learning.
Approach for selecting features by discarding nuisance and correlated ones.
New stability measures for similar features improve feature selection accuracy.
Counterexamples show HSIC feature selection misses critical features.
This paper shows feature importance remains valid even in low-performing models.
New algorithms select and rank features from MTS without feature extraction.
Pipeline learns topological features for protein stability prediction.
In text classification, dictionaries can be used to define human-comprehensible features. We propose an improvement to dictionary features called smoothed dictionary features. These features recognize document contexts instead of n-grams. We describe a principled methodology to solicit dictionary features from a teache…
FeAT improves OOD generalization by learning richer features.
The paper defines and analyzes feature complexity in DNNs, proposing metrics for feature disentanglement and evaluation.
Conventional mutual information (MI) based feature selection (FS) methods are unable to handle heterogeneous feature subset selection properly because of data format differences or estimation methods of MI between feature subset and class label. A way to solve this problem is feature transformation (FT). In this study,…
A new method for measuring conditional feature importance using generative models.
CAN approximates explicit feature interactions for CTR prediction.
Proposes a new feature selection method integrating feature relationships.
Online feature selection has been an active research area in recent years. We propose a novel diverse online feature selection method based on Determinantal Point Processes (DPP). Our model aims to provide diverse features which can be composed in either a supervised or unsupervised framework. The framework aims to pro…
New method disentangles feature importance scores in machine learning.
Feature selection has been proven a powerful preprocessing step for high-dimensional data analysis. However, most state-of-the-art methods tend to overlook the structural correlation information between pairwise samples, which may encapsulate useful information for refining the performance of feature selection. Moreove…
GRANITE unifies feature-based explanation methods to reduce disagreement.