Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

13263952 · Jun 202019922001200920182026
48 results for geographic locations

Paper improves geographic location embeddings using Flickr tags and structured data.

problem Lack of integration between Flickr metadata and structured scientific data.
method Learning vector space embeddings of geographic locations.
result Improved predictions of ecological features using the new method.

Study improves motor insurance claim prediction using geographic data.

problem Limited location identifiers in public actuarial datasets.
method Zone-level modeling framework with environmental and orthoimagery data.
result Geographic information improves MTPL claim prediction accuracy.

Study evaluates methods for improving model robustness to various real-world distribution shifts.

problem Improving model robustness to real-world distribution shifts like geographic changes.
method Introduced new datasets and evaluated existing methods on four types of shifts (style, blurriness, location, camera operation).
result Data augmentations and larger models can improve robustness on real-world distribution shifts, contrary to prior claims.

Study uses trajectory embedding to measure place function similarity at fine spatial granularity.

problem Measuring place function similarity at fine spatial granularity.
method Trajectory embedding to reduce dimensions and measure similarity of place functions.
result Embedding similarity can be a metric proxy for place functions at fine spatial granularity.

Study reveals centralization in Bitcoin transactions involving retail users.

problem Centralization and bias in Bitcoin transaction data.
method Heuristic classification of Bitcoin users, weekly activity pattern analysis.
result Most real transactions involve Frequent Receivers, centralizing the ecosystem.

Automated valuation model uses diverse data sources for real estate appraisal.

problem Accurate and efficient automated valuation of real estate properties.
method Web data acquisition and machine learning model combining structural and geographical data.
result The model achieves high prediction accuracy for real estate values.

GWRBoost improves GWR for better spatial relationship quantification.

problem Underfitting in GWR for complex data and lack of explainable quantification.
method Geographically weighted gradient boosting model using localized additive model and gradient boosting optimization.
result Significant improvement in RMSE and AICc compared to classic GWR.

Interpretable ML models predict recidivism as well as non-interpretable methods and are more fair.

problem Improving recidivism prediction models for fairness and interpretability.
method Trained interpretable ML models on two recidivism datasets, compared to existing methods, and analyzed fairness.
result Interpretable ML models can predict recidivism as well as non-interpretable methods and are more fair.

This paper develops synthetic mobility datasets to protect privacy while maintaining realism.

problem Privacy concerns restrict sharing real mobility datasets, leading to lack of reproducibility.
method The paper benchmarks RNNs, GANs, and copulas to generate realistic synthetic trajectories.
result The approach generates trajectories that are statistically and semantically similar to real-world data.

Which song will Smith listen to next? Which restaurant will Alice go to tomorrow? Which product will John click next? These applications have in common the prediction of user trajectories that are in a constant state of flux over a hidden network (e.g. website links, geographic location). What users are doing now may b…

2015-11-03abs ↗pdf ↗

Space2Vec learns multi-scale spatial representations from grid cell insights.

problem Encoding spatial features with varying scales from GIS data.
method Proposes Space2Vec, a multi-scale representation learning model using grid cell insights.
result Space2Vec outperforms baselines in predicting POI types and image classification with geo-locations.

Study predicts climate data at distant locations using machine learning.

problem Predict climate variables at distant locations where comprehensive data collection is not feasible.
method Uses reservoir computing and vector autoregression models for prediction.
result Machine learning improves prediction accuracy for highly correlated data.

Spatially-aware model improves earthquake hazard assessment accuracy.

problem Misrepresentation of seismic effects across diverse landscapes.
method Causal Bayesian network with Gaussian Processes and normalizing flows.
result Achieves up to 35.2% AUC improvement over existing methods.

Proposes D2D-LSTM for predicting mobile social network content diffusion paths.

problem Lack of accurate content popularity prediction considering time and location in mobile social networks.
method D2D-LSTM, a deep neural network combining user social features and files features.
result Significantly improved prediction accuracy (up to 85.858%) and faster convergence (less than 100 steps).

We present an analysis of the credit market of Japan. The analysis is performed by investigating the bipartite network of banks and firms which is obtained by setting a link between a bank and a firm when a credit relationship is present in a given time window. In our investigation we focus on a community detection alg…

2014-07-21abs ↗pdf ↗

A model predicts solar irradiance without local data using satellite and weather forecasts.

problem Forecasting solar irradiance without local measurements for geographically dispersed solar generators.
method Uses satellite data and weather forecasts with a deep neural network trained on a subset of ground data.
result Proposed model performs as well or better than local models across 25 locations and prediction horizons.

We investigate the community structure of the global ownership network of transnational corporations. We find a pronounced organization in communities that cannot be explained by randomness. Despite the global character of this network, communities reflect first of all the geographical location of firms, while the indu…

2013-01-11abs ↗pdf ↗

How are economic activities linked to geographic locations? To answer this question, we use a data-driven approach that builds on the information about location, ownership and economic activities of the world's 3,000 largest firms and their almost one million subsidiaries. From this information we generate a bipartite …

2015-12-09abs ↗pdf ↗

This paper pretends to analyze the importance which the natural advantages and local resources are in the manufacturing industry location, in relation with the "spillovers" effects and industrial policies. To this, we estimate the Rybczynski equation matrix for the various manufacturing industries in Portugal, at regio…

2011-10-25abs ↗pdf ↗

Graph-partitioning-based DCRNN improves traffic forecasting for large highways.

problem Challenges in accurately forecasting traffic on large highway networks.
method Graph-partitioning method to decompose large networks into smaller, independent networks.
result Demonstrated improved traffic forecasting on a large California highway network.

The problem of predicting the location of users on large social networks like Twitter has emerged from real-life applications such as social unrest detection and online marketing. Twitter user geolocation is a difficult and active research topic with a vast literature. Most of the proposed methods follow either a conte…

2017-12-21abs ↗pdf ↗

We propose a latent self-exciting point process model that describes geographically distributed interactions between pairs of entities. In contrast to most existing approaches that assume fully observable interactions, here we consider a scenario where certain interaction events lack information about participants. Ins…

2013-02-12abs ↗pdf ↗

Weather derivatives help farmers hedge against crop yield risks.

problem High basis risks in weather derivatives pricing models.
method Machine learning ensemble technique to determine yield-weather relationships; mean-reverting model with local temperature dependence.
result Average temperature is the most significant weather variable affecting maize yield.

We investigate the tendency for financial instruments to form clusters when there are multiple factors influencing the correlation structure. Specifically, we consider a stock portfolio which contains companies from different industrial sectors, located in several different countries. Both sector membership and geograp…

2015-05-07abs ↗pdf ↗

A model for POI recommendation using relation embedding.

problem Challenges in POI recommendation due to sparse user-POI matrix and varying context.
method Translation-based relation embedding using Knowledge Graph Embedding techniques, combined matrix factorization framework.
result Demonstrates effectiveness of the proposed model on real-world datasets.

STICC clusters geographic objects considering both spatial contiguity and attributes.

problem Discovering repeated geographic patterns with spatial contiguity.
method Spatial Toeplitz Inverse Covariance-Based Clustering (STICC) method.
result STICC significantly outperforms baseline methods in adjusted rand index and macro-F1 score.

Tropical cyclone wind-intensity prediction is a challenging task considering drastic changes climate patterns over the last few decades. In order to develop robust prediction models, one needs to consider different characteristics of cyclones in terms of spatial and temporal characteristics. Transfer learning incorpora…

2017-08-22abs ↗pdf ↗

Post-estimation smoothing improves prediction accuracy with structural indices.

problem Using natural structural indices in machine learning without losing robustness.
method A post-estimation smoothing operator that separates from the original predictor.
result Post-estimation smoothing improves accuracy over original predictors under simple conditions.

Predict and explain service failures in supply-chain networks using data models.

problem Predict and explain service failures in supply-chain networks, particularly last-mile pickup and delivery.
method Used supervised classification with Random Forests and Association Rules on a dataset of 500,000 services.
result Classifier reaches an average sensitivity of 0.7 and specificity of 0.7 for 5 types of failure.