Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

0111 · Aug 200419922001200920172026
14 results for microdata

Microdata improves inflation forecasts after major shocks, study finds.

problem Forecasting inflation in a non-stationary environment with microeconomic data.
method Developed a scan test to detect periods of micro forecast outperformance, combined with adaptive machine learning.
result Micro forecasts improve inflation predictions after major shocks, especially after 2020.

New method calibrates ABMs using graph neural networks for microdata.

problem Calibrating ABMs to granular microdata with high-dimensional learning tasks.
method Temporal graph neural networks for learning parameter posteriors.
result Graph neural networks offer inductive biases for Bayesian inference with ABM microstates.

Framework generates precise synthetic populations for scalable modeling.

problem Generating accurate synthetic populations without personal data.
method Constraint-programming framework encoding aggregated statistics and structural relations.
result Exact control of demographic profiles without requiring microdata.

We investigate the shape of the Italian personal income distribution using microdata from the Survey on Household Income and Wealth, made publicly available by the Bank of Italy for the years 1977--2002. We find that the upper tail of the distribution is consistent with a Pareto-power law type distribution, while the r…

2004-08-03abs ↗pdf ↗

Study analyzes household capital risk and poverty trapping, deriving a new function for capital deficit distribution.

problem Analyzing the risk of household capital falling into poverty.
method Introduced a new Gerber-Shiu function to model trapping time and capital deficit distribution.
result Derived a model for capital deficit distribution at trapping using GB distributions.

FinDiff generates synthetic financial data for regulatory tasks.

problem Sharing microdata for research due to privacy regulations.
method Diffusion model using embedding encodings for mixed modality financial data.
result FinDiff excels in generating high-fidelity, privacy-preserving synthetic financial data.

Preserving the privacy of individuals by protecting their sensitive attributes is an important consideration during microdata release. However, it is equally important to preserve the quality or utility of the data for at least some targeted workloads. We propose a novel framework for privacy preservation based on the …

2017-11-05abs ↗pdf ↗

New method reconstructs data subsets from limited published statistics.

problem Reconstructing tabular data from aggregate statistics when full datasets are not possible.
method Generates and verifies subsets of rows and columns that are guaranteed to be correct.
result Privacy violations can persist even with sparse published statistics.

Synthetic tabular data synthesis models balance utility and risk.

problem Generating synthetic tabular data for regulated domains.
method Latent flow models with various learning targets, paths, and sampling methods.
result Velocity and posterior matching objectives yield higher utility, while score and noise matching achieve lower risk.

Study models parking duration using machine learning and interpretable methods.

problem Parking issues in developing countries like India.
method Artificial neural networks (ANNs) for capturing relationships; Garson algorithm and LIME for model interpretation.
result LIME shows higher prediction accuracy and can be universally adopted.