Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

1.1%2.2%3.3%4.4% · Jan 202019922001200920182026
48 results for Social Sciences

The study reveals nations drive scientific research for social and economic interests.

problem Why do nations produce scientific research?
method Synthesizes previous concepts of science and scientific research, defines them, and identifies key drivers.
result Scientific research is driven by nations' social and economic interests, not just for philosophical inquiries.

Complexity science integrates natural and social sciences to understand socio-economic dynamics.

problem Chronic poverty and lack of understanding of societal development.
method Use of tools from natural sciences to analyze socio-economic factors.
result Integration of natural and social sciences reveals synergistic mechanisms driving societal evolution.

SINN combines social science and deep learning for predicting opinion dynamics.

problem Predicting opinion dynamics in social networks using traditional models requires extensive calibration with real data.
method SINN integrates theoretical models and social media data using physics-informed neural networks (PINNs) and matrix factorization.
result SINN outperforms six baseline methods in predicting opinion dynamics on real-world and synthetic datasets.

We discuss several multi-agent models that have their origin in the kinetic exchange theory of statistical mechanics and have been recently applied to a variety of problems in the social sciences. This class of models can be easily adapted for simulations in areas other than physics, such as the modeling of income and …

2013-05-03abs ↗pdf ↗

This paper analyzes social influence using causal data science.

problem Separating genuine causal processes from spurious correlations in social influence data.
method The approach involves partitioning data into groups with minimal contradiction, followed by constrained MLE for causal topology learning.
result The method can retrieve genuine causal arcs and improve influence spread prediction.

Computer science scans LLMs to understand and manipulate their economic forecasts.

problem Understanding and controlling the reasoning of large language models in economics.
method Brain scanning techniques applied to LLMs to identify and manipulate underlying concepts.
result LLMs can be steered to generate forecasts with specific biases, allowing for correction or simulation.

Study uses trillion internet observations to analyze social science insights.

problem Understanding social science insights from internet data.
method Unified dataset of over 1.5 trillion observations, applied to urban growth, sleep duration, and economic productivity.
result Internet growth reaches saturation at 1 IP per 3 people, taking 16.1 years.

Social media enhances or diminishes scientific status, depending on usage.

problem Impact of social media on scientific stratification and mobility.
method Logistic Attribution Analysis combining statistical and machine learning methods.
result Social media promotes stratification and mobility, but beyond a threshold, it negatively impacts status.

New method uses imperfect LLM annotations for valid statistical inference in social science.

problem Inaccurate large language model annotations in social science research.
method Design-based supervised learning (DSL) combining imperfect LLM surrogates with gold-standard labels.
result DSL provides valid statistical inference with comparable predictive accuracy to existing methods.

Develops c-GNF for personalized social science policy analysis.

problem Challenges in estimating causal effects and counterfactual inference in social sciences.
method causal-Graphical Normalizing Flow (c-GNF) method.
result c-GNF performs well in estimating causal effects and counterfactual inference.

The study explores how machine learning can enhance scientific research.

problem Improving scientific models with machine learning.
method Analysis of data-driven models versus manually added variables in regression.
result Complex models may not always improve over simpler ones in scientific contexts.

Fine-tuned open-source LLMs match or exceed closed-source models in social science research.

problem Limited scalability and high costs of large LLMs in social science research.
method Fine-tuning open-source models for specific tasks, exploring training set size effects, proposing hybrid workflow.
result Small, fine-tuned open-source LLMs achieve equal or superior performance to commercial alternatives.

This thesis explores how opportunity and information flow through weak ties in the labor market.

problem The impact of unemployment on individual and national well-being.
method Leveraging computational social science, network science, and data-driven theories to measure opportunity/information flow.
result Opportunity/information flow through weak ties is a key determinant of unemployment length.

New estimator improves statistical validity of synthetic data integration.

problem Combining synthetic data generated by large language models with real data for valid inference.
method Generalized method of moments estimator with theoretical guarantees.
result Improves estimates of target parameter through interactions between synthetic and real data.

StepMix estimates mixture models with covariates for social science applications.

problem Estimating latent classes with covariates in social science models.
method Pseudo-likelihood estimation using one-, two-, and three-step approaches.
result Unified framework for expectation-maximization subroutines.

Study examines financial services, economic growth, and well-being using four prongs.

problem Understanding the components of well-being and their relation to economic growth.
method Four-pronged approach: Uncertainty Principle, Fiscal Responsibilities, Smaller Organizations, Redirecting Growth.
result Holistic understanding of well-being beyond economic growth indicators.

Fisher et al. extend multi-VAR for better modeling of heterogeneous time series.

problem Modeling structurally heterogeneous processes in social, health, and behavioral sciences.
method Adaptive weighting schemes for penalized estimation of multiple-subject multivariate time series.
result Improved estimation performance compared to alternative estimators.

Study compares methods for improving document retrieval accuracy.

problem Improving document retrieval accuracy from large corpora.
method Comparison of query expansion, topic models, and active learning.
result Active learning outperforms keyword lists in most settings.

Quantum mechanics models human perception and decision-making, offering a new approach to understanding social dynamics.

problem Understanding the complex interactions between individuals and groups in social networks.
method Developed a simple computational code based on quantum mechanics principles to model human perception and decision-making.
result Quantum-inspired models can help explain differences in individual and group behavior.

In this paper we focus on the beneficial role of random strategies in social sciences by means of simple mathematical and computational models. We briefly review recent results obtained by two of us in previous contributions for the case of the Peter principle and the efficiency of a Parliament. Then, we develop a new …

2012-09-26abs ↗pdf ↗

Data science redefines causal inference from observational data, classifying tasks into description, prediction, and counterfactual prediction.

problem Widespread misunderstandings about data science's role in causal inference from observational data.
method Organizing data science tasks into three classes: Description, prediction, and counterfactual prediction (including causal inference).
result The necessity of subject-matter expert knowledge for causal analyses in data science.

New tool detects weak and strong Islamophobic hate speech on social media.

problem Detecting Islamophobic hate speech on social media is challenging due to its varied nature.
method Built a multi-class classifier distinguishing between non-Islamophobic, weak Islamophobic, and strong Islamophobic content using GloVe word embeddings.
result Accuracy of 77.6% and balanced accuracy of 83% on a dataset of 109,488 tweets.

New research shows graph embeddings fail to capture key network properties.

problem Graph embeddings fail to capture salient properties of complex networks.
method Mathematical proof and empirical study of various embedding techniques.
result Any successful graph embedding must have a rank nearly linear in the number of vertices.

Optimizes control interventions in real-world networks using deep-learning and network science.

problem Optimizing control over socioeconomic networks subject to constraints.
method Integrates optimization tools from deep-learning with network science.
result Characterizes vulnerability of corporate networks to takeovers.

New method improves active statistical inference by reducing noise.

problem Inaccurate uncertainty estimates in active sampling lead to noisy results.
method Robust sampling strategies that interpolate between uniform and active sampling based on uncertainty scores.
result The robust sampling ensures that the estimator is never worse than uniform sampling and usually outperforms active inference.

The paper uses a graph autoencoder to learn unbiased plant-pollinator interaction embeddings.

problem Sampling bias in citizen science data affects ecological network analysis.
method Bipartite graph variational autoencoder with HSIC for fairness.
result The method mitigates sampling bias and provides unbiased embeddings.

New findings show AI models can't be validated in complex social systems.

problem AI models in complex social systems can't be validated due to data collection practices.
method Formal impossibility results using the MovieLens benchmark.
result AI models in complex social systems are invalid under current data collection practices.

This work studies fairness in systems of multiple algorithms, addressing pitfalls and constructing fair compositions.

problem Fairness of scoring and classification algorithms in systems of multiple algorithms.
method Identifying and addressing pitfalls of naive composition, constructing fair compositions for individual and group fairness.
result Fairness properties of systems of multiple fair algorithms are not necessarily preserved under composition.

Paper presents a machine learning method to improve significance tests for misspecified linear models.

problem Misspecification of linear assumptions in social science models leads to inaccurate significance levels.
method Apply machine learning to fit ground truth function, calculate linear approximation, and adjust the estimator.
result The method significantly outperforms linear regression for non-linear ground truth functions.

P.W. Anderson proposed the concept of complexity in order to describe the emergence and growth of macroscopic collective patterns out of the simple interactions of many microscopic agents. In the physical sciences this paradigm was implemented systematically and confirmed repeatedly by successful confrontation with rea…

2008-03-14abs ↗pdf ↗