The prominent inequality of wealth and income is a huge concern especially in the United States. The likelihood of diminishing poverty is one valid reason to reduce the world's surging level of economic inequality. The principle of universal moral equality ensures sustainable development and improve the economic stabil…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New datasets improve fairness research by revealing UCI Adult's limitations.
The paper examines bias in ML models using the Adult dataset.
Study shows explanation disparities in machine learning models are influenced by data and model properties.
Unified framework interprets SSL models, revealing biases.
Model predicts cannabis use disorder risk for adolescents and young adults.
Securely trains fair models using homomorphic encryption.
Machine learning models are widely adopted in scenarios that directly affect people. The development of software systems based on these models raises societal and legal concerns, as their decisions may lead to the unfair treatment of individuals based on attributes like race or gender. Data preparation is key in any ma…
In this paper, we present a new explainability formalism designed to shed light on how each input variable of a test set impacts the predictions of machine learning models. Hence, we propose a group explainability formalism for trained machine learning decision rules, based on their response to the variability of the i…
Interpretability and fairness are critical in computer vision and machine learning applications, in particular when dealing with human outcomes, e.g. inviting or not inviting for a job interview based on application materials that may include photographs. One promising direction to achieve fairness is by learning data …
Enhanced synthetic dataset improves asset allocation analysis.
By using methods of statistical physics, we focus on the quantitative analysis of the economic income data descending from different databases. To explain our approach, we introduce the necessary theoretical background, the extended Yakovenko et al. (EY) model. This model gives an analytical description of the annual h…
A new method uses LLMs to discover causal pathways that affect fairness in machine learning.
To construct interpretable explanations that are consistent with the original ML model, counterfactual examples---showing how the model's output changes with small perturbations to the input---have been proposed. This paper extends the work in counterfactual explanations by addressing the challenge of feasibility of su…
In the context of the Dragulescu-Yakovenko (2000) model, we show that empirical income distribution with truncated datasets, cannot be properly modeled by the one-parameter exponential distribution. However, a truncated version characterized by an exponential distribution with two parameters gives an accurate fit.
We investigate the shape of the Italian personal income distribution using microdata from the Survey on Household Income and Wealth, made publicly available by the Bank of Italy for the years 1977--2002. We find that the upper tail of the distribution is consistent with a Pareto-power law type distribution, while the r…
Paper examines how income support affects retirement decisions for low-income individuals.
One shirt size cannot fit everybody, while we cannot make a unique shirt that fits perfectly for everyone because of resource limitation. This analogy is true for the policy making. Policy makers cannot establish a single policy to solve all problems for all regions because each region has its own unique issue. In the …
Study cost-effective fairness audits with partial feedback, improving over random exploration.
Missing data imputation can help improve the performance of prediction models in situations where missing data hide useful information. This paper compares methods for imputing missing categorical data for supervised classification tasks. We experiment on two machine learning benchmark datasets with missing categorical…
Given a dataset of careers and incomes, how large a difference of income between any pair of careers would be? Given a dataset of travel time records, how long do we need to spend more when choosing a public transportation mode instead of to travel? In this paper, we propose a framework that is able to infer or…
Image normalization is a critical step in medical imaging. This step is often done on a per-dataset basis, preventing current segmentation algorithms from the full potential of exploiting jointly normalized information across multiple datasets. To solve this problem, we propose an adversarial normalization approach for…
As virtually all aspects of our lives are increasingly impacted by algorithmic decision making systems, it is incumbent upon us as a society to ensure such systems do not become instruments of unfair discrimination on the basis of gender, race, ethnicity, religion, etc. We consider the problem of determining whether th…
The electronic calendar is a valuable resource nowadays for managing our daily life appointments or schedules, also known as events, ranging from professional to highly personal. Researchers have studied various types of calendar events to predict smartphone user behavior for incoming mobile communications. However, th…
Automated method selects eye tracking variables for categorization tasks.
ADHD is being recognized as a diagnosis which persists into adulthood impacting economic, occupational, and educational outcomes. There is an increased need to accurately diagnose and recommend interventions for this population. One consideration is the development and implementation of reliable and valid outcome measu…
The aim of this work is to establish the personal income distribution from the elementary constituents of a free market; products of a representative good and agents forming the economic network. The economy is treated as a self-organized system. Based on the idea that the dynamics of an economy is governed by slow mod…
Herein, we applied statistical physics to study incomes of three (low-, medium- and high-income) society classes instead of the two (low- and medium-income)classes studied so far. In the frame of the threshold nonlinear Langevin dynamics and its threshold Fokker-Planck counterpart, we derived a unified formula for desc…
This paper explores several types of income which have not been explored so far by authors who tackled income and wealth distribution using Statistical Physics. The main types of income we plan to analyze are income before redistribution (or gross income), income of retired people (or pensions), and income of active pe…
Income redistribution is the transfer of income from some individuals to others directly or indirectly by means of social mechanisms, such as taxation, public services and so on. Employing a spatial public goods game, we study the influence of income redistribution on the evolution of cooperation. Two kinds of evolutio…
Accurate facial expression analysis is an essential step in various clinical applications that involve physical and mental health assessments of older adults (e.g. diagnosis of pain or depression). Although remarkable progress has been achieved toward developing robust facial landmark detection methods, state-of-the-ar…
Two new methods improve monotonic constraint enforcement in regression and classification trees.
This dataset contains the annual aggregated income taxes of all the Italian municipalities over the years 2007-2011. Data are clustered over the Italian regions and provinces. The source of the data is the Italian Ministry of Economics and Finance. The administrative variations in Italy over the quinquennium have been …
We use distributionally-robust optimization for machine learning to mitigate the effect of data poisoning attacks. We provide performance guarantees for the trained model on the original data (not including the poison records) by training the model for the worst-case distribution on a neighbourhood around the empirical…
We found a unified formula for description of the household incomes of all society classes, for instance, for the European Union in years 2005-2010. The formula is more general than well known that of Yakovenko et al. because, it satisfactorily describes not only the household incomes of low- and medium-income society …
Analyzes how many people can receive stable income in a pooled annuity fund.
Capital usually leads to income, and income is more accurately and easily measured. Thus we summarize income distributions in USA, Germany, etc.
Developing a scientific understanding of cities in a fast urbanizing world is essential for planning sustainable urban systems. Recently, it was shown that income and wealth creation follow increasing returns, scaling superlinearly with city size. We study scaling of per capita incomes for separate census defined incom…
Study shows income inequality increases with city size, affecting only the wealthiest deciles.
We found a unified formula for description of the household incomes of all society classes, for instance, of those of the European Union in year 2007. This formula is a stationary solution of the threshold Fokker-Planck equation (derived from the threshold nonlinear Langevin one). The formula is more general than the w…
Income and wealth distribution affect stability of a society to a large extent and high inequality affects it negatively. Moreover, in the case of developed countries, recently has been proven that inequality is closely related to all negative phenomena affecting society. So far, Econophysics papers tried to analyse in…
Model shows significant income inequality emerges from equal opportunities in a simple economy.
We investigate the Japanese personal income distribution in the high income range over the 112 years 1887-1998, and that in the middle income range over the 44 years 1955-98. It is observed that the distribution pattern of the lognormal with power law tail is the universal structure. However the indexes specifying the …
Maximal correlation framework improves fairness in machine learning algorithms.
We prove that the refined approach -- our extension of the Yakovenko et al. formalism -- is universal in the sense that it describes well both household incomes in the European Union and the individual incomes in the United States for social classes of any income. This formalism allowed the study of the impact of the r…
We investigate the Brazilian personal income distribution using data from National Household Sample Survey (PNAD), an annual research available by the Brazilian Institute of Geography and Statistics (IBGE). It provides general characteristics of the country's population. Using PNAD data background we also confirm the e…
We analyze three sets of income data: the US Panel Study of Income Dynamics PSID), the British Household Panel Survey (BHPS), and the German Socio-Economic Panel (GSOEP). It is shown that the empirical income distribution is consistent with a two-parameter lognormal function for the low-middle income group (97%-99% of …
Building machine learning models that are fair with respect to an unprivileged group is a topical problem. Modern fairness-aware algorithms often ignore causal effects and enforce fairness through modifications applicable to only a subset of machine learning models. In this work, we propose a new definition of fairness…