We characterize the class of exchangeable feature allocations assigning probability to a feature allocation of individuals, displaying features with counts for these features. Each element of this class is parametrized by a countable matrix …
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We address the problem of estimating the parameters of a time-homogeneous Markov chain given only noisy, aggregate data. This arises when a population of individuals behave independently according to a Markov chain, but individual sample paths cannot be observed due to limitations of the observation process or the need…
New approach for estimating individual treatment effects in low compliance settings.
Paper proves identifiability and consistency of hub model for network inference.
This work develops a model to distinguish network and covariate information.
This paper presents a dynamic model to study the impact on the economic outcomes in different societies during the Malthusian Era of individualism (time spent working alone) and collectivism (complementary time spent working with others). The model is driven by opposing forces: a greater degree of collectivism provides…
Algorithm learns similarity metrics for individual fairness.
We give a microscopic representation of the stock-market in which the microscopic agents are the individual traders and their capital. Their basic dynamics consists in the auto-catalysis of the individual capital and in the global competition/cooperation between the agents mediated by the total wealth invested in the s…
Study finds AUC is most consistent across different prevalence in binary classification.
We propose nonparametric methods for individual calibration in regression models.
Framework generates precise synthetic populations for scalable modeling.
We find the optimal investment strategy to minimize the expected time that an individual's wealth stays below zero, the so-called {\it occupation time}. The individual consumes at a constant rate and invests in a Black-Scholes financial market consisting of one riskless and one risky asset, with the risky asset's price…
Currently, pension providers are running into trouble mainly due to the ultra-low interest rates and the guarantees associated to some pension benefits. With the aim of reducing the pension volatility and providing adequate pension levels with no guarantees, we carry out mathematical analysis of a new pension design in…
Unified approach to aggregating models and preferences.
Interpreting predictions from tree ensemble methods such as gradient boosting machines and random forests is important, yet feature attribution for trees is often heuristic and not individualized for each prediction. Here we show that popular feature attribution methods are inconsistent, meaning they can lower a featur…
Matrix estimation improves individual fairness without sacrificing performance.
Rank aggregation systems collect ordinal preferences from individuals to produce a global ranking that represents the social preference. Rank-breaking is a common practice to reduce the computational complexity of learning the global ranking. The individual preferences are broken into pairwise comparisons and applied t…
DGSAM improves domain generalization by minimizing individual sharpness.
System recommends workouts and predicts success rates using RNNs.
We consider the predictive problem of supervised ranking, where the task is to rank sets of candidate items returned in response to queries. Although there exist statistical procedures that come with guarantees of consistency in this setting, these procedures require that individuals provide a complete ranking of all i…
The study analyzes how deep neural networks treat instances with regular and irregular patterns.
The authors propose a parametric model called the arena model for prediction in paired competitions, i.e. paired comparisons with eliminations and bifurcations. The arena model has a number of appealing advantages. First, it predicts the results of competitions without rating many individuals. Second, it takes full adv…
Study efficient inference for network quantile causal effects with partial interference.
New method finds balanced clusters in graphs using auxiliary information.
A number of machine learning (ML) methods have been proposed recently to maximize model predictive accuracy while enforcing notions of group parity or fairness across sub-populations. We propose a desirable property for these procedures, slack-consistency: For any individual, the predictions of the model should be mono…
Consistent spectral clustering with fairness constraints on representation graphs.
This note fills the gap in market-consistent valuation of lifelong health insurance products.
Confidentiality of patient information is an essential part of Electronic Health Record System. Patient information, if exposed, can cause a serious damage to the privacy of individuals receiving healthcare. Hence it is important to remove such details from physician notes. A system is proposed which consists of a deep…
Develops methods for near-optimal personalized treatment recommendations.
ClusterSC improves synthetic control by selecting relevant donor groups.
This paper introduces new methods for analysing the extreme and erratic behaviour of time series to evaluate the impact of COVID-19 on cryptocurrency market dynamics. Across 51 cryptocurrencies, we examine extreme behaviour through a study of distribution extremities, and erratic behaviour through structural breaks. Fi…
Paper uses stochastic algorithms to estimate systemic risk measures.
We propose a new algorithm for training generative adversarial networks that jointly learns latent codes for both identities (e.g. individual humans) and observations (e.g. specific photographs). By fixing the identity portion of the latent codes, we can generate diverse images of the same subject, and by fixing the ob…
New method estimates optimal dose intervals for personalized treatment.
We analyze the distribution of income and income tax of individuals in Japan for the fiscal year 1998. From the rank-size plots we find that the accumulated probability distribution of both data obey a power law with a Pareto exponent very close to -2. We also present an analysis of the distribution of the debts owed b…
This paper improves deep learning model consistency through ensemble methods.
Estimating the largest community in a mixed population via sequential sampling.
New model clusters cells and individuals, revealing genetic influences on cell types.
GWIB improves counterfactual regression by balancing latent distributions and reducing selection bias.
In recent years rank aggregation has received significant attention from the machine learning community. The goal of such a problem is to combine the (partially revealed) preferences over objects of a large population into a single, relatively consistent ordering of those objects. However, in many cases, we might not w…
Bayesian approach clusters survival data for better risk prediction.
CausalBGM uses AI to infer causal effects from complex data.
Research in several fields now requires the analysis of data sets in which multiple high-dimensional types of data are available for a common set of objects. In particular, The Cancer Genome Atlas (TCGA) includes data from several diverse genomic technologies on the same cancerous tumor samples. In this paper we introd…
We consider the problem of learning predictive models from longitudinal data, consisting of irregularly repeated, sparse observations from a set of individuals over time. Such data often exhibit {\em longitudinal correlation} (LC) (correlations among observations for each individual over time), {\em cluster correlation…
SD-SCMs generate counterfactual data for causal inference benchmarks.
Person re-identification (re-id), an emerging problem in visual surveillance, deals with maintaining entities of individuals whilst they traverse various locations surveilled by a camera network. From a visual perspective re-id is challenging due to significant changes in visual appearance of individuals in cameras wit…
Proposes a new estimator for weak instrumental variables in panel data models.
With ever-increasing available data, predicting individuals' preferences and helping them locate the most relevant information has become a pressing need. Understanding and predicting preferences is also important from a fundamental point of view, as part of what has been called a "new" computational social science. He…