Study proposes an active subsampling method for estimating individualized thresholds in high-dimensional data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Clustering explores meaningful patterns in the non-labeled data sets. Cluster Ensemble Selection (CES) is a new approach, which can combine individual clustering results for increasing the performance of the final results. Although CES can achieve better final results in comparison with individual clustering algorithms…
Based on the daily data of American and Chinese stock markets, the dynamic behavior of a financial network with static and dynamic thresholds is investigated. Compared with the static threshold, the dynamic threshold suppresses the large fluctuation induced by the cross-correlation of individual stock prices, and leads…
I show the equivalence between a model of financial contagion and the threshold model of global cascades proposed by Watts (2002). The model financial network comprises banks that hold risky external assets as well as interbank assets. It is shown that a simple threshold model can replicate the size and the frequency o…
Matrix estimation improves individual fairness without sacrificing performance.
Adaptive algorithm for outlier detection by balancing arm exploration and threshold estimation.
CV inference can be invalid for relatively unstable model comparisons.
Locally adaptive clustering for tree delineation.
This paper considers the problem of estimating multiple related Gaussian graphical models from a -dimensional dataset consisting of different classes. Our work is based upon the formulation of this problem as group graphical lasso. This paper proposes a novel hybrid covariance thresholding algorithm that can effecti…
Study finds AUC is most consistent across different prevalence in binary classification.
We solve an optimal consumption problem with habit formation constraints.
In this paper, we investigate a multivariate multi-response (MVMR) linear regression problem, which contains multiple linear regression models with differently distributed design matrices, and different regression and output vectors. The goal is to recover the support union of all regression vectors using -reg…
New diagnostics detect variability in individual risk estimates from machine learning models in healthcare.
Binary classification rules based on covariates typically depend on simple loss functions such as zero-one misclassification. Some cases may require more complex loss functions. For example, individual-level monitoring of HIV-infected individuals on antiretroviral therapy (ART) requires periodic assessment of treatment…
We investigate the probability distributions of the recurrence intervals between consecutive 1-min returns above a positive threshold or below a negative threshold of two indices and 20 individual stocks in China's stock market. The distributions of recurrence intervals for positive and negative thresho…
There has been much discussion recently about how fairness should be measured or enforced in classification. Individual Fairness [Dwork, Hardt, Pitassi, Reingold, Zemel, 2012], which requires that similar individuals be treated similarly, is a highly appealing definition as it gives strong guarantees on treatment of in…
The causal effect of a treatment can vary from person to person based on their individual characteristics and predispositions. Mining for patterns of individual-level effect differences, a problem known as heterogeneous treatment effect estimation, has many important applications, from precision medicine to recommender…
Azure (the cloud service provided by Microsoft) is composed of physical computing units which are called nodes. These nodes are controlled by a software component called Fabric Controller (FC), which can consider the nodes to be in one of many different states such as Ready, Unhealthy, Booting, etc. Some of these state…
Paper finds exact recovery threshold in general hypergraph model.
Complex performance measures, beyond the popular measure of accuracy, are increasingly being used in the context of binary classification. These complex performance measures are typically not even decomposable, that is, the loss evaluated on a batch of samples cannot typically be expressed as a sum or average of losses…
We consider the effects of the global financial crisis through a local Korean financial market around the 2008 crisis. We analyze 185 individual stock prices belonging to the KOSPI (Korea Composite Stock Price Index), cosidering three time periods: the time before, during, and after the crisis. The complex networks gen…
Neurally Augmented ALISTA improves sparse reconstruction performance.
We investigate the trading behavior of Finnish individual investors trading the stocks selected to compute the OMXH25 index in 2003 by tracking the individual daily investment decisions. We verify that the set of investors is a highly heterogeneous system under many aspects. We introduce a correlation based method that…
Study community detection in multi-view data with various types of information.
The presence of noisy instances in mobile phone data is a fundamental issue for classifying user phone call behavior (i.e., accept, reject, missed and outgoing), with many potential negative consequences. The classification accuracy may decrease and the complexity of the classifiers may increase due to the number of re…
New technique reduces gender discrimination in credit lending models.
New method estimates optimal dose intervals for personalized treatment.
This paper presents a simple agent-based model of an economic system, populated by agents playing different games according to their different view about social cohesion and tax payment. After a first set of simulations, correctly replicating results of existing literature, a wider analysis is presented in order to stu…
Model compares altruism and individualism in wealth dynamics.
Investigates timing and asset allocation for life insurance in uncertain financial planning.
In open set recognition (OSR), almost all existing methods are designed specially for recognizing individual instances, even these instances are collectively coming in batch. Recognizers in decision either reject or categorize them to some known class using empirically-set threshold. Thus the decision threshold plays a…
Subspace clustering refers to the problem of clustering high-dimensional data points into a union of low-dimensional linear subspaces, where the number of subspaces, their dimensions and orientations are all unknown. In this paper, we propose a variation of the recently introduced thresholding-based subspace clustering…
Deep ROC analysis improves model selection and interpretation in medical and AI applications.
The Wisdom of Crowds (WOC), as a theory in the social science, gets a new paradigm in computer science. The WOC theory explains that the aggregate decision made by a group is often better than those of its individual members if specific conditions are satisfied. This paper presents a novel framework for unsupervised an…
White matter hyperintensity (WMH) is commonly found in elder individuals and appears to be associated with brain diseases. U-net is a convolutional network that has been widely used for biomedical image segmentation. Recently, U-net has been successfully applied to WMH segmentation. Random initialization is usally used…
Securely analyzes survival data across multiple institutions without revealing individual patient records.
The paper tackles fairness in scoring functions for binary classification.
Introduces Rashomon Capacity to measure predictive multiplicity in probabilistic classifiers.
Framework for controlling multiple risks in AI models.
Paper presents a probabilistic model to improve LLM cascade performance.
Study phase transitions in identifying infected individuals using group testing.
New method improves reliability of selecting individuals based on predicted treatment effects.
New method provides calibrated feature importance explanations for regression models.
Study optimizes interbank lending and borrowing to reduce systemic risk.
Improved reliability of machine learning predictions using variational auto-encoders.
The receiver operating characteristic (ROC) curve is a very useful tool for analyzing the diagnostic/classification power of instruments/classification schemes as long as a binary-scale gold standard is available. When the gold standard is continuous and there is no confirmative threshold, ROC curve becomes less useful…
Study connects database alignment and planted matching using Gaussian features.
The paper tackles fair classification with multiple sensitive features.