An Atlas model is a rank-based system of continuous semimartingales for which the steady-state values of the processes follow a power law, or Pareto distribution. For a power law, the log-log plot of these steady-state values versus rank is a straight line. Zipf's law is a power law for which the slope of this line is …
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study consumption-investment problem in markets with rank-based returns.
Constructs rank-based continuous semimartingales for financial markets.
Rank-based metrics are some of the most widely used criteria for performance evaluation of computer vision models. Despite years of effort, direct optimization for these metrics remains a challenge due to their non-differentiable and non-decomposable nature. We present an efficient, theoretically sound, and general met…
Study shows how market firm capitalization models converge to stochastic PDE solutions.
Reward collapse occurs when ranking-based reward models yield uniform rewards for different prompts.
New gossip algorithms improve robustness of rank-based statistics in decentralized systems.
In the seminal work [9], several macroscopic market observables have been introduced, in an attempt to find characteristics capturing the diversity of a financial market. Despite the crucial importance of such observables for investment decisions, a concise mathematical description of their dynamics has been missing. W…
New method improves compatibility of risk stratification models without sacrificing accuracy.
Develops a new fairness learning approach for multi-task regression models.
New method for learning causal relationships in PNL models.
This paper tackles ranking-based performance normalization for optimization algorithms.
QS-BO optimizes functions using only rank-based feedback.
There has been an increasing interest in testing the equality of large Pearson's correlation matrices. However, in many applications it is more important to test the equality of large rank-based correlation matrices since they are more robust to outliers and nonlinearity. Unlike the Pearson's case, testing the equality…
We consider systems of diffusion processes ("particles") interacting through their ranks (also referred to as "rank-based models" in the mathematical finance literature). We show that, as the number of particles becomes large, the process of fluctuations of the empirical cumulative distribution functions converges to t…
This article describes the R package varrank. It has a flexible implementation of heuristic approaches which perform variable ranking based on mutual information. The package is particularly suitable for exploring multivariate datasets requiring a holistic analysis. The core functionality is a general implementation of…
New algorithm achieves almost exact graph matching in almost quadratic time.
This paper improves model robustness to underrepresented groups using ranking metrics and reweighting.
LxCIM metric improves binary classification performance evaluation.
Paper discusses extending Gini score for tied rankings and case weights.
In this work, we take a closer look at the evaluation of two families of methods for enriching information from knowledge graphs: Link Prediction and Entity Alignment. In the current experimental setting, multiple different scores are employed to assess different aspects of model performance. We analyze the informative…
Paper introduces a new method to compute pseudoinverse for ELM with large datasets.
RCPO uses ranked choice modeling for better LLM alignment.
New model predicts stock performance in large equity markets.
Modern retrieval systems are often driven by an underlying machine learning model. The goal of such systems is to identify and possibly rank the few most relevant items for a given query or context. Thus, such systems are typically evaluated using a ranking-based performance metric such as the area under the precision-…
The problem of active diagnosis arises in several applications such as disease diagnosis, and fault diagnosis in computer networks, where the goal is to rapidly identify the binary states of a set of objects (e.g., faulty or working) by sequentially selecting, and observing, (noisy) responses to binary valued queries. …
A new Bayesian optimization method using Poisson process for better noise robustness.
A new algorithm difFOCI improves feature learning from data.
We discuss a natural game of competition and solve the corresponding mean field game with \emph{common noise} when agents' rewards are \emph{rank dependent}. We use this solution to provide an approximate Nash equilibrium for the finite player game and obtain the rate of convergence.
Gini index needs auto-calibration for consistent decision-making.
Rank-based Bayesian Optimization improves molecule selection in chemical systems.
E-values enhance conformal prediction methods.
Proposes a new method to improve Bayesian computation accuracy using flexible classification.
Cross-modal hashing has been receiving increasing interests for its low storage cost and fast query speed in multi-modal data retrievals. However, most existing hashing methods are based on hand-crafted or raw level features of objects, which may not be optimally compatible with the coding process. Besides, these hashi…
We propose a semiparametric approach, named nonparanormal skeptic, for estimating high dimensional undirected graphical models. In terms of modeling, we consider the nonparanormal family proposed by Liu et al (2009). In terms of estimation, we exploit nonparametric rank-based correlation coefficient estimators includin…
Throughout science and technology, receiver operating characteristic (ROC) curves and associated area under the curve (AUC) measures constitute powerful tools for assessing the predictive abilities of features, markers and tests in binary classification problems. Despite its immense popularity, ROC analysis has been su…
This paper evaluates various loss functions for Transformer models in stock ranking.
New metrics reveal oversmoothing in GNNs more accurately than traditional methods.
This paper addresses the problem of emotion recognition from physiological signals. Features are extracted and ranked based on their effect on classification accuracy. Different classifiers are compared. The inter-subject variability and the personalization effect are thoroughly investigated, through trial-based and su…
Object ranking or "learning to rank" is an important problem in the realm of preference learning. On the basis of training data in the form of a set of rankings of objects represented as feature vectors, the goal is to learn a ranking function that predicts a linear order of any new set of objects. In this paper, we pr…
Dual-ISL improves implicit generative model training with convex optimization and explicit density approximation.
Extreme Multi-label classification (XML) is an important yet challenging machine learning task, that assigns to each instance its most relevant candidate labels from an extremely large label collection, where the numbers of labels, features and instances could be thousands or millions. XML is more and more on demand in…
Efficient method for vertex embedding and community detection.
We study a mean-field version of rank-based models of equity markets such as the Atlas model introduced by Fernholz in the framework of Stochastic Portfolio Theory. We obtain an asymptotic description of the market when the number of companies grows to infinity. Then, we discuss the long-term capital distribution. We r…
We propose a bootstrap-based robust high-confidence level upper bound (Robust H-CLUB) for assessing the risks of large portfolios. The proposed approach exploits rank-based and quantile-based estimators, and can be viewed as a robust extension of the H-CLUB method (Fan et al., 2015). Such an extension allows us to hand…
Study ranks of elliptic curves via prime averages.
We analyze different re-ranking algorithms for diversification and show that majority of them are based on maximizing submodular/modular functions from the class of parameterized concave/linear over modular functions. We study the optimality of such algorithms in terms of the `total curvature'. We also show that by adj…
Unified multitask learning framework for mixed-type outcomes.