Constructs rank-based continuous semimartingales for financial markets.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study consumption-investment problem in markets with rank-based returns.
Study shows how market firm capitalization models converge to stochastic PDE solutions.
We consider systems of diffusion processes ("particles") interacting through their ranks (also referred to as "rank-based models" in the mathematical finance literature). We show that, as the number of particles becomes large, the process of fluctuations of the empirical cumulative distribution functions converges to t…
We study the limiting behaviour of the empirical measure of a system of diffusions interacting through their ranks when the number of diffusions tends to infinity. We prove that the limiting dynamics is given by a McKean-Vlasov evolution equation. Moreover, we show that in a wide range of cases the evolution of the cum…
An Atlas model is a rank-based system of continuous semimartingales for which the steady-state values of the processes follow a power law, or Pareto distribution. For a power law, the log-log plot of these steady-state values versus rank is a straight line. Zipf's law is a power law for which the slope of this line is …
Rank-based metrics are some of the most widely used criteria for performance evaluation of computer vision models. Despite years of effort, direct optimization for these metrics remains a challenge due to their non-differentiable and non-decomposable nature. We present an efficient, theoretically sound, and general met…
New gossip algorithms improve robustness of rank-based statistics in decentralized systems.
QS-BO optimizes functions using only rank-based feedback.
Reward collapse occurs when ranking-based reward models yield uniform rewards for different prompts.
This paper tackles ranking-based performance normalization for optimization algorithms.
There has been an increasing interest in testing the equality of large Pearson's correlation matrices. However, in many applications it is more important to test the equality of large rank-based correlation matrices since they are more robust to outliers and nonlinearity. Unlike the Pearson's case, testing the equality…
In the seminal work [9], several macroscopic market observables have been introduced, in an attempt to find characteristics capturing the diversity of a financial market. Despite the crucial importance of such observables for investment decisions, a concise mathematical description of their dynamics has been missing. W…
Develops a new fairness learning approach for multi-task regression models.
New method improves compatibility of risk stratification models without sacrificing accuracy.
Paper introduces a new method to compute pseudoinverse for ELM with large datasets.
New method for learning causal relationships in PNL models.
New algorithm achieves almost exact graph matching in almost quadratic time.
This article describes the R package varrank. It has a flexible implementation of heuristic approaches which perform variable ranking based on mutual information. The package is particularly suitable for exploring multivariate datasets requiring a holistic analysis. The core functionality is a general implementation of…
This paper improves model robustness to underrepresented groups using ranking metrics and reweighting.
LxCIM metric improves binary classification performance evaluation.
The problem of active diagnosis arises in several applications such as disease diagnosis, and fault diagnosis in computer networks, where the goal is to rapidly identify the binary states of a set of objects (e.g., faulty or working) by sequentially selecting, and observing, (noisy) responses to binary valued queries. …
We discuss a natural game of competition and solve the corresponding mean field game with \emph{common noise} when agents' rewards are \emph{rank dependent}. We use this solution to provide an approximate Nash equilibrium for the finite player game and obtain the rate of convergence.
E-values enhance conformal prediction methods.
In this work, we take a closer look at the evaluation of two families of methods for enriching information from knowledge graphs: Link Prediction and Entity Alignment. In the current experimental setting, multiple different scores are employed to assess different aspects of model performance. We analyze the informative…
Proposes a new method to improve Bayesian computation accuracy using flexible classification.
Paper discusses extending Gini score for tied rankings and case weights.
Cross-modal hashing has been receiving increasing interests for its low storage cost and fast query speed in multi-modal data retrievals. However, most existing hashing methods are based on hand-crafted or raw level features of objects, which may not be optimally compatible with the coding process. Besides, these hashi…
Modern retrieval systems are often driven by an underlying machine learning model. The goal of such systems is to identify and possibly rank the few most relevant items for a given query or context. Thus, such systems are typically evaluated using a ranking-based performance metric such as the area under the precision-…
A new methodology has been introduced to clean the correlation matrix of single stocks returns based on a constrained principal component analysis using financial data. Portfolios were introduced, namely "Fundamental Maximum Variance Portfolios", to capture in an optimal way the risks defined by financial criteria ("Bo…
A new Bayesian optimization method using Poisson process for better noise robustness.
Object ranking or "learning to rank" is an important problem in the realm of preference learning. On the basis of training data in the form of a set of rankings of objects represented as feature vectors, the goal is to learn a ranking function that predicts a linear order of any new set of objects. In this paper, we pr…
RCPO uses ranked choice modeling for better LLM alignment.
Extreme Multi-label classification (XML) is an important yet challenging machine learning task, that assigns to each instance its most relevant candidate labels from an extremely large label collection, where the numbers of labels, features and instances could be thousands or millions. XML is more and more on demand in…
Efficient method for vertex embedding and community detection.
A new algorithm difFOCI improves feature learning from data.
Throughout science and technology, receiver operating characteristic (ROC) curves and associated area under the curve (AUC) measures constitute powerful tools for assessing the predictive abilities of features, markers and tests in binary classification problems. Despite its immense popularity, ROC analysis has been su…
New model predicts stock performance in large equity markets.
Study ranks of elliptic curves via prime averages.
New metrics reveal oversmoothing in GNNs more accurately than traditional methods.
We analyze different re-ranking algorithms for diversification and show that majority of them are based on maximizing submodular/modular functions from the class of parameterized concave/linear over modular functions. We study the optimality of such algorithms in terms of the `total curvature'. We also show that by adj…
Unified multitask learning framework for mixed-type outcomes.
Rank-based Bayesian Optimization improves molecule selection in chemical systems.
Gini index needs auto-calibration for consistent decision-making.
This paper addresses the problem of emotion recognition from physiological signals. Features are extracted and ranked based on their effect on classification accuracy. Different classifiers are compared. The inter-subject variability and the personalization effect are thoroughly investigated, through trial-based and su…
We propose a semiparametric approach, named nonparanormal skeptic, for estimating high dimensional undirected graphical models. In terms of modeling, we consider the nonparanormal family proposed by Liu et al (2009). In terms of estimation, we exploit nonparametric rank-based correlation coefficient estimators includin…
A new framework for consistent segmentation evaluation reduces operating losses.
We present Zeno, a technique to make distributed machine learning, particularly Stochastic Gradient Descent (SGD), tolerant to an arbitrary number of faulty workers. Zeno generalizes previous results that assumed a majority of non-faulty nodes; we need assume only one non-faulty worker. Our key idea is to suspect worke…