New gossip algorithms improve robustness of rank-based statistics in decentralized systems.
problem Ensuring robustness in decentralized AI and edge intelligence systems, especially in the presence of corrupted or adversarial data.
method Developed asynchronous gossip algorithms for computing rank-based statistics.
result First convergence rate bound for asynchronous gossip-based rank estimation.
There has been an increasing interest in testing the equality of large Pearson's correlation matrices. However, in many applications it is more important to test the equality of large rank-based correlation matrices since they are more robust to outliers and nonlinearity. Unlike the Pearson's case, testing the equality…
This paper tackles ranking-based performance normalization for optimization algorithms.
problem Ranking optimization algorithms across diverse numerical scales disrupts performance comparisons.
method Introduces absolute ranking and a sampling-based computational method to address numerical scale variation.
result Provides a more robust framework for assessing performance across multiple algorithms and problems.
New algorithm achieves almost exact graph matching in almost quadratic time.
problem Graph matching under correlated Erdős-Rényi models.
method Rank-based graph matching using local tree correlation tests.
result Achieves almost exact recovery in almost quadratic time complexity.
New method for learning causal relationships in PNL models.
problem Learning causal relationships from empirical observations in PNL models.
method Rank-based methods to estimate non-linear functions, disentangling from independence tests.
result Consistent method for PNL causal discovery, validated in experiments.
Proposes a new method to improve Bayesian computation accuracy using flexible classification.
problem Bayesian computations accuracy check using rank-based simulation-based calibration has limitations.
method Replaces marginal rank test with a flexible classification approach that learns from data.
result Improves statistical power and provides an interpretable divergence measure of miscalibration.
Develops a new fairness learning approach for multi-task regression models.
problem Fairness in multi-task regression models with biased datasets.
method Uses rank-based non-parametric independence test (Mann Whitney U statistic) and reformulates as non-convex optimization problem.
result Outperforms state-of-the-art methods on fairness metrics.
The paper critiques the ambiguity of rank-based evaluation methods for entity alignment and link prediction.
problem Ambiguity in evaluating entity alignment and link prediction methods using rank-based scores.
method Analysis of multiple evaluation measures and demonstration of their shortcomings.
result Existing scores cannot reliably compare model performance across different datasets.
This paper improves model robustness to underrepresented groups using ranking metrics and reweighting.
problem Underrepresented groups suffer from low accuracy in models trained via ERM.
method Proposes Discounted Cumulative Gain (DCG) and Discounted Rank Upweighting (DRU) methods.
result Models trained with DRU show superior generalization to unseen groups.
QS-BO optimizes functions using only rank-based feedback.
problem Optimizing expensive functions with unreliable or unavailable metric values.
method Quantile-scaling pipeline to convert ranks into Gaussian targets.
result QS-BO consistently achieves lower objective values and is statistically significant.
We optimize rank-based metrics using blackbox differentiation.
problem Challenges in directly optimizing rank-based metrics due to their non-differentiable and non-decomposable nature.
method Efficient, theoretically sound, and general method for differentiating rank-based metrics with mini-batch gradient descent.
result Competitive performance on standard image retrieval datasets and improved performance on object detectors.
An Atlas model is a rank-based system of continuous semimartingales for which the steady-state values of the processes follow a power law, or Pareto distribution. For a power law, the log-log plot of these steady-state values versus rank is a straight line. Zipf's law is a power law for which the slope of this line is …
Study consumption-investment problem in markets with rank-based returns.
problem Consumption-investment problem in markets with rank-based returns.
method Derives an HJB equation with Neumann boundary conditions for the value function and proves a corresponding verification theorem.
result Explicit solutions for unconstrained, open market constraints, and fully invested cases.
Constructs rank-based continuous semimartingales for financial markets.
problem Model financial markets using rank-based diffusions.
method Uses Dirichlet forms and Feller property to construct semimartingales.
result Establishes nonexistence of triple collisions and simplified rank process dynamics.
New ROC tools assess predictive abilities for any linearly ordered outcomes.
problem Fundamental restriction in ROC analysis for non-dichotomous outcomes.
method ROC movies and UROC curves for linearly ordered outcomes.
result CPA equals AUC for binary outcomes and relates to Spearman's coefficient for pairwise distinct outcomes.
E-values enhance conformal prediction methods.
problem Distribution-free uncertainty quantification.
method Reformulation of conformal prediction using e-values.
result E-values offer new theoretical and practical capabilities.
Study shows how market firm capitalization models converge to stochastic PDE solutions.
problem Understanding convergence of rank-based models with common noise to stochastic PDE solutions.
method Analysis of mean field limit, martingale problem, and pathwise entropy solutions.
result Empirical cumulative distribution function converges to solution of a stochastic PDE under certain conditions.
Reward collapse occurs when ranking-based reward models yield uniform rewards for different prompts.
problem Reward collapse in aligning large language models with human preferences.
method Introduced a prompt-aware optimization scheme to derive closed-form expressions for reward distributions.
result Our prompt-aware utility functions significantly alleviate reward collapse during training.
Framework monitors insurance pricing models for drift and recalibration.
problem Maintaining predictive performance of pricing models in evolving insurance portfolios.
method Formalizes deviance loss and Murphy's score, studies Gini score, develops monitoring framework.
result Framework guides decisions on refitting or recalibrating pricing models.
In the seminal work [9], several macroscopic market observables have been introduced, in an attempt to find characteristics capturing the diversity of a financial market. Despite the crucial importance of such observables for investment decisions, a concise mathematical description of their dynamics has been missing. W…
New method tests independence using ROC analysis and bipartite ranking.
problem Testing independence of two random variables with unknown marginals.
method Nonparametric framework based on ROC analysis and bipartite ranking.
result The method detects small departures from independence in high dimensions.
New method improves compatibility of risk stratification models without sacrificing accuracy.
problem Compatibility issues arise when updating clinical machine learning models.
method Proposes rank-based compatibility measure and new loss function.
result Increased compatibility of models by 0.019 with no loss in discriminative performance.
Paper introduces a new method to compute pseudoinverse for ELM with large datasets.
problem Efficient computation of pseudoinverse for ELM with large datasets.
method Rank-based matrix decomposition of the hidden layer matrix.
result Optimal training time and reduced computational complexity for large hidden nodes.
We propose a novel non-parametric adaptive anomaly detection algorithm for high dimensional data based on rank-SVM. Data points are first ranked based on scores derived from nearest neighbor graphs on n-point nominal data. We then train a rank-SVM using this ranked data. A test-point is declared as an anomaly at alpha-…
We consider systems of diffusion processes ("particles") interacting through their ranks (also referred to as "rank-based models" in the mathematical finance literature). We show that, as the number of particles becomes large, the process of fluctuations of the empirical cumulative distribution functions converges to t…
This article describes the R package varrank. It has a flexible implementation of heuristic approaches which perform variable ranking based on mutual information. The package is particularly suitable for exploring multivariate datasets requiring a holistic analysis. The core functionality is a general implementation of…
We investigate the effect of the proportional hazards assumption on prognostic and predictive models of the survival time of patients suffering from amyotrophic lateral sclerosis (ALS). We theoretically compare the underlying model formulations of several variants of survival forests and implementations thereof, includ…
LxCIM metric improves binary classification performance evaluation.
problem Evaluation metrics for binary classification are often not invariant to local class exchange.
method Proposes LxCIM, a rank-based metric invariant to local class exchange.
result LxCIM addresses limitations of existing metrics like AUROC.
Efficiently clusters survival curves without computationally intensive resampling.
problem Identifying clusters of survival curves efficiently and scalably.
method Log-rank test combined with k-means clustering.
result Achieves comparable results to bootstrap-based methods but with improved efficiency.
The problem of active diagnosis arises in several applications such as disease diagnosis, and fault diagnosis in computer networks, where the goal is to rapidly identify the binary states of a set of objects (e.g., faulty or working) by sequentially selecting, and observing, (noisy) responses to binary valued queries. …
We discuss a natural game of competition and solve the corresponding mean field game with \emph{common noise} when agents' rewards are \emph{rank dependent}. We use this solution to provide an approximate Nash equilibrium for the finite player game and obtain the rate of convergence.
Paper discusses extending Gini score for tied rankings and case weights.
problem Extending Gini score for tied rankings and case weights.
method Discuss and adapt Gini score for ties and case weights.
result Gini score can be used for tied rankings and case weights.
Cross-modal hashing has been receiving increasing interests for its low storage cost and fast query speed in multi-modal data retrievals. However, most existing hashing methods are based on hand-crafted or raw level features of objects, which may not be optimally compatible with the coding process. Besides, these hashi…
Modern retrieval systems are often driven by an underlying machine learning model. The goal of such systems is to identify and possibly rank the few most relevant items for a given query or context. Thus, such systems are typically evaluated using a ranking-based performance metric such as the area under the precision-…
A new Bayesian optimization method using Poisson process for better noise robustness.
problem Estimating relative rankings of candidates in noisy environments.
method Poisson process-based ranking surrogate model and tailored acquisition functions.
result PoPBO framework shows lower computation costs and better robustness to noise compared to GP-BO.
Object ranking or "learning to rank" is an important problem in the realm of preference learning. On the basis of training data in the form of a set of rankings of objects represented as feature vectors, the goal is to learn a ranking function that predicts a linear order of any new set of objects. In this paper, we pr…
RCPO uses ranked choice modeling for better LLM alignment.
problem Pairwise preference optimization limits LLM alignment.
method Unified framework combining preference optimization and ranked choice modeling.
result RCPO outperforms competitive baselines in LLM alignment.
Extreme Multi-label classification (XML) is an important yet challenging machine learning task, that assigns to each instance its most relevant candidate labels from an extremely large label collection, where the numbers of labels, features and instances could be thousands or millions. XML is more and more on demand in…
A new algorithm difFOCI improves feature learning from data.
problem Feature selection and learning from data.
method Parametric, differentiable approximation of FOCI method.
result Improves feature learning with better management of spurious correlations.
Efficient method for vertex embedding and community detection.
problem Vertex embedding and community detection.
method Normalized one-hot graph encoder and rank-based cluster size measure.
result Excellent numerical performance of graph encoder ensemble algorithm.
Linear Discriminant Analysis (LDA) is a well-known method for dimensionality reduction and classification. Previous studies have also extended the binary-class case into multi-classes. However, many applications, such as object detection and keyframe extraction cannot provide consistent instance-label pairs, while LDA …
New model predicts stock performance in large equity markets.
problem Predicting stock performance in large equity markets over long time horizons.
method Rank-based volatility stabilized models calibrated to empirical data.
result The model exhibits relative arbitrage and statistically fits empirical features.
Study ranks of elliptic curves via prime averages.
problem Classifying elliptic curves by rank.
method Average Frobenius trace over primes, data science experiments.
result Oscillating pattern in average trace values, correlates with rank.
New metrics reveal oversmoothing in GNNs more accurately than traditional methods.
problem Oversmoothing in graph neural networks reduces model performance.
method Rank-based metrics to measure oversmoothing in GNNs.
result Rank-based metrics consistently capture oversmoothing, while energy-based metrics often fail.
Researchers define quantiles on Riemannian manifolds using optimal transport.
problem Defining quantiles on nonlinear manifolds.
method Measure-transportation-based approach.
result Theoretical and empirical properties of quantile functions on manifolds.
Unified multitask learning framework for mixed-type outcomes.
problem Difficulty in formulating a unified objective for tasks with different outcomes.
method Multitask transformation framework with shared sparsity, using deep neural networks and rank-based optimization.
result Improved prediction and variable selection across continuous, binary, and mixed outcomes.
We analyze different re-ranking algorithms for diversification and show that majority of them are based on maximizing submodular/modular functions from the class of parameterized concave/linear over modular functions. We study the optimality of such algorithms in terms of the `total curvature'. We also show that by adj…
Rank-based Bayesian Optimization improves molecule selection in chemical systems.
problem Optimizing chemical compounds using traditional regression models.
method Introducing Rank-based Bayesian Optimization (RBO) using ranking models.
result RBO outperforms regression-based BO, especially for rough landscapes and activity cliffs.