New Hermite series estimator for Spearman rank correlation in non-stationary data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We study the problem of rank aggregation: given a set of ranked lists, we want to form a consensus ranking. Furthermore, we consider the case of extreme lists: i.e., only the rank of the best or worst elements are known. We impute missing ranks by the average value and generalise Spearman's ρto extreme ranks. Our main …
New unsupervised feature selection method for imbalanced datasets.
Standardizes weighted ranking correlation coefficients to maintain zero expected value.
This paper uses rank correlation methods to construct MSTs from financial returns, finding them more stable and robust.
Given a set of objects, an online ranking system outputs at each time step a full ranking of the set, observes a feedback of some form and suffers a loss. We study the setting in which the (adversarial) feedback is an element in , and the loss is the position (0th, 1st, 2nd...) of the item in the outputted r…
In the last decade, many diverse advances have occurred in the field of information extraction from data. Information extraction in its simplest form takes place in computing environments, where structured data can be extracted through a series of queries. The continuous expansion of quantities of data have therefore p…
We propose a semiparametric approach, named nonparanormal skeptic, for estimating high dimensional undirected graphical models. In terms of modeling, we consider the nonparanormal family proposed by Liu et al (2009). In terms of estimation, we exploit nonparametric rank-based correlation coefficient estimators includin…
Nonparametric correlations such as Spearman's rank correlation and Kendall's tau correlation are widely applied in scientific and engineering fields. This paper investigates the problem of computing nonparametric correlations on the fly for streaming data. Standard batch algorithms are generally too slow to handle real…
STRAPSim measures ETF portfolio similarity better than existing methods.
In the IEEE Investment ranking challenge 2018, participants were asked to build a model which would identify the best performing stocks based on their returns over a forward six months window. Anonymized financial predictors and semi-annual returns were provided for a group of anonymized stocks from 1996 to 2017, which…
Several tasks in machine learning are evaluated using non-differentiable metrics such as mean average precision or Spearman correlation. However, their non-differentiability prevents from using them as objective functions in a learning framework. Surrogate and relaxation methods exist but tend to be specific to a given…
We introduce a new family of minmax rank aggregation problems under two distance measures, the Kendall τ and the Spearman footrule. As the problems are NP-hard, we proceed to describe a number of constant-approximation algorithms for solving them. We conclude with illustrative applications of the aggregation methods on…
Paper introduces differentiable sorting and ranking with time complexity.
The Chirikov standard map and the 2D Froeschlé map are investigated. A few thousand values of the Hurst exponent (HE) and the maximal Lyapunov exponent (mLE) are plotted in a mixed space of the nonlinear parameter versus the initial condition. Both characteristic exponents reveal remarkably similar structures in this s…
Study examines stock price correlations between Indonesian holding companies and their subsidiaries.
The paper introduces a framework to select efficient datasets for preserving model rankings.
A new sparse benchmark metabench identifies key abilities from large benchmarks.
ChatGPT predicts stock market movements based on Bloomberg headlines, showing a positive correlation over short to medium terms.
A simple method reduces bias in LLM auto-evaluators by controlling output length.
Paper proposes Coalitional BAE to improve explainability of unsupervised deep learning models.
This paper proves a generalization bound for complex-valued neural networks scaling with spectral complexity.
This paper fills in local bounds for Spearman's footrule and Gini's gamma measures of association.
Novel fusion of autoencoders predicts sleepiness from speech.
In this paper, we propose a semiparametric approach, named nonparanormal skeptic, for efficiently and robustly estimating high dimensional undirected graphical models. To achieve modeling flexibility, we consider Gaussian Copula graphical models (or the nonparanormal) as proposed by Liu et al. (2009). To achieve estima…
Eliciting semantic similarity between concepts in the biomedical domain remains a challenging task. Recent approaches founded on embedding vectors have gained in popularity as they risen to efficiently capture semantic relationships The underlying idea is that two words that have close meaning gather similar contexts. …
Developers of text-to-speech synthesizers (TTS) often make use of human raters to assess the quality of synthesized speech. We demonstrate that we can model human raters' mean opinion scores (MOS) of synthesized speech using a deep recurrent neural network whose inputs consist solely of a raw waveform. Our best models …
Parallel deep learning architectures like fine-tuned BERT and MT-DNN, have quickly become the state of the art, bypassing previous deep and shallow learning methods by a large margin. More recently, pre-trained models from large related datasets have been able to perform well on many downstream tasks by just fine-tunin…
Efficiently predict LLM benchmarks using feature selection and regression.
This paper proposes a new class of copulas which characterize the set of all twice continuously differentiable copulas. We show that our proposed new class of copulas is a new generalized copula family that include not only asymmetric copulas but also all smooth copula families available in the current literature. Spea…
Unified framework detects overfitting in crash classification models.
An investigation is presented of how a comprehensive choice of five most important measures of concordance (namely Spearman's rho, Kendall's tau, Gini's gamma, Blomqvist's beta, and their weaker counterpart Spearman's footrule) relate to non-exchangeability, i.e., asymmetry on copulas. Besides these results, the method…
An automated metric to evaluate dialogue quality is vital for optimizing data driven dialogue management. The common approach of relying on explicit user feedback during a conversation is intrusive and sparse. Current models to estimate user satisfaction use limited feature sets and rely on annotation schemes with low …
Learning knowledge representation is an increasingly important technology applicable in many domain-specific machine learning problems. We discuss the effectiveness of traditional Link Prediction or Knowledge Graph Completion evaluation protocol when embedding knowledge representation for categorised multi-relational d…
We investigate the relative information content of six measures of dependence between two random variables and for large or extreme events for several models of interest for financial time series. The six measures of dependence are respectively the linear correlation and Spearman's rho conditio…
The paper describes correlations of spectra for higher rank Anosov representations.
An investigation is presented of how a comprehensive choice of four most important measures of concordance (namely Spearman's rho, Kendall's tau, Spearman's footrule, and Gini's gamma) relate to the fifth one, i.e., the Blomqvist's beta. In order to work out these results we present a novel method of estimating the val…
Improved DeepONet variants using Transformer cross-conditioning enhance PDE solution efficiency.
A new method for Gaussian Processes handles mixed continuous and categorical inputs.
Proposes a method to enhance multi-view learning by maximizing higher order correlations.
Enhances labels from unlabeled data using sample correlations.
New methods rank players using covariates and comparisons, outperforming existing algorithms.
A self-supervised debiasing method using rank regularization mitigates spurious correlations in neural networks.
LLM forecasting benchmarks suffer from information leakage, which confounds model performance.
Forest tree species mapped with high accuracy using satellite data.
Optimal Word2Vec hyper-parameters improve NLP tasks.
There has been an increasing interest in testing the equality of large Pearson's correlation matrices. However, in many applications it is more important to test the equality of large rank-based correlation matrices since they are more robust to outliers and nonlinearity. Unlike the Pearson's case, testing the equality…
Throughout science and technology, receiver operating characteristic (ROC) curves and associated area under the curve (AUC) measures constitute powerful tools for assessing the predictive abilities of features, markers and tests in binary classification problems. Despite its immense popularity, ROC analysis has been su…