A new Bayesian optimization method using Poisson process for better noise robustness.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Rank-based Bayesian Optimization improves molecule selection in chemical systems.
QS-BO optimizes functions using only rank-based feedback.
We optimize rank-based metrics using blackbox differentiation.
Several tasks in machine learning are evaluated using non-differentiable metrics such as mean average precision or Spearman correlation. However, their non-differentiability prevents from using them as objective functions in a learning framework. Surrogate and relaxation methods exist but tend to be specific to a given…
An Atlas model is a rank-based system of continuous semimartingales for which the steady-state values of the processes follow a power law, or Pareto distribution. For a power law, the log-log plot of these steady-state values versus rank is a straight line. Zipf's law is a power law for which the slope of this line is …
Study consumption-investment problem in markets with rank-based returns.
Constructs rank-based continuous semimartingales for financial markets.
New gossip algorithms improve robustness of rank-based statistics in decentralized systems.
The paper critiques the ambiguity of rank-based evaluation methods for entity alignment and link prediction.
Study shows how market firm capitalization models converge to stochastic PDE solutions.
Reward collapse occurs when ranking-based reward models yield uniform rewards for different prompts.
This paper tackles ranking-based performance normalization for optimization algorithms.
There has been an increasing interest in testing the equality of large Pearson's correlation matrices. However, in many applications it is more important to test the equality of large rank-based correlation matrices since they are more robust to outliers and nonlinearity. Unlike the Pearson's case, testing the equality…
We consider the predictive problem of supervised ranking, where the task is to rank sets of candidate items returned in response to queries. Although there exist statistical procedures that come with guarantees of consistency in this setting, these procedures require that individuals provide a complete ranking of all i…
In the seminal work [9], several macroscopic market observables have been introduced, in an attempt to find characteristics capturing the diversity of a financial market. Despite the crucial importance of such observables for investment decisions, a concise mathematical description of their dynamics has been missing. W…
Develops a new fairness learning approach for multi-task regression models.
New method improves compatibility of risk stratification models without sacrificing accuracy.
Paper introduces a new method to compute pseudoinverse for ELM with large datasets.
We consider systems of diffusion processes ("particles") interacting through their ranks (also referred to as "rank-based models" in the mathematical finance literature). We show that, as the number of particles becomes large, the process of fluctuations of the empirical cumulative distribution functions converges to t…
New method for learning causal relationships in PNL models.
New algorithm achieves almost exact graph matching in almost quadratic time.
ORAT improves model robustness against outliers and adversarial attacks.
This article describes the R package varrank. It has a flexible implementation of heuristic approaches which perform variable ranking based on mutual information. The package is particularly suitable for exploring multivariate datasets requiring a holistic analysis. The core functionality is a general implementation of…
New ROC tools assess predictive abilities for any linearly ordered outcomes.
This paper improves model robustness to underrepresented groups using ranking metrics and reweighting.
LxCIM metric improves binary classification performance evaluation.
The problem of active diagnosis arises in several applications such as disease diagnosis, and fault diagnosis in computer networks, where the goal is to rapidly identify the binary states of a set of objects (e.g., faulty or working) by sequentially selecting, and observing, (noisy) responses to binary valued queries. …
We discuss a natural game of competition and solve the corresponding mean field game with \emph{common noise} when agents' rewards are \emph{rank dependent}. We use this solution to provide an approximate Nash equilibrium for the finite player game and obtain the rate of convergence.
E-values enhance conformal prediction methods.
Proposes a new method to improve Bayesian computation accuracy using flexible classification.
Paper discusses extending Gini score for tied rankings and case weights.
Cross-modal hashing has been receiving increasing interests for its low storage cost and fast query speed in multi-modal data retrievals. However, most existing hashing methods are based on hand-crafted or raw level features of objects, which may not be optimally compatible with the coding process. Besides, these hashi…
Modern retrieval systems are often driven by an underlying machine learning model. The goal of such systems is to identify and possibly rank the few most relevant items for a given query or context. Thus, such systems are typically evaluated using a ranking-based performance metric such as the area under the precision-…
New bounds show polyhedral surrogates are optimal for generalization.
Object ranking or "learning to rank" is an important problem in the realm of preference learning. On the basis of training data in the form of a set of rankings of objects represented as feature vectors, the goal is to learn a ranking function that predicts a linear order of any new set of objects. In this paper, we pr…
RCPO uses ranked choice modeling for better LLM alignment.
Extreme Multi-label classification (XML) is an important yet challenging machine learning task, that assigns to each instance its most relevant candidate labels from an extremely large label collection, where the numbers of labels, features and instances could be thousands or millions. XML is more and more on demand in…
A new algorithm difFOCI improves feature learning from data.
Efficient method for vertex embedding and community detection.
We present a framework for automatically structuring and training fast, approximate, deep neural surrogates of stochastic simulators. Unlike traditional approaches to surrogate modeling, our surrogates retain the interpretable structure and control flow of the reference simulator. Our surrogates target stochastic simul…
Randomizing the Fourier-transform (FT) phases of temporal-spatial data generates surrogates that approximate examples from the data-generating distribution. We propose such FT surrogates as a novel tool to augment and analyze training of neural networks and explore the approach in the example of sleep-stage classificat…
We formalize and study the natural approach of designing convex surrogate loss functions via embeddings, for problems such as classification, ranking, or structured prediction. In this approach, one embeds each of the finitely many predictions (e.g.\ rankings) as a point in , assigns the original loss val…
The paper proposes a scalable framework for uncertainty quantification and propagation in surrogate-based Bayesian inference.
Paper proposes hybrid modeling to improve surrogate accuracy using multiple data sources.
New model predicts stock performance in large equity markets.
Improving predictive understanding of Earth system variability and change requires data-model integration. Efficient data-model integration for complex models requires surrogate modeling to reduce model evaluation time. However, building a surrogate of a large-scale Earth system model (ESM) with many output variables i…
The paper proposes a method to assess surrogate heterogeneity in non-randomized data.