Progress of machine learning in critical care has been difficult to track, in part due to absence of public benchmarks. Other fields of research (such as computer vision and natural language processing) have established various competitions and public benchmarks. Recent availability of large clinical datasets has enabl…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper proposes a new daily benchmark for post-GFC government bond CIP deviations.
RED-2400 is a public benchmark of trading events from a Solana exchange, labeled by algorithmic rejection.
Widely-used public benchmarks are of huge importance to computer vision and machine learning research, especially with the computational resources required to reproduce state of the art results quickly becoming untenable. In medical image computing, the wide variety of image modalities and problem formulations yields a…
New private learning algorithms improve utility in tasks with public features.
Benchmarking recursive collapse claims with a new framework under false-positive control.
Along with the advance of opinion mining techniques, public mood has been found to be a key element for stock market prediction. However, how market participants' behavior is affected by public mood has been rarely discussed. Consequently, there has been little progress in leveraging public mood for the asset allocatio…
Paper addresses challenges in benchmarking stream learning algorithms with real-world data.
Publicly pretraining models on Web data may undermine differential privacy.
Public transport is one of the major forms of transportation in the world. This makes it vital to ensure that public transport is efficient. This research presents a novel real-time GPS bus transit data for over 500 routes of buses operating in New Delhi. The data can be used for modeling various timetable optimization…
KT models improved slightly with synthetic student data.
In recent years, an active field of research has developed around automated machine learning (AutoML). Unfortunately, comparing different AutoML systems is hard and often done incorrectly. We introduce an open, ongoing, and extensible benchmark framework which follows best practices and avoids common mistakes. The fram…
Predicts language model performance from public models without training.
ISOMORPH creates a digital twin for supply chain logistics, advancing time-series forecasting benchmarks.
We introduce a transformer-based GNN model, named UGformer, to learn graph representations. In particular, we present two UGformer variants, wherein the first variant (publicized in September 2019) is to leverage the transformer on a set of sampled neighbors for each input node, while the second (publicized in May 2021…
Introduces AMLB, an open benchmark for AutoML frameworks.
Deployment-complete benchmarking assesses if evidence leads to consistent deployment actions.
FLBench automates federated learning benchmarking.
Efficiently identifies promising hyperparameters for online learning models.
Proposes CLRS benchmark to evaluate algorithmic reasoning.
We introduce Graph-Sparse Logistic Regression, a new algorithm for classification for the case in which the support should be sparse but connected on a graph. We val- idate this algorithm against synthetic data and benchmark it against L1-regularized Logistic Regression. We then explore our technique in the bioinformat…
A benchmark for simulation-based inference methods.
Since the introduction and the public availability of the \textsc{ucr} time series benchmark data sets, numerous Time Series Classification (TSC) methods has been designed, evaluated and compared to each others. We suggest a critical view of TSC performance evaluation protocols put in place in recent TSC literature. Th…
Pythae is a Python library for benchmarking VAE models.
Molecular machine learning has been maturing rapidly over the last few years. Improved methods and the presence of larger datasets have enabled machine learning algorithms to make increasingly accurate predictions about molecular properties. However, algorithmic progress has been limited due to the lack of a standard b…
BTZSC benchmarks zero-shot text classification across diverse models.
Quantization based techniques are the current state-of-the-art for scaling maximum inner product search to massive databases. Traditional approaches to quantization aim to minimize the reconstruction error of the database points. Based on the observation that for a given query, the database points that have the largest…
New simulation shows trading algorithms' performance varies with parallelism.
Risk, including economic risk, is increasingly a concern for public policy and management. The possibility of dealing effectively with risk is hampered, however, by lack of a sound empirical basis for risk assessment and management. The paper demonstrates the general point for cost and demand risks in urban rail projec…
Proposes CLRS-Text, a new benchmark for evaluating LM reasoning capabilities.
Implementing large-scale information and communication technology (IT) projects carries large risks and easily might disrupt operations, waste taxpayers' money, and create negative publicity. Because of the high risks it is important that government leaders manage the attendant risks. We analysed a sample of 1,355 publ…
We created financial benchmarks for distribution shifts in crude oil prices and volatility.
Deep Learning approaches for solving Inverse Problems in imaging have become very effective and are demonstrated to be quite competitive in the field. Comparing these approaches is a challenging task since they highly rely on the data and the setup that is used for training. We provide a public dataset of computed tomo…
Because the choice and tuning of the optimizer affects the speed, and ultimately the performance of deep learning, there is significant past and recent research in this area. Yet, perhaps surprisingly, there is no generally agreed-upon protocol for the quantitative and reproducible evaluation of optimization strategies…
Learning new tasks with few samples using related task evaluations.
New benchmark for earthquake forecasting models shows current neural point processes are not yet suitable.
Study proposes a multi-agent system using LLMs for REIT trading, outperforming benchmarks.
Performing controlled experiments on noisy data is essential in understanding deep learning across noise levels. Due to the lack of suitable datasets, previous research has only examined deep learning on controlled synthetic label noise, and real-world label noise has never been studied in a controlled setting. This pa…
Benchmark assesses LLMs' causal inference skills, revealing significant limitations.
Characterizing the dynamic interactive patterns of complex systems helps gain in-depth understanding of how components interrelate with each other while performing certain functions as a whole. In this study, we present a novel multimodal data fusion approach to construct a complex network, which models the interaction…
We present a novel approach to learn binary classifiers when only positive and unlabeled instances are available (PU learning). This problem is routinely cast as a supervised task with label noise in the negative set. We use an ensemble of SVM models trained on bootstrap resamples of the training data for increased rob…
Study uses neural networks to predict firm earnings, outperforming benchmarks and analysts.
We consider learning problems where the training set consists of two types of examples: private and public. The goal is to design a learning algorithm that satisfies differential privacy only with respect to the private examples. This setting interpolates between private learning (where all examples are private) and cl…
Study public-data assisted private stochastic optimization with labeled or unlabeled public data.
We derive an explicit solution for deterministic market impact parameters in the Graewe and Horst (2017) portfolio liquidation model. The model allows to combine various forms of market impact, namely instantaneous, permanent and temporary. We show that the solutions to the two benchmark models of Almgren and Chris (20…
Public pretraining improves private model training even in extreme distribution shift scenarios.
Advancements in neural machinery have led to a wide range of algorithmic solutions for molecular property prediction. Two classes of models in particular have yielded promising results: neural networks applied to computed molecular fingerprints or expert-crafted descriptors, and graph convolutional neural networks that…
Private estimation with public data reduces sample complexity.