This paper evaluates AI ethics guidelines and their implementation.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This work simplifies choosing reinforcement-learning algorithms.
Annotation guidelines used to guide the annotation of training and evaluation datasets can have a considerable impact on the quality of machine learning models. In this study, we explore the effects of annotation guidelines on the quality of app feature extraction models. As a main result, we propose several changes to…
Framework for applying GPs to real-world data with scalability guidelines.
Proposes guidelines for developing medical AI products.
Guidelines for using explainable ML to avoid misuse.
Paper derives policy rules from observational data for hepatitis C treatment.
Guidelines for deploying deep learning models on smartphones are developed.
Bayesian optimisation has gained great popularity as a tool for optimising the parameters of machine learning algorithms and models. Somewhat ironically, setting up the hyper-parameters of Bayesian optimisation methods is notoriously hard. While reasonable practical solutions have been advanced, they can often fail to …
Consistently checking the statistical significance of experimental results is one of the mandatory methodological steps to address the so-called "reproducibility crisis" in deep reinforcement learning. In this tutorial paper, we explain how the number of random seeds relates to the probabilities of statistical errors. …
Scale of data and scale of computation infrastructures together enable the current deep learning renaissance. However, training large-scale deep architectures demands both algorithmic improvement and careful system configuration. In this paper, we focus on employing the system approach to speed up large-scale training.…
Sherpa.ai framework combines federated learning and differential privacy for edge AI services.
Deep neural networks (DNNs) have been proven to have many redundancies. Hence, many efforts have been made to compress DNNs. However, the existing model compression methods treat all the input samples equally while ignoring the fact that the difficulties of various input samples being correctly classified are different…
The study analyzes convergence rates for sparse pivotal estimators in high-dimensional regression.
This chapter introduces reproducibility in machine learning for medical imaging.
Survival analysis models predict economic convergence across Americas.
The paper analyzes SMOTE for imbalanced classification, providing theoretical bounds and guidelines.
The study identifies patient subgroups with enhanced or diminished opioid treatment effects.
Robust low-rank matrix estimation is a topic of increasing interest, with promising applications in a variety of fields, from computer vision to data mining and recommender systems. Recent theoretical results establish the ability of such data models to recover the true underlying low-rank matrix when a large portion o…
Paper offers guidelines for using ML in cyber security, focusing on botnet detection.
The paper provides guidelines for choosing between SBI methods in complex biological models.
New framework quantifies and reduces concept-based models' leakage.
System exposes study population descriptions in clinical guidelines.
SaML guides ML models to avoid survey biases.
Paper analyzes and simplifies deep learning framework parallelism for better performance.
The study examines how experimental design choices affect machine learning model performance.
Codebook for Institutional Grammar 2.0 simplifies policy encoding.
This study improves lookahead Bayesian optimization using rollout approximation.
For a long time, designing neural architectures that exhibit high performance was considered a dark art that required expert hand-tuning. One of the few well-known guidelines for architecture design is the avoidance of exploding gradients, though even this guideline has remained relatively vague and circumstantial. We …
Improves interpretability of anomaly scores in GBRBM-based detection.
Approximate Incremental Value-at-Risk formulae provide an easy-to-use preliminary guideline for risk allocation. Both the cases of risk adding and risk pooling are examined and beta-based formulae achieved. Results highlight how much the conditions for adding new risky positions are stronger than those required for ris…
New tuning rules for Metropolis algorithms derived from Bayesian large-sample asymptotics.
Infinite-dimensional diffusion models tackle generative tasks for complex data.
Self-attention models benefit equally from width and depth, but beyond a certain point, depth becomes less efficient.
Making applications aware of the mobility experienced by the user can open the door to a wide range of novel services in different use-cases, from smart parking to vehicular traffic monitoring. In the literature, there are many different studies demonstrating the theoretical possibility of performing Transportation Mod…
This study evaluates different normalizing flow architectures for MCMC.
This manuscript addresses the problem of the automatic lesion boundary detection in dermoscopy, using deep neural networks. An approach is based on the adaptation of the U-net convolutional neural network with skip connections for lesion boundary segmentation task. I hope this paper could serve, to some extent, as an e…
We present in this paper an empirical framework motivated by the practitioner point of view on stability. The goal is to both assess clustering validity and yield market insights by providing through the data perturbations we propose a multi-view of the assets' clustering behaviour. The perturbation framework is illust…
Adjusted for chance measures are widely used to compare partitions/clusterings of the same data set. In particular, the Adjusted Rand Index (ARI) based on pair-counting, and the Adjusted Mutual Information (AMI) based on Shannon information theory are very popular in the clustering community. Nonetheless it is an open …
We present Generative Adversarial Capsule Network (CapsuleGAN), a framework that uses capsule networks (CapsNets) instead of the standard convolutional neural networks (CNNs) as discriminators within the generative adversarial network (GAN) setting, while modeling image data. We provide guidelines for designing CapsNet…
We conduct an extensive evaluation of price jump tests based on high-frequency financial data. After providing a concise review of multiple alternative tests, we document the size and power of all tests in a range of empirically relevant scenarios. Particular focus is given to the robustness of test performance to the …
This paper proposes a use of an ordinal classifier to evaluate the financial solidity of non-life insurance companies as strong, moderate, weak, and insolvency. This study constructed an efficient classification model that can be used by regulators to evaluate the financial solidity and to determine the priority of fur…
Study differential and integral calculus on noncommutative C*-algebras.
In this paper, simple mathematical models from Control Theory are applied to three very important economic paradigms, namely (a) minimum wages in self-regulating markets, (b) market-versus-true values and currency rates, and (c) government spending and taxation levels. Analytical solutions are provided in all three par…
Multi-task learning (MTL) has led to successes in many applications of machine learning, from natural language processing and speech recognition to computer vision and drug discovery. This article aims to give a general overview of MTL, particularly in deep neural networks. It introduces the two most common methods for…
Used to estimate the risk of an estimator or to perform model selection, cross-validation is a widespread strategy because of its simplicity and its apparent universality. Many results exist on the model selection performances of cross-validation procedures. This survey intends to relate these results to the most recen…
The paper examines the reliability of limit order book representations in the face of data perturbation.
Clustering is an essential data mining tool that aims to discover inherent cluster structure in data. For most applications, applying clustering is only appropriate when cluster structure is present. As such, the study of clusterability, which evaluates whether data possesses such structure, is an integral part of clus…