Community moderation drifts towards majority, study finds.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Deep learning model detects and corrects outliers in crowd-sourced weather data.
Crowd-sourcing is a cheap and popular means of creating training and evaluation datasets for machine learning, however it poses the problem of `truth inference', as individual workers cannot be wholly trusted to provide reliable annotations. Research into models of annotation aggregation attempts to infer a latent `tru…
BUDS balances privacy and utility by shuffling data, achieving strong privacy with minimal loss.
The paper tackles ranking experts based on their answers to questions, considering statistical and computational challenges.
Proposes a method for interpreting time-varying causal effect moderation in high-dimensional data.
New algorithms find all ε-good arms in stochastic bandits.
Enhances content moderation with culturally-aware models.
New lower bounds show challenges in clustering in moderate dimensions.
A multi-agent system is trialed as a means of crowd-sourcing inexpensive but high quality streams of predictions. Each agent is a microservice embodying statistical models and endowed with economic self-interest. The ability to fork and modify simple agents is granted to a large number of employees in a firm and empiri…
This paper considers extractive summarisation in a comparative setting: given two or more document groups (e.g., separated by publication time), the goal is to select a small number of documents that are representative of each group, and also maximally distinguishable from other groups. We formulate a set of new object…
Optimizes variance reduction in Heston model using large and moderate deviations.
Importance sampling has become an important tool for the computation of tail-based risk measures. Since such quantities are often determined mainly by rare events standard Monte Carlo can be inefficient and importance sampling provides a way to speed up computations. This paper considers moderate deviations for the wei…
Wide-AdGraph detects ads and trackers using a graph of resource requests.
Python models predict stock sentiment for market-beating returns.
We extend previous large deviations results for the randomised Heston model to the case of moderate deviations. The proofs involve the Gärtner-Ellis theorem and sharp large deviations tools.
Optimal learning via moderate deviations theory improves statistical accuracy.
Computes invariants distinguishing between immersions and embeddings of doodles and blobs on surfaces.
New method interprets deep learning for causal effects, separating prognostic and moderating covariates.
We provide a unifying treatment of pathwise moderate deviations for models commonly used in financial applications, and for related integrated functionals. Suitable scaling allows us to transfer these results into small-time, large-time and tail asymptotics for diffusions, as well as for option prices and realised vari…
Unified approach to stochastic Volterra systems' deviations.
Stochastic Gradient Descent shows directional bias with moderate learning rates, impacting optimization outcomes.
We consider call option prices in diffusion models close to expiry, in an asymptotic regime ("moderately out of the money") that interpolates between the well-studied cases of at-the-money options and out-of-the-money fixed-strike options. First and higher order small-time moderate deviation estimates of call prices an…
A new graph-based clustering method for moderate-dimensional data.
We evaluated the effectiveness of an automated bird sound identification system in a situation that emulates a realistic, typical application. We trained classification algorithms on a crowd-sourced collection of bird audio recording data and restricted our training methods to be completely free of manual intervention.…
We report the first, to the best of our knowledge, hand-in-hand collaboration between human rights activists and machine learners, leveraging crowd-sourcing to study online abuse against women on Twitter. On a technical front, we carefully curate an unbiased yet low-variance dataset of labeled tweets, analyze it to acc…
Unintended bias in Machine Learning can manifest as systemic differences in performance for different demographic groups, potentially compounding existing challenges to fairness in society at large. In this paper, we introduce a suite of threshold-agnostic metrics that provide a nuanced view of this unintended bias, by…
The paper analyzes the dynamics of tokens in transformer models at moderate interaction levels.
Paper optimizes change-point detection using learned distributions from training sequences.
In this paper we address a classification problem where two sources of labels with different levels of fidelity are available. Our approach is to combine data from both sources by applying a co-kriging schema on latent functions, which allows the model to account item-dependent labeling discrepancy. We provide an exten…
We define the intrinsic scale at which a network begins to reveal its identity as the scale at which subgraphs in the network (created by a random walk) are distinguishable from similar sized subgraphs in a perturbed copy of the network. We conduct an extensive study of intrinsic scale for several networks, ranging fro…
LLMs overestimate stock returns and are less accurate at predicting extreme outcomes.
Q-learning with cSMART data assesses cAI tailoring variables.
Data analysis require a pairwise proximity measure over objects. Recent work has extended this to situations where the distance information between objects is given as comparison results of distances between three objects (triplets). Humans find the comparison tasks much easier than the exact distance computation and s…
Mosquitoes are the only known vector of malaria, which leads to hundreds of thousands of deaths each year. Understanding the number and location of potential mosquito vectors is of paramount importance to aid the reduction of malaria transmission cases. In recent years, deep learning has become widely used for bioacous…
A framework for optimizing prompt selection in generative language models.
Cluster analysis of very high dimensional data can benefit from the properties of such high dimensionality. Informally expressed, in this work, our focus is on the analogous situation when the dimensionality is moderate to small, relative to a massively sized set of observations. Mathematically expressed, these are the…
A framework uses proxies to prioritize treatment without estimating causal effects.
Recognizing a hotel from an image of a hotel room is important for human trafficking investigations. Images directly link victims to places and can help verify where victims have been trafficked, and where their traffickers might move them or others in the future. Recognizing the hotel from images is challenging becaus…
Study shows noisy data collection in ImageNet leads to biased model performance.
Deep neural networks (DNN) are able to successfully process and classify speech utterances. However, understanding the reason behind a classification by DNN is difficult. One such debugging method used with image classification DNNs is activation maximization, which generates example-images that are classified as one o…
We discuss a general dynamic replication approach to counterparty credit risk modeling. This leads to a fundamental jump-process backward stochastic differential equation (BSDE) for the credit risk adjusted portfolio value. We then reduce the fundamental BSDE to a continuous BSDE. Depending on the close out value conve…
Yelp has been one of the most popular local service search engine in US since 2004. It is powered by crowd-sourced text reviews and photo reviews. Restaurant customers and business owners upload photo images to Yelp, including reviewing or advertising either food, drinks, or inside and outside decorations. It is obviou…
Houdini finds high-dimensional saddle points under few constraints.
Solving real-world problems, particularly with deep learning, relies on the availability of abundant, quality data. In this paper we develop a novel framework that maximises the utility of time-series datasets that contain only small quantities of expertly-labelled data, larger quantities of weakly (or coarsely) labell…
Study describes frequencies of geodesics on hyperbolic surfaces as genus grows.
The reproducing kernel Hilbert space (RKHS) embedding of distributions offers a general and flexible framework for testing problems in arbitrary domains and has attracted considerable amount of attention in recent years. To gain insights into their operating characteristics, we study here the statistical performance of…
The Lax-Hopf formula simplifies the value function of an intertemporal optimization (infinite dimensional) problem associated with a convex transaction-cost function which depends only on the transactions (velocities) of a commodity evolution: it states that the value function is equal to the marginal fonction of a fin…