High throughput screening of compounds (chemicals) is an essential part of drug discovery [7], involving thousands to millions of compounds, with the purpose of identifying candidate hits. Most statistical tools, including the industry standard B-score method, work on individual compound plates and do not exploit cross…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Bayesian learning improves reliability of molecular predictions for hit compound discovery.
KANEL combines models for early hit enrichment in virtual screening.
SILVR generates new molecules fitting protein binding sites.
CSLVAE generates large chemical libraries efficiently.
Deep learning uses ROC cost functions to improve virtual screening accuracy.
In this work, we attempt to solve the Hit Song Science problem, which aims to predict which songs will become chart-topping hits. We constructed a dataset with approximately 1.8 million hit and non-hit songs and extracted their audio features using the Spotify Web API. We test four models on our dataset. Our best model…
Minimal hitting time on origami equals diophantine type for certain slopes.
Record companies invest billions of dollars in new talent around the globe each year. Gaining insight into what actually makes a hit song would provide tremendous benefits for the music industry. In this research we tackle this question by focussing on the dance hit song classification problem. A database of dance hit …
Price limit trading rules are adopted in some stock markets (especially emerging markets) trying to cool off traders' short-term trading mania on individual stocks and increase market efficiency. Under such a microstructure, stocks may hit their up-limits and down-limits from time to time. However, the behaviors of pri…
This paper improves bond market making by adjusting hit-ratios for client flow quality.
Study bond market making with hit-ratio target using optimal control and HJB equations.
Paper analyzes Hit-and-Run's convergence rates and applies similar methods to randomized Kaczmarz.
In this paper, we investigate the cooling-off effect (opposite to the magnet effect) from two aspects. Firstly, from the viewpoint of dynamics, we study the existence of the cooling-off effect by following the dynamical evolution of some financial variables over a period of time before the stock price hits its limit. S…
The hitting measure is singular and has dimension less than 1 for cocompact Fuchsian groups.
Study estimates Medallion's compounded return before fees at 31.8%.
Researchers prove hitting measure singularity for most Fuchsian and Kleinian groups.
Characterizes measures preserving compound mixed renewal process properties.
The paper improves competitive and dynamic regret bounds for smoothed online learning.
This chapter is an attempt to present a mathematical theory of compound fractional Poisson processes. The chapter begins with the characterization of a well-known Lévy process: The compound Poisson process. The semi-Markov extension of the compound Poisson process naturally leads to the compound fractional Poisson proc…
Large unweighted directed graphs are commonly used to capture relations between entities. A fundamental problem in the analysis of such networks is to properly define the similarity or dissimilarity between any two vertices. Despite the significance of this problem, statistical characterization of the proposed metrics …
Stochastic gradient Langevin dynamics (SGLD) is a fundamental algorithm in stochastic optimization. Recent work by Zhang et al. [2017] presents an analysis for the hitting time of SGLD for the first and second order stationary points. The proof in Zhang et al. [2017] is a two-stage procedure through bounding the Cheege…
Generates natural product-like compounds using GPT models.
We empirically investigated the relationships between the degree of efficiency and the predictability in financial time-series data. The Hurst exponent was used as the measurement of the degree of efficiency, and the hit rate calculated from the nearest-neighbor prediction method was used for the prediction of the dire…
In this paper, we introduce a new model for the risk process based on general compound Hawkes process (GCHP) for the arrival of claims. We call it risk model based on general compound Hawkes process (RMGCHP). The Law of Large Numbers (LLN) and the Functional Central Limit Theorem (FCLT) are proved. We also study the ma…
One of the most important problems of data processing in high energy and nuclear physics is the event reconstruction. Its main part is the track reconstruction procedure which consists in looking for all tracks that elementary particles leave when they pass through a detector among a huge number of points, so-called hi…
In this paper we consider finite volume hyperbolic manifolds X with non-empty totally geodesic boundary. We consider the distribution of the times for the geodesic flow to hit the boundary and derive a formula for the moments of the associated random variable in terms of the orthospectrum. We show that the the first tw…
Normalized compound random measures are flexible nonparametric priors for related distributions. We consider building general nonparametric regression models using normalized compound random measure mixture models. Posterior inference is made using a novel pseudo-marginal Metropolis-Hastings sampler for normalized comp…
Model financial default cascades on sparse graphs via hitting times.
Compound Finance optimizes risk metrics for V3 protocol using Chainrisk simulations.
This study compares two neural models for financial forecasting, showing their superiority.
STMT predicts compounds in unknown areas with trend reflection.
New Riemannian geometry for Compound Gaussian distributions applied to efficient change detection.
Paper reconciles different Ricci flow approaches and proves weak solutions.
Holomorphic map connects Hitchin components to character varieties.
The paper analyzes McKean-Vlasov equations with hitting times, proving global solvability.
ChemGrapher uses deep learning to automatically convert chemical compound images into accurate graphs.
New algorithm determines dimensions of hit spaces in polynomial algebra.
A model for hit song prediction can be used in the pop music industry to identify emerging trends and potential artists or songs before they are marketed to the public. While most previous work formulates hit song prediction as a regression or classification problem, we present in this paper a convolutional neural netw…
One of the most important problems of data processing in high energy and nuclear physics is the event reconstruction. Its main part is the track reconstruction procedure which consists in looking for all tracks that elementary particles leave when they pass through a detector among a huge number of points, so-called hi…
The paper simulates Lévy processes and their extremum and hitting time.
In this paper, we study the classical problem of the first passage hitting density of an Ornstein--Uhlenbeck process. We give two complementary (forward and backward) formulations of this problem and provide semi-analytical solutions for both. The corresponding problems are comparable in complexity. By using the method…
This study deals with the problem of pricing compound options when the underlying asset follows a mixed fractional Brownian motion with jumps. An analytic formula for compound options is derived under the risk neutral measure. Then, these results are applied to value extendible options. Moreover, some special cases of …
SurvSurf predicts first hitting times for intermittent events without monotonic violations.
Parrot learns optimal cache replacement policies using imitation learning.
The study improves compound selection in in silico screening by focusing on model's ability to predict desirable outcomes.
Supervised learning models, also known as quantitative structure-activity regression (QSAR) models, are increasingly used in assisting the process of preclinical, small molecule drug discovery. The models are trained on data consisting of a finite dimensional representation of molecular structures and their correspondi…
A new method solves complex financial problems using deep learning.