Algorithm learns from offline data to improve performance in target environment.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We propose an offline-online procedure for Fourier transform based option pricing. The method supports the acceleration of such essential tasks of mathematical finance as model calibration, real-time pricing, and, more generally, risk assessment and parameter risk estimation. We adapt the empirical magic point interpol…
Due to recent technical and scientific advances, we have a wealth of information hidden in unstructured text data such as offline/online narratives, research articles, and clinical reports. To mine these data properly, attributable to their innate ambiguity, a Word Sense Disambiguation (WSD) algorithm can avoid numbers…
Framework optimizes battery storage for markets by separating long-term degradation from short-term market dynamics.
Linear optimization is many times algorithmically simpler than non-linear convex optimization. Linear optimization over matroid polytopes, matching polytopes and path polytopes are example of problems for which we have simple and efficient combinatorial algorithms, but whose non-linear convex counterpart is harder and …
Paper addresses RLHF alignment challenges with novel algorithms.
Efficiently learns and transports posterior densities for real-time inference.
Machine learning has started to be deployed in fields such as healthcare and finance, which propelled the need for and growth of privacy-preserving machine learning (PPML). We propose an actively secure four-party protocol (4PC), and a framework for PPML, showcasing its applications on four of the most widely-known mac…
A new method estimates time-varying parameters in earth system models using offline and online data assimilation.
This paper bridges offline and online RL by studying policy finetuning with a reference policy.
Algorithm reduces online regret by leveraging offline data in linear bandits.