This thesis advances algorithms and software for QMC, GP, and sciML.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Unified framework connects physical laws and machine learning.
Novel framework for ML-assisted inference valid for any statistical task.
Financial institutions have massive computations to carry out overnight which are very demanding in terms of the consumed CPU. The challenge is to price many different products on a cluster-like architecture. We have used the Premia software to valuate the financial derivatives. In this work, we explain how Premia can …
Differentiable programming aids in solving differential equations and their sensitivities.
The scientific literature is a rich source of information for data mining with conceptual knowledge graphs; the open science movement has enriched this literature with complementary source code that implements scientific models. To exploit this new resource, we construct a knowledge graph using unsupervised learning me…
arfpy simplifies data generation with adversarial random forests.
In this paper we form a general conservation law that unifies a class of physics field theories. For this we first introduce the notion of a general field as a formal sum differential forms on a Minkowski manifold. Thereafter, we employ the action principle to define the conservation law for such general fields. By con…
Targeted Learning uses robust statistics for reproducible research.
In this paper we study the adaptive learnability of decision trees of depth at most from membership queries. This has many applications in automated scientific discovery such as drugs development and software update problem. Feldman solves the problem in a randomized polynomial time algorithm that asks $\tilde O(2^…
Develops neural networks for reductive Lie groups, enhancing symmetry respect.
Python package for functional data analysis.
The reproducibility of scientific research has become a point of critical concern. We argue that openness and transparency are critical for reproducibility, and we outline an ecosystem for open and transparent science that has emerged within the human neuroimaging community. We discuss the range of open data sharing re…
CARVE validates clustering results using resampling and stability analysis.
DRIFT uses RL to automate functional software testing efficiently.
Multimodal deep learning improves flaw detection in software programs.
KnotPlot helps beginners and veterans use software for visualizing knots.
Brief history and challenges of interpretable machine learning.
Can engineering neural networks be approached in a disciplined way similar to how engineers build software for civil aircraft? We present nn-dependability-kit, an open-source toolbox to support safety engineering of neural networks for autonomous driving systems. The rationale behind nn-dependability-kit is to consider…
Research proposes an ensemble learning model for efficient software defect prediction.
Improving software quality through effective organizational learning.
This paper introduces Sigma, a domain-specific computational representation for collaboration in large-scale for the field of economics. A computational representation is not a programming language or a software platform. A computational representation is a domain-specific representation system based on three specific …
Improved software flaw detection using NAS on multimodal DL models.
Method predicts hardware resource usage by control software with guaranteed linear convergence.
Software development effort estimation is considered a fundamental task for software development life cycle as well as for managing project cost, time and quality. Therefore, accurate estimation is a substantial factor in projects success and reducing the risks. In recent years, software effort estimation has received …
Existing language models such as n-grams for software code often fail to capture a long context where dependent code elements scatter far apart. In this paper, we propose a novel approach to build a language model for software code to address this particular issue. Our language model, partly inspired by human memory, i…
This paper tackles co-design of neural hardware and software to improve efficiency.
The public package registry npm is one of the biggest software registry. With its 216 911 software packages, it forms a big network of software dependencies. In this paper we evaluate various methods for finding similar packages in the npm network, using only the structure of the graph. Namely, we want to find a way of…
Although software analytics has experienced rapid growth as a research area, it has not yet reached its full potential for wide industrial adoption. Most of the existing work in software analytics still relies heavily on costly manual feature engineering processes, and they mainly address the traditional classification…
The purpose of this study is to introduce new design-criteria for next-generation hyperparameter optimization software. The criteria we propose include (1) define-by-run API that allows users to construct the parameter search space dynamically, (2) efficient implementation of both searching and pruning strategies, and …
ParaDRAM automates parallel MCMC simulations across languages.
Survey of software developers' experience with Github Copilot tool.
Existing malware detectors on safety-critical devices have difficulties in runtime detection due to the performance overhead. In this paper, we introduce PROPEDEUTICA, a framework for efficient and effective real-time malware detection, leveraging the best of conventional machine learning (ML) and deep learning (DL) te…
A deep-learning inference accelerator is synthesized from a C-language software program parallelized with Pthreads. The software implementation uses the well-known producer/consumer model with parallel threads interconnected by FIFO queues. The LegUp high-level synthesis (HLS) tool synthesizes threads into parallel FPG…
RAMANMETRIX simplifies Raman spectroscopy data analysis.
Estimating mutual information from i.i.d. samples drawn from an unknown joint density function is a basic statistical problem of broad interest with multitudinous applications. The most popular estimator is one proposed by Kraskov and Stögbauer and Grassberger (KSG) in 2004, and is nonparametric and based on the distan…
Software helps teach latent variable methods in multivariate data analytics.
Estimates yearly improvement rates for nearly all technologies using US patent data.
A new method avoids noise amplification when subtracting or dividing stochastic signals.
Software estimates inequality in random systems with changing communities.
DUE framework models unknown equations from data using deep learning.
xVal tokenizes numbers continuously for better scientific model training.
AEC Games model represents software MARL environments better than POSGs.
Galactica learns from scientific literature to help researchers.
Fenrir efficiently estimates Bayesian MLN-DLMs for scalable inference.
Certain research strands can yield "forbidden knowledge". This term refers to knowledge that is considered too sensitive, dangerous or taboo to be produced or shared. Discourses about such publication restrictions are already entrenched in scientific fields like IT security, synthetic biology or nuclear physics researc…
Forest management relies on the evaluation of silviculture practices. The increase in natural risk due to climate change makes it necessary to consider evaluation criteria that take natural risk into account. Risk integration in existing software requires advanced programming skills.We propose a user-friendly software …
Data science models, although successful in a number of commercial domains, have had limited applicability in scientific problems involving complex physical phenomena. Theory-guided data science (TGDS) is an emerging paradigm that aims to leverage the wealth of scientific knowledge for improving the effectiveness of da…