DisCoveR efficiently discovers declarative process models from event logs.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We describe DyNet, a toolkit for implementing neural network models based on dynamic declaration of network structure. In the static declaration strategy that is used in toolkits like Theano, CNTK, and TensorFlow, the user first defines a computation graph (a symbolic representation of the computation), and then exampl…
In this paper, we consider sequential online prediction (SOP) for streaming data in the presence of outliers and change points. We propose an INstant TEmporal structure Learning (INTEL) algorithm to address this problem. Our INTEL algorithm is developed based on a full consideration of the duality between online predic…
We consider an insurance entity endowed with an initial capital and a surplus process modelled as a Brownian motion with drift. It is assumed that the company seeks to maximise the cumulated value of expected discounted dividends, which are declared or paid in a foreign currency. The currency fluctuation is modelled as…
We study the pricing problem for corporate defaultable bond from the viewpoint of the investors outside the firm that could not exactly know about the information of the firm. We consider the problem for pricing of corporate defaultable bond in the case when the firm value is only declared in some fixed discrete time a…
In this work we present Ludwig, a flexible, extensible and easy to use toolbox which allows users to train deep learning models and use them for obtaining predictions without writing code. Ludwig implements a novel approach to deep learning model building based on two main abstractions: data types and declarative confi…
In this paper, we study the concept of Parisian ruin under the hybrid observation scheme model introduced by Li et al. \cite{binetal2016}. Under this model, the process is observed at Poisson arrival times whenever the business is financially healthy and it is continuously observed when it goes below . The Parisian …
Forecastability measures predictive information across horizons.
We propose a general approach to modeling semi-supervised learning (SSL) algorithms. Specifically, we present a declarative language for modeling both traditional supervised classification tasks and many SSL heuristics, including both well-known heuristics such as co-training and novel domain-specific heuristics. In ad…
Probabilistic programming allows specification of probabilistic models in a declarative manner. Recently, several new software systems and languages for probabilistic programming have been developed on the basis of newly developed and improved methods for approximate inference in probabilistic models. In this contribut…
We propose a technique for declaratively specifying strategies for semi-supervised learning (SSL). The proposed method can be used to specify ensembles of semi-supervised learning, as well as agreement constraints and entropic regularization constraints between these learners, and can be used to model both well-known h…
We analyse perception and memory, using mathematical models for knowledge graphs and tensors, to gain insights into the corresponding functionalities of the human mind. Our discussion is based on the concept of propositional sentences consisting of \textit{subject-predicate-object} (SPO) triples for expressing elementa…
A central tenet of probabilistic programming is that a model is specified exactly once in a canonical representation which is usable by inference algorithms. We describe JointDistributions, a family of declarative representations of directed graphical models in TensorFlow Probability.
In this paper we examine the process involved in the design and implementation of a port-graph model to be used for the analysis of an agent-based rational negligence model. Rational negligence describes the phenomenon that occurred during the financial crisis of 2008 whereby investors chose to trade asset-backed secur…
We address the problem of non-parametric multiple model comparison: given candidate models, decide whether each candidate is as good as the best one(s) or worse than it. We propose two statistical tests, each controlling a different notion of decision errors. The first test, building on the post selection inference…
New architecture separates object state and behavior for better game dynamics.
LIC compiles probabilistic models to generate efficient MCMC proposals.
We discuss memory models which are based on tensor decompositions using latent representations of entities and events. We show how episodic memory and semantic memory can be realized and discuss how new memory traces can be generated from sensory input: Existing memories are the basis for perception and new memories ar…
Consider a Bayesian inference problem where a variable of interest does not take values in a Euclidean space. These "non-standard" data structures are in reality fairly common. They are frequently used in problems involving latent discrete factor models, networks, and domain specific problems such as sequence alignment…
This article presents the use of Answer Set Programming (ASP) to mine sequential patterns. ASP is a high-level declarative logic programming paradigm for high level encoding combinatorial and optimization problem solving as well as knowledge representation and reasoning. Thus, ASP is a good candidate for implementing p…
In situations like tax declarations or analyzes of household budgets we would like to automatically evaluate credibility of exogenous variable (declared income) based on some available (endogenous) variables - we want to build a model and train it on provided data sample to predict (conditional) probability distributio…
Data application developers and data scientists spend an inordinate amount of time iterating on machine learning (ML) workflows -- by modifying the data pre-processing, model training, and post-processing steps -- via trial-and-error to achieve the desired model performance. Existing work on accelerating machine learni…
New algorithms detect anomalies in processes with minimal delay.
Shifts in environment between development and deployment cause classical supervised learning to produce models that fail to generalize well to new target distributions. Recently, many solutions which find invariant predictive distributions have been developed. Among these, graph-based approaches do not require data fro…
We build deep RL agents that execute declarative programs expressed in formal language. The agents learn to ground the terms in this language in their environment, and can generalize their behavior at test time to execute new programs that refer to objects that were not referenced during training. The agents develop di…
A new method for math reasoning that allows for iterative correction.
In this dissertation, I derive a new method to estimate the Vapnik-Chervonenkis Dimension (VCD) for the class of linear functions. This method is inspired by the technique developed by Vapnik et al. Vapnik et al. (1994). My contribution rests on the approximation of the expected maximum difference between two empirical…
CUBE explains models by balanced experiments and contrasts.
Comment on a theorem about surface determination from map data.
System detects financial forecasts in tweets, achieving high precision.
With pressure to increase graduation rates and reduce time to degree in higher education, it is important to identify at-risk students early. Automated early warning systems are therefore highly desirable. In this paper, we use unsupervised clustering techniques to predict the graduation status of declared majors in fi…
Combining deep neural networks with structured logic rules is desirable to harness flexibility and reduce uninterpretability of the neural models. We propose a general framework capable of enhancing various types of neural networks (e.g., CNNs and RNNs) with declarative first-order logic rules. Specifically, we develop…
Today, the dominant paradigm for training neural networks involves minimizing task loss on a large dataset. Using world knowledge to inform a model, and yet retain the ability to perform end-to-end training remains an open question. In this paper, we present a novel framework for introducing declarative knowledge to ne…
OpenAlpha validates decentralized capital strategies using game theory and market aggregation.
The Euclidean cone metrics coming from q-differentials on a closed surface of genus g > 1 define an equivalence relation on homotopy classes of closed curves declaring two to be equivalent if they have the equal length in every such metric. We prove an analog of the result of Randol for hyperbolic metrics (building on …
Reinforcement learning approaches have long appealed to the data management community due to their ability to learn to control dynamic behavior from raw system performance. Recent successes in combining deep neural networks with reinforcement learning have sparked significant new interest in this domain. However, pract…
There are no solid arguments to sustain that digital currencies are the future of online payments or the disruptive technology that some of its former participants declared when used to face critiques. This paper aims to solve the cryptocurrency puzzle from a behavioral finance perspective by finding the parallelism be…
Improved portfolio optimization reduces sensitivity to neural network initialization.
Incorporating constraints is a major concern in probabilistic machine learning. A wide variety of problems require predictions to be integrated with reasoning about constraints, from modelling routes on maps to approving loan predictions. In the former, we may require the prediction model to respect the presence of phy…
In his 1954 paper about the initial value problem for 2D hyperbolic nonlinear PDEs, P. Lax declared that he had "a strong reason to believe" that there must exist a well-defined class of "not genuinely nonlinear" nonlinear PDEs. In 1978 G. Boillat coined the term "completely exceptional" to denote it. In the case of $2…
Asymmetry PRISM outperforms CPU and GPU solvers for institutional rebalancing.
Study explores how labour income impacts optimal bankruptcy strategy.
Tree-AMP simplifies inference in complex tree-structured models.
Examines SOFR derivatives pricing and hedging post-LIBOR discontinuation.
We propose a non-parametric anomaly detection algorithm for high dimensional data. We score each datapoint by its average -NN distance, and rank them accordingly. We then train limited complexity models to imitate these scores based on the max-margin learning-to-rank framework. A test-point is declared as an anomaly…
Study compares DSPy teleprompter algorithms for aligning LLM evaluations with human annotations.
Automated method bounds causal effects in discrete data.
Mathematical models help keep vaccine prices low.