Theoretical justification for asymmetric actor-critic algorithms in reinforcement learning.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Deep neural networks justify medical diagnoses with textual explanations.
Data-driven decision-making often overestimates benefits due to the winner's curse.
Review of latest DRL algorithms with theoretical and practical insights.
In this paper we address the problem of discovering a small set of frequent serial episodes from sequential data so as to adequately characterize or summarize the data. We discuss an algorithm based on the Minimum Description Length (MDL) principle and the algorithm is a slight modification of an earlier method, called…
Theoretical justification for image inpainting using diffusion models.
Study on random matrices in deep neural networks using Gaussian data.
In lexicon-based classification, documents are assigned labels by comparing the number of words that appear from two opposed lexicons, such as positive and negative sentiment. Creating such words lists is often easier than labeling instances, and they can be debugged by non-experts if classification performance is unsa…
The paper gives a simple algebraic description, and background justification, for the Bowley Ratio, the relative returns to labour and capital, in a simple economy.
New approach to topic modelling with covariates for large text corpora.
This paper develops a theory for group Lasso using a concept called strong group sparsity. Our result shows that group Lasso is superior to standard Lasso for strongly group-sparse signals. This provides a convincing theoretical justification for using group sparse regularization when the underlying group structure is …
New approach links machine learning reliability to epistemic uncertainty.
DACE estimates covariance from compressed data, improving accuracy.
Extends neural network approximations to guarantee continuity of real-world learning tasks.
This paper considers the problem of subspace clustering under noise. Specifically, we study the behavior of Sparse Subspace Clustering (SSC) when either adversarial or random noise is added to the unlabelled input data points, which are assumed to be in a union of low-dimensional subspaces. We show that a modified vers…
Word2vec analysis reveals spectral underpinnings.
POWSS simplifies Q-value estimation in POMDPs with continuous observations.
New theory explains GAN's high quality but low diversity.
Proposes a Bayesian approach to explain, justify, and quantify uncertainty in DNNs.
Observational studies are rising in importance due to the widespread accumulation of data in fields such as healthcare, education, employment and ecology. We consider the task of answering counterfactual questions such as, "Would this patient have lower blood sugar had she received a different medication?". We propose …
We introduce a framework for analyzing transductive combination of Gaussian process (GP) experts, where independently trained GP experts are combined in a way that depends on test point location, in order to scale GPs to big data. The framework provides some theoretical justification for the generalized product of GP e…
This note justifies approximations of arithmetic forwards using weighted averages of overnight forwards.
Enhances reinforcement learning from sparse data.
First we express the holonomy along a boundary curve as the integral on the domain, of an expression which is linear in the curvature. Then we provide a rigorous justification of the definition of curvature in Regge calculus.
Post-hoc calibration of neural networks using g-Layers proves theoretical justification.
In this paper, we study the nonnegative matrix factorization problem under the separability assumption (that is, there exists a cone spanned by a small subset of the columns of the input nonnegative data matrix containing all columns), which is equivalent to the hyperspectral unmixing problem under the linear mixing mo…
The paper provides a theoretical justification for using stable SSM blocks in deep sequential models.
To analyse a very large data set containing lengthy variables, we adopt a sequential estimation idea and propose a parallel divide-and-conquer method. We conduct several conventional sequential estimation procedures separately, and properly integrate their results while maintaining the desired statistical properties. A…
We provide justifications for two questions on special maps on subgroups of the reals. We will show that the questions can be treated from different points of view. We also discuss two versions of Anderson's Involution Conjecture.
This work incorporates topological features via persistence diagrams to classify point cloud data arising from materials science. Persistence diagrams are multisets summarizing the connectedness and holes of given data. A new distance on the space of persistence diagrams generates relevant input features for a classifi…
Mapper classifier improves robustness over CNNs.
Traditionally, practitioners initialize the {\tt k-means} algorithm with centers chosen uniformly at random. Randomized initialization with uneven weights ({\tt k-means++}) has recently been used to improve the performance over this strategy in cost and run-time. We consider the k-means problem with semi-supervised inf…
This paper contains the motivation for the study of critical surfaces. In previous work the only justification given for the definition of this new class of surfaces is the strength of the results. However, when viewed as the topological analogue to index 2 minimal surfaces, critical surfaces become quite natural.
We refine the analysis of hedging strategies for options under the SABR model carried out in [2]. In particular, we provide a theoretical justification of the empirical observation made in [2] that the modified delta ("Bartlett's delta") introduced there provides a more accurate and robust hedging strategy than the con…
New formulation of MIL using shapelets for better classifier of bags.
This technical report constructs a theoretical framework to relate standard Taylor approximation based optimisation methods with Natural Gradient (NG), a method which is Fisher efficient with probabilistic models. Such a framework will be shown to also provide mathematical justification to combine higher order methods …
We discuss a general method to learn data representations from multiple tasks. We provide a justification for this method in both settings of multitask learning and learning-to-learn. The method is illustrated in detail in the special case of linear feature learning. Conditions on the theoretical advantage offered by m…
Paper justifies ST estimator using pWGF and proposes an improved variant.
The paper considers a general semi-Markov model for Limit Order Books with two states, which incorporates price changes that are not fixed to one tick. Furthermore, we introduce an even more general case of the semi-Markov model for LimitOrder Books that incorporates an arbitrary number of states for the price changes.…
Enhances SDR via Hellinger correlation for better data dependency understanding.
We describe a method for removing the effect of confounders in order to reconstruct a latent quantity of interest. The method, referred to as half-sibling regression, is inspired by recent work in causal inference using additive noise models. We provide a theoretical justification and illustrate the potential of the me…
Paper introduces a new robust method for estimating Pareto tail index from grouped data.
Classically time is kept fixed for infinitesimal variations in problems in mechanics. Apparently, there appears to be no mathematical justification in the literature for this standard procedure. This can be explained canonically by unveiling the intrinsic mathematical structure of time in Lagrangian mechanics. Moreover…
The paper provides theoretical guarantees for transformation-based models in variational inference.
Fine-grained pretraining improves neural network's ability to learn rare features.
The modern scale of data has brought new challenges to Bayesian inference. In particular, conventional MCMC algorithms are computationally very expensive for large data sets. A promising approach to solve this problem is embarrassingly parallel MCMC (EP-MCMC), which first partitions the data into multiple subsets and r…
Contrastive divergence (CD) is a promising method of inference in high dimensional distributions with intractable normalizing constants, however, the theoretical foundations justifying its use are somewhat shaky. This document proposes a framework for understanding CD inference, how/when it works, and provides multiple…
Large-scale kernel approximation is an important problem in machine learning research. Approaches using random Fourier features have become increasingly popular [Rahimi and Recht, 2007], where kernel approximation is treated as empirical mean estimation via Monte Carlo (MC) or Quasi-Monte Carlo (QMC) integration [Yang …