New estimator improves off-policy evaluation for large action spaces.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Bayesian deep learning predicts satellite collisions.
We consider a Black-Scholes market in which a number of stocks and an index are traded. The simplified Capital Asset Pricing Model is the conjunction of the usual Capital Asset Pricing Model, or CAPM, and the statement that the appreciation rate of the index is equal to its squared volatility plus the interest rate. (T…
Cross-sectional signatures of market panic were recently discussed on daily time scales in [1], extended here to a study of cross-sectional properties of stocks on intra-day time scales. We confirm specific intra-day patterns of dispersion and kurtosis, and find that the correlation across stocks increases in times of …
Space debris warnings follow a predictable pattern, allowing timely satellite maneuvers.
Typical spoken language understanding systems provide narrow semantic parses using a domain-specific ontology. The parses contain intents and slots that are directly consumed by downstream domain applications. In this work we discuss expanding such systems to handle compound entities and intents by introducing a domain…
Query2box embeds complex queries as boxes to handle logical operations in large KGs.
In many healthcare settings, intuitive decision rules for risk stratification can help effective hospital resource allocation. This paper introduces a novel variant of decision tree algorithms that produces a chain of decisions, not a general tree. Our algorithm, -Carving Decision Chain (ACDC), sequentially carves o…
We examine how recently documented, fundamental phenomena in deep learning models subject to pruning are affected by changes in the pruning procedure. Specifically, we analyze differences in the connectivity structure and learning dynamics of pruned models found through a set of common iterative pruning techniques, to …
Yield curve modeling is an essential problem in finance. In this work, we explore the use of Bayesian statistical methods in conjunction with Nelson-Siegel model. We present the hierarchical Bayesian model for the parameters of the Nelson-Siegel yield function. We implement the MAP estimates via BFGS algorithm in rstan…
New method estimates treatment effects over time with unobserved confounders.
We propose a novel exponentially-modified Gaussian (EMG) mixture residual model. The EMG mixture is well suited to model residuals that are contaminated by a distribution with positive support. This is in contrast to commonly used robust residual models, like the Huber loss or , which assume a symmetric contami…
We propose the Gaussian attention model for content-based neural memory access. With the proposed attention model, a neural network has the additional degree of freedom to control the focus of its attention from a laser sharp attention to a broad attention. It is applicable whenever we can assume that the distance in t…
This paper introduces the combinatorial Boolean model (CBM), which is defined as the class of linear combinations of conjunctions of Boolean attributes. This paper addresses the issue of learning CBM from labeled data. CBM is of high knowledge interpretability but naïve learning of it requires exponentially large compu…
Numerous empirical evidence has corroborated that the noise plays a crucial rule in effective and efficient training of neural networks. The theory behind, however, is still largely unknown. This paper studies this fundamental problem through training a simple two-layer convolutional neural network model. Although trai…
Restricted Boltzmann Machines (RBMs) are one of the fundamental building blocks of deep learning. Approximate maximum likelihood training of RBMs typically necessitates sampling from these models. In many training scenarios, computationally efficient Gibbs sampling procedures are crippled by poor mixing. In this work w…
Bayesian model for discrete data with conditional transformations.
While great progress has been made at making neural networks effective across a wide range of visual tasks, most models are surprisingly vulnerable. This frailness takes the form of small, carefully chosen perturbations of their input, known as adversarial examples, which represent a security threat for learned vision …
We investigate connections between information-theoretic and estimation-theoretic quantities in vector Poisson channel models. In particular, we generalize the gradient of mutual information with respect to key system parameters from the scalar to the vector Poisson channel model. We also propose, as another contributi…
Study on list learning with noisy data, showing limits and some learnable cases.
This paper presents a new family of backpropagation-free neural architectures, Gated Linear Networks (GLNs). What distinguishes GLNs from contemporary neural networks is the distributed and local nature of their credit assignment mechanism; each neuron directly predicts the target, forgoing the ability to learn feature…
For a learning task, data can usually be collected from different sources or be represented from multiple views. For example, laboratory results from different medical examinations are available for disease diagnosis, and each of them can only reflect the health state of a person from a particular aspect/view. Therefor…
Bayesian Context Trees model improves financial time series forecasting.
A simple regularization technique speeds up training of Neural ODEs.
This paper introduces a new method to better understand financial market causality.
Machine learning competition predicts spacecraft collision risks.
Paper proposes ensemble distillation for well-calibrated structured prediction.
We present a general numerical approach for learning unknown dynamical systems using deep neural networks (DNNs). Our method is built upon recent studies that identified the residue network (ResNet) as an effective neural network structure. In this paper, we present a generalized ResNet framework and broadly define res…
GMED edits stored examples to improve continual learning.
Develops a new GLM framework for claims reserving with adaptive estimation.
A Bayesian network is a graphical model that encodes probabilistic relationships among variables of interest. When used in conjunction with statistical techniques, the graphical model has several advantages for data analysis. One, because the model encodes dependencies among all variables, it readily handles situations…
TMLE improves causal effect estimation in missing data scenarios with various positivity violations.
We address the problem of general supervised learning when data can only be accessed through an (indefinite) similarity function between data points. Existing work on learning with indefinite kernels has concentrated solely on binary/multi-class classification problems. We propose a model that is generic enough to hand…
We consider the problem of designing a sparse Gaussian process classifier (SGPC) that generalizes well. Viewing SGPC design as constructing an additive model like in boosting, we present an efficient and effective SGPC design method to perform a stage-wise optimization of a predictive loss function. We introduce new me…
Bayesian approach sparsifies neural networks efficiently.
It is becoming increasingly important to understand the vulnerability of machine learning models to adversarial attacks. In this paper we study the feasibility of robust learning from the perspective of computational learning theory, considering both sample and computational complexity. In particular, our definition of…
Normalization techniques have only recently begun to be exploited in supervised learning tasks. Batch normalization exploits mini-batch statistics to normalize the activations. This was shown to speed up training and result in better models. However its success has been very limited when dealing with recurrent neural n…
We obtain an expression for the curvature of the Lie group SDiff and use it to derive Lukatskii's formula for the case where is locally Euclidean. We discuss qualitatively some previous findings for SDiff in conjunction with our result.
The article classifies cubiquitous sublattices and applies them to branched covers.
In this work, we introduce a new method for imitation learning from video demonstrations. Our method, Relational Mimic (RM), improves on previous visual imitation learning methods by combining generative adversarial networks and relational learning. RM is flexible and can be used in conjunction with other recent advanc…
We propose a method based on finite mixture models for classifying a set of observations into number of different categories. In order to demonstrate the method, we show how the component densities for the mixture model can be derived by using the maximum entropy method in conjunction with conservation of Pythagorean m…
Multi-step ahead forecasting is still an open challenge in time series forecasting. Several approaches that deal with this complex problem have been proposed in the literature but an extensive comparison on a large number of tasks is still missing. This paper aims to fill this gap by reviewing existing strategies for m…
Paper proposes consistent estimators for learning to defer decisions to experts.
We prove that there are examples of finitely generated groups G together with group ring elements Q \in \bbQ G for which the von Neumann dimension \dim_{LG}\ker Q is irrational, so (in conjunction with other known results) answering a question of Atiyah.
The paper introduces closed-form expressions for interpreting Tsetlin Machines.
Optimizes experimental design using synthetic controls for better outcomes.
We follow the main stocks belonging to the New York Stock Exchange and to Nasdaq from 2003 to 2012, through years of normality and of crisis, and study the dynamics of networks built on two measures expressing relations between those stocks: correlation, which is symmetric and measures how similar two stocks behave, an…
Deep metric learning has been demonstrated to be highly effective in learning semantic representation and encoding information that can be used to measure data similarity, by relying on the embedding learned from metric learning. At the same time, variational autoencoder (VAE) has widely been used to approximate infere…