Truncated backpropagation through time (TBPTT) is a popular method for learning in recurrent neural networks (RNNs) that saves computation and memory at the cost of bias by truncating backpropagation after a fixed number of lags. In practice, choosing the optimal truncation length is difficult: TBPTT will not converge …
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Adaptive Nucleus Truncation Improves Long-Form Reasoning
Adaptive truncation improves privacy in online Bayesian estimation.
The positivity assumption, or the experimental treatment assignment (ETA) assumption, is important for identifiability in causal inference. Even if the positivity assumption holds, practical violations of this assumption may jeopardize the finite sample performance of the causal estimator. One of the consequences of pr…
UDN adapts depth to data complexity, outperforming standard neural networks.
Paper proposes approximate Stein classes for efficient truncated density estimation.
Current variational inference methods for hierarchical Bayesian nonparametric models can neither characterize the correlation structure among latent variables due to the mean-field setting, nor infer the true posterior dimension because of the universal truncation. To overcome these limitations, we propose the conditio…
We consider large scale empirical risk minimization (ERM) problems, where both the problem dimension and variable size is large. In these cases, most second order methods are infeasible due to the high cost in both computing the Hessian over all samples and computing its inverse in high dimensions. In this paper, we pr…
Efficiently estimate Boolean product distribution parameters from truncated samples.
In this paper, we develop a general theory of truncated inverse binomial sampling. In this theory, the fixed-size sampling and inverse binomial sampling are accommodated as special cases. In particular, the classical Chernoff-Hoeffding bound is an immediate consequence of the theory. Moreover, we propose a rigorous and…
A new algorithm speeds up elliptical slice sampling for truncated multivariate normals.
As in standard linear regression, in truncated linear regression, we are given access to observations whose dependent variable equals , where is some fixed unknown vector of interest and is independent noise; except we are only given an observation if its dep…
Stochastic gradient descent (SGD) is commonly used for optimization in large-scale machine learning problems. Langford et al. (2009) introduce a sparse online learning method to induce sparsity via truncated gradient. With high-dimensional sparse data, however, the method suffers from slow convergence and high variance…
Recent work has demonstrated the effectiveness of gradient descent for directly recovering the factors of low-rank matrices from random linear measurements in a globally convergent manner when initialized properly. However, the performance of existing algorithms is highly sensitive in the presence of outliers that may …
Adaptive Monte Carlo methods are recent variance reduction techniques. In this work, we propose a mathematical setting which greatly relaxes the assumptions needed by for the adaptive importance sampling techniques presented by Vazquez-Abad and Dufresne, Fu and Su, and Arouna. We establish the convergence and asymptoti…
Improves FI-PINNs by combining re-sampling and subset simulation for better failure probability estimation.
Derives equations for deep learning biases and weights, showing data complexity reduction.
Graph-structured data arise ubiquitously in many application domains. A fundamental problem is to quantify their similarities. Graph kernels are often used for this purpose, which decompose graphs into substructures and compare these substructures. However, most of the existing graph kernels do not have the property of…
Paper tackles unbounded density ratio estimation for covariate shift adaptation.
ALTBI enhances outlier detection by maximizing the inlier-memorization effect.
Recommender systems are widely used to recommend the most appealing items to users. These recommendations can be generated by applying collaborative filtering methods. The low-rank matrix completion method is the state-of-the-art collaborative filtering method. In this work, we show that the skewed distribution of rati…
Computing partition functions, the normalizing constants of probability distributions, is often hard. Variants of importance sampling give unbiased estimates of a normalizer Z, however, unbiased estimates of the reciprocal 1/Z are harder to obtain. Unbiased estimates of 1/Z allow Markov chain Monte Carlo sampling of "d…
The problem of an arbitrary truncated Levy flight description using the method of cumulant approach has been solved. The set of cumulants of the truncated Levy distribution given the assumption of arbitrary truncation has been found. The influence of truncation shape on the truncated Levy flight properties in the Gauss…
In the paper "On Truncated Variation of Brownian Motion with Drift" (Bull. Pol. Acad. Sci. Math. 56 (2008), no.4, 267 - 281) we defined truncated variation of Brownian motion with drift, where is a standard Brownian motion. Truncated variation differs from regular variation by neglect…
Study develops smart contract framework for procurement under demand variability.
Survey of extreme value modeling techniques for insurance.
In this paper we propose a new kind of high order numerical scheme for backward stochastic differential equations(BSDEs). Unlike the traditional -scheme, we reduce truncation errors by taking carefully for every subinterval according to the characteristics of integrands. We give error estimates of this nonlinear…
TSNPE improves SBI efficiency and scalability.
Optimal algorithm learns Gaussian under halfspace truncation with minimal samples.
New method for constructing truncated vine copulas.
Non-negative matrix factorization (NMF) minimizes the Euclidean distance between the data matrix and its low rank approximation, and it fails when applied to corrupted data because the loss function is sensitive to outliers. In this paper, we propose a Truncated CauchyNMF loss that handle outliers by truncating large e…
New algorithm improves inference for flexible models with infinite latent features.
Variance-Calibrated Modulation (VCM) addresses the likelihood trap in LLMs by reshaping the probability distribution before truncation.
Paper defines new risk measures for elliptical distributions.
Adaptive algorithm for multi-objective optimization with binary constraints.
New DP framework using data truncation for efficient estimation.
Unified framework for mean testing under truncation bias.
Score matching method improves density estimation for truncated data on manifolds.
This article reviews and explains HMC-based methods for sampling constrained continuous distributions.
Truncated densities are probability density functions defined on truncated domains. They share the same parametric form with their non-truncated counterparts up to a normalizing constant. Since the computation of their normalizing constants is usually infeasible, Maximum Likelihood Estimation cannot be easily applied t…
The method approximates stationary distributions of Markov models by truncating irrelevant states.
Paper tackles overestimation bias in continuous control, improving performance by 25%.
We consider an appoximation of a catenoid constructed from "odd" truncated cones that maintains minimality in a certain sense. Thorough this procedure, we obtain a discrete curve approximating a catenary by exploiting the fact that it is the function that generates a catenoid. In this investigation, the theory of the G…
Estimates domain truncation error for option pricing PDEs.
Choppy optimizes ranked list truncation using Transformer architecture.
New COS method formula improves option pricing accuracy.
The generalized correlation approach, which has been successfully used in statistical radio physics to describe non-Gaussian random processes, is proposed to describe stochastic financial processes. The generalized correlation approach has been used to describe a non-Gaussian random walk with independent, identically d…
Lower bound shows super-polynomial gap for estimating truncated Gaussian means.