We design a new algorithm for the Euclidean -means problem that operates in the local model of differential privacy. Unlike in the non-private literature, differentially private algorithms for the -means objective incur both additive and multiplicative errors. Our algorithm significantly reduces the additive erro…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A fast Monte Carlo method for additive processes and option pricing.
Multiplicative noise models are often used instead of additive noise models in cases in which the noise variance depends on the state. Furthermore, when Poisson distributions with relatively small counts are approximated with normal distributions, multiplicative noise approximations are straightforward to implement. Th…
Passive investing can incur hidden costs due to market timing inefficiencies.
Inexact subgradient methods work well for semialgebraic functions with additive errors.
First order discretizations of Langevin diffusion can achieve better generalization error with additional smoothness assumptions.
Efficient classifier error estimation without re-training.
Bayesian Additive Regression Trees (BART) is a fully Bayesian approach to modeling with ensembles of trees. BART can uncover complex regression functions with high dimensional regressors in a fairly automatic way and provide Bayesian quantification of the uncertainty through the posterior. However, BART assumes IID nor…
LSTM Networks accurately forecast COVID-19 cases in Turkey with lower error than other methods.
The paper develops a minimax optimal method for high-dimensional regression using auxiliary data.
One-pass algorithm finds small subset for subspace approximation with additive error.
Tensors are becoming prevalent in modern applications such as medical imaging and digital marketing. In this paper, we propose a sparse tensor additive regression (STAR) that models a scalar response as a flexible nonparametric function of tensor covariates. The proposed model effectively exploits the sparse and low-ra…
This work studies scaling laws for low-precision training in high-dimensional linear regression.
Gradient-free optimization for additive models achieves optimal error.
HARFE approximates sparse additive functions using random features and ridge regression.
Study high-dimensional logistic regression with missing data, providing exact error characterizations.
New estimators improve sparse semiparametric additive modeling.
This paper examines fundamental error characteristics for a general class of matrix completion problems, where the matrix of interest is a product of two a priori unknown matrices, one of which is sparse, and the observations are noisy. Our main contributions come in the form of minimax lower bounds for the expected pe…
This work finds a point with small test error in polynomial time for mildly overparameterized neural nets.
Bayesian Additive Regression Networks use neural networks for regression tasks.
Robust variable selection for high-dimensional data with missing and measurement errors.
The paper introduces a new model to correct bias in treatment effect estimates due to sample selection.
We consider the problem of estimating the class prior in an unlabeled dataset. Under the assumption that an additional labeled dataset is available, the class prior can be estimated by fitting a mixture of class-wise data distributions to the unlabeled data distribution. However, in practice, such an additional labeled…
Paper develops Euler scheme for fractional delay diff. eqs with additive noise.
We consider general non-Euclidean distance measures between real world objects that need to be classified. It is assumed that objects are represented by distances to other objects only. Conditions for zero-error dissimilarity based classifiers are derived. Additional conditions are given under which the zero-error deci…
The main aim of this paper is to provide an analysis of gradient descent (GD) algorithms with gradient errors that do not necessarily vanish, asymptotically. In particular, sufficient conditions are presented for both stability (almost sure boundedness of the iterates) and convergence of GD with bounded, (possibly) non…
Gaussian graphical model is a graphical representation of the dependence structure for a Gaussian random vector. It is recognized as a powerful tool in different applied fields such as bioinformatics, error-control codes, speech language, information retrieval and others. Gaussian graphical model selection is a statist…
In this paper, we obtain asymptotic formulas with error estimates for the implied volatility associated with a European call pricing function. We show that these formulas imply Lee's moment formulas for the implied volatility and the tail-wing formulas due to Benaim and Friz. In addition, we analyze Pareto-type tails o…
Study on natural actor-critic for POMDPs with finite memory.
Study selective classification with halfspaces, achieving error bounds under Gaussian distributions.
Method detects errors in numerical data using regression models.
TCE measures calibration error with a test-based approach.
This paper improves entropy bounds for ranking time-series complexity.
The paper analyzes CycleGAN's error components for unpaired data generation.
Algorithm learns decision trees from noisy data.
Uniform stability of a learning algorithm is a classical notion of algorithmic stability introduced to derive high-probability bounds on the generalization error (Bousquet and Elisseeff, 2002). Specifically, for a loss function with range bounded in , the generalization error of a -uniformly stable learning a…
In healthcare applications, predictive uncertainty has been used to assess predictive accuracy. In this paper, we demonstrate that predictive uncertainty estimated by the current methods does not highly correlate with prediction error by decomposing the latter into random and systematic errors, and showing that the for…
Proposes a model combining difference-attention and error-correction LSTMs for improved time series prediction.
The sigma-point filters, such as the UKF, which exploit numerical quadrature to obtain an additional order of accuracy in the moment transformation step, are popular alternatives to the ubiquitous EKF. The classical quadrature rules used in the sigma-point filters are motivated via polynomial approximation of the integ…
Improved sampling in generative models using CLDs with a hyperparameter.
We present a nonparametric method for selecting informative features in high-dimensional clustering problems. We start with a screening step that uses a test for multimodality. Then we apply kernel density estimation and mode clustering to the selected features. The output of the method consists of a list of relevant f…
In this paper, we prove that some Gaussian structural equation models with dependent errors having equal variances are identifiable from their corresponding Gaussian distributions. Specifically, we prove identifiability for the Gaussian structural equation models that can be represented as Andersson-Madigan-Perlman cha…
In this paper, we study the problem of approximately computing the product of two real matrices. In particular, we analyze a dimensionality-reduction-based approximation algorithm due to Sarlos [1], introducing the notion of nuclear rank as the ratio of the nuclear norm over the spectral norm. The presented bound has i…
In the regression setting, given a set of hyper-parameters, a model-estimation procedure constructs a model from training data. The optimal hyper-parameters that minimize generalization error of the model are usually unknown. In practice they are often estimated using split-sample validation. Up to now, there is an ope…
We present a new method for high-dimensional linear regression when a scale parameter of the additive errors is unknown. The proposed estimator is based on a penalized Huber -estimator, for which theoretical results on estimation error have recently been proposed in high-dimensional statistics literature. However, t…
A new sampling method called Restart improves both speed and quality of generative processes.
Sources of variability in experimentally derived data include measurement error in addition to the physical phenomena of interest. This measurement error is a combination of systematic components, originating from the measuring instrument, and random measurement errors. Several novel biological technologies, such as ma…
Estimates shared linear subspace from noisy data with multiple users.