PUMA augments models to remove unique data points without performance loss.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A recently proposed clustering method, called the Nearest Descent (ND), can organize the whole dataset into a sparsely connected graph, called the In-tree. This ND-based Intree structure proves able to reveal the clustering structure underlying the dataset, except one imperfect place, that is, there are some undesired …
New framework for attributing online marketing touchpoints.
Proposes a new k-NN algorithm to improve classification accuracy by removing noise and pseudo-neighbours.
New fairness approach removes direct effects of unprivileged groups through causal regularization.
Selective removal of data subsets can efficiently unlearn unwanted distributions.
Study removes bias from chest X-ray embeddings using orthogonalization.
While great progress has been made recently in automatic image manipulation, it has been limited to object centric images like faces or structured scene datasets. In this work, we take a step towards general scene-level image editing by developing an automatic interaction-free object removal model. Our model learns to …
Machine learning confound removal biases results, leading to misleading predictions.
We describe a method for removing the effect of confounders in order to reconstruct a latent quantity of interest. The method, referred to as half-sibling regression, is inspired by recent work in causal inference using additive noise models. We provide a theoretical justification and illustrate the potential of the me…
We analyze the effect of adding, removing, and moving basepoints on link Floer homology. We prove that adding or removing basepoints via a procedure called quasi-stabilization is a natural operation on a certain version of link Floer homology, which we call . We consider the effect on the full link Flo…
Sources of variability in experimentally derived data include measurement error in addition to the physical phenomena of interest. This measurement error is a combination of systematic components, originating from the measuring instrument, and random measurement errors. Several novel biological technologies, such as ma…
Simple attack bypasses state-of-the-art DNN watermarking.
Unified theory explains how data augmentation improves deep learning models.
Improves GAN performance by identifying and removing harmful training instances.
Study on dropout in neural networks using percolation theory.
New framework removes harmful momentum effect for long-tailed classification.
An important problem in networked systems is detection and removal of suspected malicious nodes. A crucial consideration in such settings is the uncertainty endemic in detection, coupled with considerations of network connectivity, which impose indirect costs from mistakely removing benign nodes as well as failing to r…
Paper develops efficient mechanisms for estimating variance and covariance under differential privacy in the add-remove model.
Algorithm removes specific training data from models efficiently in high-dimensional settings.
Influence functions estimate the effect of removing a training point on a model without the need to retrain. They are based on a first-order Taylor approximation that is guaranteed to be accurate for sufficiently small changes to the model, and so are commonly used to study the effect of individual points in large data…
New method removes interference bias in causal models.
Study on removing sets and uniqueness of diffusion operators on various spaces.
Multiplicative noise, including dropout, is widely used to regularize deep neural networks (DNNs), and is shown to be effective in a wide range of architectures and tasks. From an information perspective, we consider injecting multiplicative noise into a DNN as training the network to solve the task with noisy informat…
A statistical framework for removing unwanted data domains in machine learning.
This paper describes a novel deep learning-based method for mitigating the effects of atmospheric distortion. We have built an end-to-end supervised convolutional neural network (CNN) to reconstruct turbulence-corrupted video sequence. Our framework has been developed on the residual learning concept, where the spatio-…
Unsupervised method removes satellite noise without paired data.
Algorithm removes backdoor watermarks from neural networks robustly.
A new unsupervised method removes CT metal artifacts using beta-CycleGAN and attention.
Matrix factorization is a simple and effective solution to the recommendation problem. It has been extensively employed in the industry and has attracted much attention from the academia. However, it is unclear what the low-dimensional matrices represent. We show that matrix factorization can actually be seen as simult…
Improved preterm prediction using synthetic EHG signals.
A new algorithm removes unexpected correlations in biased data for better clustering.
A denoising algorithm seeks to remove noise, errors, or perturbations from a signal. Extensive research has been devoted to this arena over the last several decades, and as a result, today's denoisers can effectively remove large amounts of additive white Gaussian noise. A compressed sensing (CS) reconstruction algorit…
Theorem proves minimal hypersurfaces in nonnegative scalar curvature manifolds are smooth.
Recent advances in Representation Learning and Adversarial Training seem to succeed in removing unwanted features from the learned representation. We show that demographic information of authors is encoded in -- and can be recovered from -- the intermediate representations learned by text-based neural classifiers. The …
Causalfe estimates treatment effects in panel data with fixed effects.
New method removes hidden confounders for unbiased treatment effect estimation.
NTL protects AI models by restricting their generalization ability to specific domains.
Pruning is a standard technique for removing unnecessary structure from a neural network to reduce its storage footprint, computational demands, or energy consumption. Pruning can reduce the parameter-counts of many state-of-the-art neural networks by an order of magnitude without compromising accuracy, meaning these n…
MAGIC method optimally estimates model predictions changes.
Bank behaviour is important for pricing XVA because it links different counterparties and thus breaks the usual XVA pricing assumption of counterparty independence. Consider a typical case of a bank hedging a client trade via a CCP. On client default the hedge (effects) will be removed (rebalanced). On the other hand, …
New method removes unwanted information from representations efficiently.
We investigate the problem of learning representations that are invariant to certain nuisance or sensitive factors of variation in the data while retaining as much of the remaining information as possible. Our model is based on a variational autoencoding architecture with priors that encourage independence between sens…
Predictive models learned from historical data are widely used to help companies and organizations make decisions. However, they may digitally unfairly treat unwanted groups, raising concerns about fairness and discrimination. In this paper, we study the fairness-aware ranking problem which aims to discover discriminat…
A new method detects outliers in dirty data using a leave-out strategy.
Cluster analysis and outlier detection are strongly coupled tasks in data mining area. Cluster structure can be easily destroyed by few outliers; on the contrary, outliers are defined by the concept of cluster, which are recognized as the points belonging to none of the clusters. Unfortunately, most existing studies do…
Good data stewardship requires removal of data at the request of the data's owner. This raises the question if and how a trained machine-learning model, which implicitly stores information about its training data, should be affected by such a removal request. Is it possible to "remove" data from a machine-learning mode…
We consider the non-parametric regression problem under Huber's -contamination model, in which an fraction of observations are subject to arbitrary adversarial noise. We first show that a simple local binning median step can effectively remove the adversary noise and this median estimator is minimax optimal up t…