Many real datasets contain values missing not at random (MNAR). In this scenario, investigators often perform list-wise deletion, or delete samples with any missing values, before applying causal discovery algorithms. List-wise deletion is a sound and general strategy when paired with algorithms such as FCI and RFCI, b…
The paper compares two methods for handling missing data in causal discovery.
problem Handling missing data in causal discovery algorithms.
method Test-wise deletion and multiple imputation.
result Multiple imputation is more challenging for causal discovery than for estimation.
New approach protects privacy of deleted records in machine learning.
problem Privacy of deleted records in machine learning models.
method Sound deletion guarantee and noisy gradient descent algorithm.
result Privacy of existing records is necessary for deleted records' privacy.
Paper tackles adaptive deletion of data points from trained models.
problem Adaptive deletion of data points from trained models.
method Reduction from adaptive to non-adaptive deletion guarantees using differential privacy and max information.
result Strong provable deletion guarantees for adaptive deletion sequences.
This research tackles data deletion in linear regression with noisy SGD, finding perfect deleted points.
problem Finding points to delete from a dataset without significantly affecting the training result.
method Signal-to-noise ratio and an algorithm based on it.
result The perfect deleted point is crucial for maintaining model performance and privacy budget.
We propose the Insertion-Deletion Transformer, a novel transformer-based neural architecture and training method for sequence generation. The model consists of two phases that are executed iteratively, 1) an insertion phase and 2) a deletion phase. The insertion phase parameterizes a distribution of insertions on the c…
Efficient algorithms for deleting data from machine learning models without significantly affecting performance.
problem Deleting data from machine learning models while maintaining performance.
method Leveraging convex optimization and reservoir sampling, the paper introduces algorithms for handling long sequences of adversarial updates.
result First data deletion algorithms that promise steady-state error not growing with the length of the update sequence.
Paper proposes a fast method for approximate data deletion in generative models.
problem Efficient data deletion in unsupervised learning models is an open problem.
method Density-ratio-based framework for generative models, fast method for approximate data deletion, statistical test.
result Theoretical guarantees and empirical demonstrations of the proposed methods across various generative models.
New method for efficiently deleting data from ML models.
problem Efficiently removing data from trained ML models without retraining.
method Approximate deletion method for linear and logistic models.
result Significantly faster than existing methods, with linear time dependence on feature dimension.
Study on deleting user data in linear regression models to maintain limited memory.
problem Deleting user data in a limited time frame for statistical models.
method Proposed FIFD-OLS and FIFD-Adaptive Ridge algorithms for low-dimensional and online settings.
result Demonstrated effectiveness of FIFD-Adaptive Ridge in maintaining statistical efficiency.
The paper develops algorithms to find a robust summary of data under deletion, achieving good approximation guarantees.
problem Finding a summary of data that remains valuable even after some elements are deleted.
method Constant-factor approximation algorithms for deletion robust submodular maximization under matroid constraints.
result The algorithms provide good approximation guarantees for both centralized and streaming settings.
New examples show deletion type admissible pairs can be rigid under rational saturation.
problem Rigidity of admissible pairs of rational homogeneous spaces of Picard number one.
method Application of Mok's general criterion for non-subdiagram type admissible pairs.
result Examples of deletion type admissible pairs are rigid under rational saturation.
ID-ExpO fine-tunes neural networks for more faithful explanations.
problem Improving the faithfulness of explanations for complex machine learning models.
method Differentiable insertion/deletion metric-aware regularizers for optimization.
result Fine-tuned predictors produce more faithful explanations.
We observe the effects of the three different events that cause spread changes in the order book, namely trades, deletions and placement of limit orders. By looking at the frequencies of the relative amounts of price changing events, we discover that deletions of orders open the bid-ask spread of a stock more often tha…
We show that deleting an edge of a 3-cycle in an intrinsically knotted graph gives an intrinsically linked graph.
Intense recent discussions have focused on how to provide individuals with control over when their data can and cannot be used --- the EU's Right To Be Forgotten regulation is an example of this effort. In this paper we initiate a framework studying what to do when it is no longer permissible to deploy models derivativ…
DaRE forests enable efficient data deletion from random forests.
problem Efficiently removing data from machine learning models.
method Random Forests with data deletion enabled (DaRE).
result Data deletion from DaRE models is orders of magnitude faster than retraining.
The paper tackles robust submodular maximization under matroid constraints, providing approximation algorithms for summary extraction.
problem Maximizing submodular functions while ensuring high value even after deletions.
method Constant-factor approximation algorithms for centralized and streaming settings, considering both non-monotone and monotone objectives.
result Approximation algorithms with space complexity depending on matroid rank and deleted elements, achieving improved factors in monotone cases.
We use a variation on the commutator collection process to characterize those pure braids which become trivial when any one strand is deleted, or, more generally, those pure braids which become trivial when all the strands in any one of a list of sets of strands is deleted.
RS-Del provides robustness for sequence classifiers against edit distance attacks.
problem Certifying robustness of discrete sequence classifiers against edit distance attacks.
method Randomized deletion (RS-Del) for discrete sequence classifiers, focusing on edit distance-bounded adversaries.
result Achieved a certified accuracy of 91% at an edit distance radius of 128 bytes on malware detection.
Graph pruning improves neural network performance by addressing squashing and smoothing issues.
problem Over-squashing and over-smoothing in Graph Neural Networks.
method Proposes edge deletions to simultaneously address over-squashing and over-smoothing, optimizing spectral gap.
result Edge deletions improve generalization and distinguishability of nodes of different classes.
Gordon and Litherland showed that all compact, unoriented, possibly non-orientable surfaces in S3 bounded by a link are realted by attaching/deleting tubes and half twisted bands. In this note we give an elementary proof for this result.
The configuration space F2(M) of ordered pairs of distinct points in a manifold M, also known as the deleted square of M, is not a homotopy invariant of M: Longoni and Salvatore produced examples of homotopy equivalent lens spaces M and N of dimension three for which F2(M) and F2(N) are not homoto…
Proposes a new jackknife method for time series hyperparameter selection.
problem Hyperparameter selection for time series models.
method Artificial delete-d jackknife approach.
result Asymptotic and finite-sample advantages demonstrated.
We propose a framework for verifying data deletion in MLaaS systems.
problem Ensuring compliance with data deletion requests in MLaaS systems.
method Formal framework based on hypothesis testing, novel backdoor-based verification mechanism.
result Demonstrated high confidence in certifying data deletion with minimal impact on ML service accuracy.
New graph shows edge deletion/contraction doesn't always result in intrinsically linked graphs.
problem Edge operations in intrinsically knotted graphs don't always produce intrinsically linked graphs.
method Presented a new intrinsically knotted graph.
result Edge operations in intrinsically knotted graphs don't always result in intrinsically linked graphs.
Recently enacted legislation grants individuals certain rights to decide in what fashion their personal data may be used, and in particular a "right to be forgotten". This poses a challenge to machine learning: how to proceed when an individual retracts permission to use data which has been part of the training process…
New algorithms delete user data from machine learning models efficiently.
problem Deleting user data from machine learning models trained with empirical risk minimization.
method Developed an online unlearning algorithm using the infinitesimal jackknife, targeting non-smooth regularizers.
result Empirically improved runtime while maintaining memory requirements and test accuracy.
Second-order optimizers retain residual information after data deletion, affecting machine unlearning.
problem Residual information in second-order optimizers after data deletion.
method Comparison of first-order and second-order learners, eigendecomposition analysis.
result Second-order optimizers retain residual information, not detectable by first-order analysis.
We investigate the problem of reliable communication between two legitimate parties over deletion channels under an active eavesdropping (aka jamming) adversarial model. To this goal, we develop a theoretical framework based on probabilistic finite-state automata to define novel encoding and decoding schemes that ensur…
Efficiently adds or deletes data in GBDT models.
problem Traditional GBDT training requires all data to be accessed simultaneously, limiting add/delete operations.
method Proposes an online learning framework for GBDT supporting incremental and decremental learning.
result First work to unify incremental and decremental learning on GBDT in-place.
Applications in machine learning, optimization, and control require the sequential selection of a few system elements, such as sensors, data, or actuators, to optimize the system performance across multiple time steps. However, in failure-prone and adversarial environments, sensors get attacked, data get deleted, and a…
New algorithms reduce matching market regret to log(T) with improved stability.
problem Minimizing regret in two-sided matching markets with bandit feedback.
method Phase-based algorithm with local arm deletion to improve stability.
result Achieves Θ(log(T)) regret for markets with uniqueness consistency.
We introduce agents that use object-oriented reasoning to consider alternate states of the world in order to more quickly find solutions to problems. Specifically, a hierarchical controller directs a low-level agent to behave as if objects in the scene were added, deleted, or modified. The actions taken by the controll…
New framework for consistent submodular maximization with insertions and deletions.
problem Maintaining near-optimal solutions in a dynamic setting with insertions and deletions.
method Developed a general framework for fully dynamic submodular maximization, instantiated for cardinality and rank-k matroid constraints.
result First constant-factor approximations with sublinear consistency for both cardinality and rank-k matroid constraints.
Bayesian models can be tricked into believing false data.
problem Vulnerability of Bayesian inference to data poisoning attacks.
method Developed attacks to manipulate Bayesian posterior through deletion and replication of data.
result Demonstrated that Bayesian inference can be steered to target distributions.
New algorithm for maximizing submodular functions in real-time data changes.
problem Maximizing submodular functions under dynamic constraints.
method Randomized algorithm with O(k2) amortized update time. result 4-approximate solution to submodular maximization problem.
We propose an approach for approximating the partition function which is based on two steps: (1) computing the partition function of a simplified model which is obtained by deleting model edges, and (2) rectifying the result by applying an edge-by-edge correction. The approach leads to an intuitive framework in which o…
Researchers develop methods to prevent GANs from generating certain types of images.
problem Pre-trained GANs sometimes produce undesirable samples.
method Post-editing GANs to prevent certain types of outputs, using three algorithms.
result Our algorithms effectively prevent GANs from generating certain types of images while maintaining high quality.
New formulas for feature importance tests in regression models.
problem Identifying important features in regression models.
method Established formulas for AUC criteria and proposed alternative metrics.
result Integrated Gradients (IG) performs nearly as well as Kernel SHAP (KS) but is faster.
Paper proposes first unlearning algorithm for MCMC models.
problem Enforcing right to be forgotten in AI causes high costs for data deletion.
method Converts MCMC unlearning to explicit optimization problem, designs MCMC influence function.
result MCMC unlearning does not compromise generalizability of models.
A well-known problem in data science and machine learning is {\em linear regression}, which is recently extended to dynamic graphs. Existing exact algorithms for updating the solution of dynamic graph regression require at least a linear time (in terms of n: the size of the graph). However, this time complexity might…
Data ownership and data protection are increasingly important topics with ethical and legal implications, e.g., with the right to erasure established in the European General Data Protection Regulation (GDPR). In this light, we investigate network embeddings, i.e., the representation of network nodes as low-dimensional …
A new algorithm for competing agents in a two-sided market setting.
problem Decentralized competition between agents in a two-sided market with unknown valuations.
method UCB-D3 algorithm for UCB with Decentralized Dominant-arm Deletion.
result UCB-D3 is order optimal and achieves a new regret lower bound.
XGES improves GES by favoring early edge deletion, outperforming GES in finite data settings.
problem Learning directed acyclic graphs from finite data.
method Extremely Greedy Equivalent Search (XGES) improves GES by favoring early edge deletion.
result XGES consistently outperforms GES in recovering the correct graphs, and is 10 times faster.
Graph Neural Networks (GNNs) have boosted the performance of many graph related tasks such as node classification and graph classification. Recent researches show that graph neural networks are vulnerable to adversarial attacks, which deliberately add carefully created unnoticeable perturbation to the graph structure. …
Study surfaces with free product fundamental groups, proving existence and properties.
problem Understanding fundamental groups of quasi-projective surfaces.
method Prove existence of admissible maps and use addition-deletion Lemmas.
result Existence of admissible maps and properties of fundamental groups.
Motivated by Tverberg-type problems in topological combinatorics and by classical results about embeddings (maps without double points), we study the question whether a finite simplicial complex K can be mapped into R^d without higher-multiplicity intersections. We focus on conditions for the existence of almost r-embe…