Local gaps in Ricci shrinkers depend only on dimension.
problem Understanding local properties of Ricci shrinkers.
method Proved local versions of Ricci curvature and entropy gap theorems.
result Local gaps depend only on dimension, not global entropy.
New method reduces regret and communication costs in federated Q-learning.
problem Worst-case regret and communication cost bounds in federated Q-learning.
method Gap-dependent analysis leveraging MDP structures.
result Achieves logT-type regret and communication cost bounds. Ray Singer torsion is a numerical invariant associated with a compact Riemannian manifold equipped with a flat bundle and a Hermitian structure on this bundle. In this note we show how one can remove the dependence on the Riemannian metric and on the Hermitian structure with the help of a base point and of an Euler str…
Algorithm removes leaves to find root in uniform trees.
problem Finding the root in large uniform attachment trees.
method Leaf-stripping algorithm recursively removes leaves.
result Set of remaining vertices contains the root with high probability.
Study on removing sets and uniqueness of diffusion operators on various spaces.
problem Determining the effect of removing small sets on the self-adjointness and uniqueness of diffusion operators.
method Analyzes symmetric diffusion operators on metric measure spaces, proving a truncation result for potentials.
result Characterizes the critical size of removed sets and their effect on operator properties.
Predictive models can fail to generalize from training to deployment environments because of dataset shift, posing a threat to model reliability and the safety of downstream decisions made in practice. Instead of using samples from the target distribution to reactively correct dataset shift, we use graphical knowledge …
Generative multitask learning mitigates confounders causing targets.
problem Unobserved confounders causing targets but not inputs.
method Generative multitask learning (GMTL) modifies inference objective to remove joint target influence.
result Improved robustness to target shift across multitask learning methods.
Importance of visual context in scene understanding tasks is well recognized in the computer vision community. However, to what extent the computer vision models for image classification and semantic segmentation are dependent on the context to make their predictions is unclear. A model overly relying on context will f…
The statistical dependencies which independent component analysis (ICA) cannot remove often provide rich information beyond the linear independent components. It would thus be very useful to estimate the dependency structure from data. While such models have been proposed, they usually concentrated on higher-order corr…
DaRE forests enable efficient data deletion from random forests.
problem Efficiently removing data from machine learning models.
method Random Forests with data deletion enabled (DaRE).
result Data deletion from DaRE models is orders of magnitude faster than retraining.
New SVM margin bound improves generalization in machine learning.
problem Improving SVM margin bounds for better generalization.
method Stable sample compression schemes to derive new data-dependent generalization bounds.
result Proves a new optimal SVM margin bound with a log factor improvement.
We investigate the problem of learning representations that are invariant to certain nuisance or sensitive factors of variation in the data while retaining as much of the remaining information as possible. Our model is based on a variational autoencoding architecture with priors that encourage independence between sens…
Procedure removes training data dependency from deep networks, improving generalization.
problem Removing dependency on training data in deep networks for better generalization.
method Deterministic and stochastic parts to ensure forgetting, leveraging activation and weight dynamics.
result New bound on information extraction from black-box networks, ensuring forgetting in activations.
A method removes treatment-covariate dependence for counterfactual prediction without adversarial training.
problem Counterfactual prediction under assignment bias.
method Information-theoretic approach learning a stochastic representation Z to minimize mutual information with outcomes.
result The method performs favorably in likelihood, counterfactual error, and policy evaluation compared to adversarial baselines.
Layer-wise Relevance Propagation (LRP) and saliency maps have been recently used to explain the predictions of Deep Learning models, specifically in the domain of text classification. Given different attribution-based explanations to highlight relevant words for a predicted class label, experiments based on word deleti…
UMFI improves feature importance methods by reducing runtime and enhancing performance.
problem Improving feature importance methods to better explain causal and associative relationships in data.
method Introducing UMFI, which uses dependence removal techniques from AI fairness literature.
result UMFI outperforms MCI, especially in complex data scenarios, and reduces runtime from exponential to super-linear.
We investigate the emergence of a structure in the correlation matrix of assets' returns as the time-horizon over which returns are computed increases from the minutes to the daily scale. We analyze data from different stock markets (New York, Paris, London, Milano) and with different methods. Result crucially depends …
Study allows removing data from machine learning models with strong guarantees.
problem Certifying removal of training data from machine learning models.
method Defined and developed a certified-removal mechanism for linear classifiers.
result Demonstrated that certified removal is possible and practical in certain learning settings.
NoiseRank reduces label noise without supervision, improving classification accuracy.
problem Label noise in datasets from noisy channels.
method NoiseRank uses Markov Random Fields to estimate and rank instances based on their noise probability.
result NoiseRank improves classification accuracy on noisy datasets.
While the authors of Batch Normalization (BN) identify and address an important problem involved in training deep networks-- Internal Covariate Shift-- the current solution has certain drawbacks. Specifically, BN depends on batch statistics for layerwise input normalization during training which makes the estimates of …
Paper removes sensitive data from IoT and Big Data for privacy.
problem Privacy concerns in IoT and Big Data.
method Develops new supervised and adversarial learning methods to remove sensitive data.
result Models maintain predictive model utility while making sensitive predictions ineffective.
We study the dependence structure of market states by estimating empirical pairwise copulas of daily stock returns. We consider both original returns, which exhibit time-varying trends and volatilities, as well as locally normalized ones, where the non-stationarity has been removed. The empirical pairwise copula for ea…
New method removes interference bias in causal models.
problem Interference bias impedes causal effect identification in real-world settings.
method Novel definition of causal models with local interference, semi-parametric assumptions.
result True Average Causal Effect can be identified in certain semi-parametric models with local interference.
We investigated financial market data to determine which factors affect information flow between stocks. Two factors, the time dependency and the degree of efficiency, were considered in the analysis of Korean, the Japanese, the Taiwanese, the Canadian, and US market data. We found that the frequency of the significant…
Unified framework for convergence of discrete diffusion models without state space size dependence.
problem Fundamental limitations in existing convergence theory for discrete diffusion models, especially under singular priors and large vocabularies.
method Unified adjoint-equation-based framework that establishes dimension-free convergence guarantees in any integral probability metric (IPM).
result First dimension-free convergence bounds applicable to both masked and uniform priors, free of state space size S. Selective removal of data subsets can efficiently unlearn unwanted distributions.
problem Efficiently removing unwanted data subsets without losing important information.
method Formalized as distributional unlearning, using Kullback-Leibler divergence constraints to select a small subset of data.
result Proposed method achieves corresponding log-loss bounds and is quadratically more sample-efficient than random removal.
We study the problem of non-explosion of diffusion processes on a manifold with time-dependent Riemannian metric. In particular we obtain that Brownian motion cannot explode in finite time if the metric evolves under backwards Ricci flow. Our result makes it possible to remove the assumption of non-explosion in the pat…
The paper shows robustness and generalization are closely connected via data-dependent bounds.
problem Connecting robustness and generalization in machine learning.
method Data-dependent generalization bounds that reduce dependence on covering number and hypothesis space.
result Proves robustness implies generalization, with near-exponential improvements in various situations.
Unified framework for explaining models by removing features.
problem Unclear relationships and preferences among existing model explanation methods.
method Removal-based explanations characterized by three dimensions.
result Unified framework unifies 26 existing methods.
UNHaP removes noise from physiological events using Hawkes processes.
problem Challenges in identifying true events from spurious ones in physiological signal analysis.
method UNHaP uses marked Hawkes processes to distinguish and unmix true events from noise.
result UNHaP significantly reduces false detection rates and enhances event understanding.
New method removes scalar curvature assumption in Ricci flow smoothing.
problem Uniform bounds on scalar curvature and other factors for Ricci flow.
method Quantitative short-time existence of Ricci flow without scalar curvature assumption.
result Ricci flow smoothing for measure space limits, Gromov-Hausdorff compactness, and topological rigidity results.
Private estimation with public data reduces sample complexity.
problem Estimating private distributions with limited public data.
method Differentially private estimation with public data under constraints of pure or concentrated DP.
result Public data can significantly reduce private sample complexity for estimation.
A new boosting model handles dependent censoring in time-to-event data.
problem Independent censoring assumption leads to biased predictions in time-to-event analysis.
method Clayton-boost, a boosting approach using Clayton copula.
result Clayton-boost outperforms other methods in handling dependent censoring.
The paper proves removable singularity for nonlocal minimal graphs.
problem Proving removable singularities for nonlocal minimal graphs.
method Analyzing (s,1)-capacity zero compact sets to ensure graphs are minimal in the entire domain. result Nonlocal minimal graphs are removable in the entire domain if they are minimal in a set of (s,1)-capacity zero. Unified framework for model explanation methods based on feature removal.
problem Unclear relationships and preferences among various model explanation methods.
method Characterizes removal-based explanations along three dimensions.
result Unified 26 existing methods, including widely used approaches.
This paper deals with multivariate regression chain graphs (MVR CGs), which were introduced by Cox and Wermuth [3,4] to represent linear causal models with correlated errors. We consider the PC-like algorithm for structure learning of MVR CGs, which is a constraint-based method proposed by Sonntag and Peña in [18]. We …
Removing spurious features can hurt model accuracy and disproportionately affect different groups.
problem Interference from spurious features in robust model performance across different groups.
method Characterization and analysis of spurious feature removal in noiseless overparameterized linear regression.
result Removal of spurious features can decrease accuracy and disproportionately affect different groups, even in balanced datasets.
New algorithm achieves asymptotically optimal regret without horizon dependence.
problem Horizon-free regret minimization for reinforcement learning.
method Proposes a new algorithm and proves a regret upper bound.
result Regret upper bound of \(\tilde O(\sqrt{SAK} + S^8A^3)\) with failure probability \(\delta\).
Removes singularities for Yang-Mills-Higgs fields in higher dimensions.
problem Yang-Mills-Higgs fields with isolated singularities.
method Establishes decay estimates and conformally invariant energy bounds.
result Removable singularity theorem for Yang-Mills-Higgs fields.
Tuning-free OR-PCA improves scalability for large datasets.
problem Dataset sensitivity of OR-PCA tuning parameters.
method Implicit regularization through modified gradient descents.
result Comparable or better performance on simulated and real-world datasets.
New algorithm for mean estimation in add-remove model achieves optimal error.
problem Mean estimation in add-remove model of differential privacy.
method Proposed new algorithm achieving min-max optimality.
result Achieves best possible constant in mean squared error for all ε.
Machine learning confound removal biases results, leading to misleading predictions.
problem Common confound removal methods in machine learning lead to misleading predictions.
method Featurewise removal of confound variance by linear regression before applying ML.
result This common deconfounding approach can leak information, amplifying null or moderate effects.
We call a singularity of a presymplectic form ω removable in its graph if its graph extends to a smooth Dirac structure over the singularity. An example for this is the symplectic form of a magnetic monopole. A criterion for the removability of singularities is given in terms of regularizing functions for pure spinor…
D3M debiases models by selectively removing problematic examples.
problem Model failures on underrepresented subgroups.
method Isolates and removes specific training examples that cause failures.
result Efficiently trains debiased classifiers with minimal example removal.
This paper proposes a simple approach to derive efficient error bounds for learning multiple components with sparsity-inducing regularization. We show that for such regularization schemes, known decompositions of the Rademacher complexity over the components can be used in a more efficient manner to result in tighter b…
A new method reduces CI tests for causal structure learning.
problem Exponential CI tests in constraint-based methods.
method Recursive Markov boundary-based approach.
result Significantly reduces CI tests compared to existing methods.
PUMA augments models to remove unique data points without performance loss.
problem Preserving model performance while removing unique training data points.
method Explicitly models data influence, reweights remaining data optimally.
result PUMA effectively removes unique data points without performance degradation.
DAC-SSM learns domain-agnostic states for better imitation learning.
problem Domain shifts hinder imitation learning in partially observable tasks.
method DAC-SSM uses adversarial training to remove domain-dependent information from states.
result DAC-SSM achieves comparable performance to experts in sparse reward tasks.