Paper tackles label flipping attacks on machine learning models.
problem Label flipping attacks can degrade machine learning model performance.
method Proposes an algorithm for optimal label flipping attacks and a detection mechanism.
result Demonstrates effectiveness of proposed detection and relabeling mechanism.
We compare the sample complexity of private learning [Kasiviswanathan et al. 2008] and sanitization~[Blum et al. 2008] under pure ε-differential privacy [Dwork et al. TCC 2006] and approximate (ε,δ)-differential privacy [Dwork et al. Eurocrypt 2006]. We show that the sample complexity of these tasks under approxima…
Method sanitizes IFM in CNN layers to control privacy loss.
problem Controlling privacy loss in CNNs using input feature maps.
method Sample-and-hold approximation scheme to sanitize IFM, unfolding tensors for independence from CNN configuration.
result Control the privacy loss by adjusting the sanitization degree.
GS-WGAN sanitizes sensitive data for machine learning with improved privacy and model quality.
problem Lack of privacy in sensitive data hinders machine learning applications.
method Gradient-sanitized Wasserstein Generative Adversarial Networks (GS-WGAN).
result GS-WGAN generates more informative samples and outperforms state-of-the-art approaches.
New attacks bypass data sanitization defenses, increasing model errors.
problem Data poisoning attacks corrupt machine learning models trained on external data.
method Developed three attacks that coordinate poisoned points and formulate as optimization problems.
result 3% poisoned data increases test error from 3% to 24% on Enron spam detection.
GANsan removes sensitive attributes from data to prevent discrimination.
problem Preventing discrimination in automated decision processes.
method Generative adversarial networks (GANs) to modify attributes without losing interpretability.
result Demonstrated effectiveness and trade-off between fairness and utility on real data.
Researchers found PP-GANs can hide sensitive data in sanitized images, undermining privacy checks.
problem Lack of formal proofs of privacy in PP-GANs for image sanitization.
method Subverted PP-GANs for facial expression recognition to hide sensitive data in sanitized images.
result It is possible to hide sensitive identification data in sanitized PP-GAN output images, even allowing reconstruction of entire input images.
KNG mechanism provides sanitized statistical summaries with strong privacy and utility guarantees.
problem Producing sanitized statistical summaries with differential privacy.
method Promotes summaries that minimize an objective function by weighting gradients, achieving utility similar to objective perturbation but with stronger privacy guarantees.
result KNG's noise is asymptotically negligible compared to statistical error for many problems.
Linear filtration helps delete training data from models.
problem Deleting training data from models when individuals request it.
method Linear filtration as a computationally efficient sanitization method.
result Demonstrates benefits in an adversarial setting over naive deletion schemes.
New algorithm for private minimum spanning tree release with improved accuracy.
problem Privacy-preserving minimum spanning tree release for graphs.
method Formal differential privacy definition for graphs, new MST algorithm, combining sanitizing mechanism and MST clustering.
result Improved accuracy in weight approximation compared to state of the art.
Study shows removing outliers from training sets improves model robustness.
problem Vulnerability of deep neural networks to adversarial examples.
method Proposed a framework to detect and remove outliers from the training set to improve model robustness.
result Demonstrated that removing outliers from the training set can enhance model robustness.
A new HAR algorithm uses U-Net for pixel-level gesture recognition.
problem Multi-class window problem in traditional HAR methods.
method U-Net network for activity labeling and prediction at each sampling point.
result Highest accuracy and F1-score compared to other methods.
Privacy-preserving method protects user speech data from cloud services.
problem Privacy compromise in cloud-based speech analysis.
method Collects and sanitizes speech data before sharing, using transformation functions and voice conversion.
result Identification of sensitive emotional state reduced by ~96%.
Deep learning improves gait biometric recognition accuracy.
problem Low accuracy in existing CSI-based gait identification systems.
method Developed an end-to-end deep CSI learning system using deep neural networks.
result Achieved a top-1 accuracy of 97.12% for a dataset of 30 people.
A new privacy-preserving mechanism for shapes on manifolds.
problem Privacy-preserving sanitization of shapes on curved manifolds.
method Developed a K-norm gradient mechanism on Riemannian manifolds.
result The K-norm gradient mechanism offers better control over sensitivity than the Laplace mechanism on positively curved manifolds.
This paper defends SVMs against poisoning attacks using DBSCAN and hardness proofs.
problem Adversarial injection of specially crafted samples into training data to misclassify SVMs.
method Two strategies: robust SVM algorithms and data sanitization (DBSCAN).
result Proves hardness of simple SVM problem and effectiveness of DBSCAN for poisoning attacks.
Paper proposes detecting video manipulation using stream descriptors.
problem Misuse of manipulated video content.
method Binary classifiers on multimedia stream descriptors.
result Scalable approach can detect high-quality manipulations.
Extends differential privacy to Riemannian manifolds, improving utility.
problem Releasing private statistical summaries on Riemannian manifolds.
method Extended Laplace or K-norm mechanism using intrinsic distances and volumes.
result Demonstrates rate optimality and utility improvement over ambient spaces.
Deep learning method segments and monitors slums from satellite imagery.
problem Slum rehabilitation and improvement in developing countries.
method Regional convolutional neural networks for instance segmentation using transfer learning.
result Maximum AP of 80.0 for slum shape and appearance learning.
Kernel analysis reveals rumor truth from diffusion patterns alone.
problem Detecting unverified rumors on Twitter using text and user identities.
method Graph kernels to extract diffusion patterns from Twitter cascade structures.
result Diffusion patterns are highly informative of rumor truth or falsehood.
Language models learn from training data and can leak private information.
problem Language models lack context understanding and can expose private data.
method Discussing the limitations of current privacy protection methods for language models.
result Existing privacy protection methods are insufficient for language models.
Study compares RL and DT-based control for hedging European call options.
problem Optimizing hedging strategies for European call options with transaction costs.
method Reinforcement Learning vs. Deep Trajectory-based Stochastic Control.
result RL and DT-based methods perform differently under stepwise mean-variance hedging.
Three new oracle-efficient algorithms for private synthetic data release.
problem Constructing private synthetic data that preserves statistical query answers.
method Oracle-efficient algorithms using optimization oracles for differential privacy.
result Better accuracy in large workload and high privacy regime compared to state-of-the-art.
HLTF generates chemically valid 3D molecules with improved topology control.
problem Generating chemically valid 3D molecules is challenging due to bond topology errors.
method HLTF uses a latent multi-scale plan for global context and a constraint-aware sampler to suppress topology-driven failures.
result HLTF achieves high validity and uniqueness on QM9 and GEOM-DRUGS datasets.
A new method to protect enterprise data privacy in AI models.
problem Enterprise data leakage risks in AI models.
method ABack, a training-free mechanism using Hidden State Model.
result Improves privacy utility by up to 15% over strong baselines.
Researchers developed a differentially private method for computing Wasserstein distances.
problem Computing divergences between distributions while preserving privacy.
method They focused on the Sliced Wasserstein Distance and added Gaussian perturbations to make it differentially private.
result They introduced a new differentially private distance, the Smoothed Sliced Wasserstein Distance, which performs well in generative models and domain adaptation.
Develops user-controlled privacy collaboration between users and data providers.
problem Ensuring data privacy while maintaining utility for users and providers.
method Collaborative learning of a sensitization function to control data sharing and privacy.
result Maintains data utility while fully protecting private information through user-controlled privacy settings.
Modeling influenza spread using feature engineering and international flow deconvolution.
problem Predicting and mitigating influenza spread through feature extraction and international flow analysis.
method Discrete Fourier Transform, matrix completion, SVM, autoencoders, PCA, deconvolution of international flow.
result Significant environmental and economic features are crucial to influenza mortality.
Study maps interdependence of SDGs, finds complex, dynamic linkages.
problem Identify which SDGs promote progress and how quickly.
method Used a balanced panel of 114 countries from 2000 to 2024, applying two estimators to recover directed interaction network and measure dynamic linkages.
result 84 goal linkages survive false-discovery control, showing both synergies and trade-offs, with no single goal acting as a universal accelerator.
Proposes a new model for noisy labels considering multiple labelers and adversarial attacks.
problem Real-world noisy label models with multiple labelers and adversarial attacks.
method Labeler-dependent noise model with adversarial attack vectors.
result State-of-the-art approaches for learning from noisy labels are defeated by adversarial label attacks.
Label smoothing improves model performance even with noisy labels.
problem Mitigating label noise in deep learning models.
method Examined label smoothing as a technique to cope with label noise and compared it to loss-correction methods.
result Label smoothing is competitive with loss-correction techniques under label noise and beneficial for distillation from noisy data.
Paper proposes a method to recover accurate labels from partially valid data in multi-label learning.
problem Tackles noisy supervision in multi-label learning with partially valid labels.
method Develops a two-stage method that estimates label enrichment and ground-truth confidences.
result Demonstrates improved performance over state-of-the-art PML methods.
CbMLC improves multi-label classification with noisy labels.
problem Evaluating multi-label classifiers with noisy labels.
method Context-Based Multi-Label Classifier (CbMLC) that handles noisy labels without additional supervision.
result CbMLC yields substantial improvements over previous methods in noisy label settings.
LNEMLC embeds label network for multi-label classification.
problem Lack of effective adaptation and preservation of generalization abilities for unseen label combinations.
method LNEMLC embeds label network to extend input space for any base multi-label classifier.
result Statistically significant improvements over simple kNN baseline classifier.
Proposes ML-GCN for multi-label network node representation learning.
problem Complex multi-label networks with correlated labels.
method Two Siamese GCNs model node-label and label-label interactions, integrated under a unified objective function.
result Effective node representation learning with preserved label interactions.
Proposes MGPLL for PL learning with non-random noise.
problem Partial label learning with non-random label noise.
method Bi-directional mapping framework, conditional noise label generation, multi-class predictor, adversarial learning.
result Demonstrates state-of-the-art performance in partial label learning.
Logistic regression can handle noisy labels effectively when labels are imperfectly assigned by multiple experts.
problem Label noise in supervised classification due to manual labelling by multiple experts.
method Using approximate posterior probabilities of class membership from multiple experts to train logistic regression models.
result Logistic regression can be robust to label noise when classification difficulty is the only source of errors.
A new method learns label correlations for better multi-label predictions.
problem Label correlations not accurately characterized by existing approaches.
method Sparse reconstruction in the label space to learn correlations, then integrate into model training.
result Our approach outperforms state-of-the-art multi-label learning methods.
Paper tackles multi-label zero-shot learning, improving label embedding projection for unseen classes.
problem Challenges in transferring knowledge from seen to unseen classes in multi-label zero-shot learning.
method Proposes a transfer-aware embedding projection approach to project label embeddings into a low-dimensional space for better inter-label relationships and explicit information transfer.
result Demonstrates the efficacy of the proposed approach through experiments on zero-shot multi-label image classification.
FLAME auto-labels mobile data efficiently on diverse processors.
problem Accurately and efficiently labeling mobile data with unknown labels on heterogeneous processors.
method Self-adaptive auto-labeling system Flame that schedules and executes workloads on mobile processors.
result Flame achieves high labeling accuracy and performance on heterogeneous mobile processors.
PML-LFC improves PML by estimating label confidence from both feature and label spaces.
problem PML challenges in real-world scenarios where only some labels are relevant.
method PML-LFC estimates label confidence using feature and label space similarities, training a predictor with these values.
result PML-LFC achieves superior performance on synthetic and real-world datasets.
Proposes methods to improve multi-label learning by addressing local label imbalance.
problem Local label imbalance within minority class examples degrades multi-label learning performance.
method Introduces a measure to assess local label imbalance and two sampling approaches (MLSOL, MLUL) to address it.
result Experimental results show MLSOL and MLUL improve performance on multi-label datasets.
An active learner is given a hypothesis class, a large set of unlabeled examples and the ability to interactively query labels to an oracle of a subset of these examples; the goal of the learner is to learn a hypothesis in the class that fits the data well by making as few label queries as possible. This work addresses…
LaMP neural networks model label interactions for multi-label classification.
problem Efficiently modeling label interactions in multi-label classification.
method Label Message Passing (LaMP) Neural Networks, treating labels as nodes on a graph, compute hidden representations conditioned on input using attention-based message passing.
result Significantly outperforms state-of-the-art multi-label classification models on seven real-world datasets.
Paper tackles label insufficiency and inaccuracy in semi-supervised learning.
problem Label insufficiency and inaccuracy in semi-supervised learning.
method Graph-based propagation for label insufficiency and label filtering for inaccuracy.
result SIIS improves performance in the presence of label noise and scarcity.
An important problem in multi-label classification is to capture label patterns or underlying structures that have an impact on such patterns. This paper addresses one such problem, namely how to exploit hierarchical structures over labels. We present a novel method to learn vector representations of a label space give…
New algorithm for XMC from aggregated labels.
problem Finding relevant labels for inputs from a large label universe.
method Developed a scalable algorithm to impute individual labels from group labels.
result Advantages over existing approaches in XMC and MIML tasks.
Enhances labels from unlabeled data using sample correlations.
problem Lack of label distributions in real-world applications.
method Proposes LESC and gLESC methods to enhance label distributions.
result Improves performance of label enhancement through sample correlations.