Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

1122 · Jan 201819922001200920182026
29 results for Sanitization

Method sanitizes IFM in CNN layers to control privacy loss.

problem Controlling privacy loss in CNNs using input feature maps.
method Sample-and-hold approximation scheme to sanitize IFM, unfolding tensors for independence from CNN configuration.
result Control the privacy loss by adjusting the sanitization degree.

GS-WGAN sanitizes sensitive data for machine learning with improved privacy and model quality.

problem Lack of privacy in sensitive data hinders machine learning applications.
method Gradient-sanitized Wasserstein Generative Adversarial Networks (GS-WGAN).
result GS-WGAN generates more informative samples and outperforms state-of-the-art approaches.

New attacks bypass data sanitization defenses, increasing model errors.

problem Data poisoning attacks corrupt machine learning models trained on external data.
method Developed three attacks that coordinate poisoned points and formulate as optimization problems.
result 3% poisoned data increases test error from 3% to 24% on Enron spam detection.

Researchers found PP-GANs can hide sensitive data in sanitized images, undermining privacy checks.

problem Lack of formal proofs of privacy in PP-GANs for image sanitization.
method Subverted PP-GANs for facial expression recognition to hide sensitive data in sanitized images.
result It is possible to hide sensitive identification data in sanitized PP-GAN output images, even allowing reconstruction of entire input images.

KNG mechanism provides sanitized statistical summaries with strong privacy and utility guarantees.

problem Producing sanitized statistical summaries with differential privacy.
method Promotes summaries that minimize an objective function by weighting gradients, achieving utility similar to objective perturbation but with stronger privacy guarantees.
result KNG's noise is asymptotically negligible compared to statistical error for many problems.

New algorithm for private minimum spanning tree release with improved accuracy.

problem Privacy-preserving minimum spanning tree release for graphs.
method Formal differential privacy definition for graphs, new MST algorithm, combining sanitizing mechanism and MST clustering.
result Improved accuracy in weight approximation compared to state of the art.

Study shows removing outliers from training sets improves model robustness.

problem Vulnerability of deep neural networks to adversarial examples.
method Proposed a framework to detect and remove outliers from the training set to improve model robustness.
result Demonstrated that removing outliers from the training set can enhance model robustness.

This paper defends SVMs against poisoning attacks using DBSCAN and hardness proofs.

problem Adversarial injection of specially crafted samples into training data to misclassify SVMs.
method Two strategies: robust SVM algorithms and data sanitization (DBSCAN).
result Proves hardness of simple SVM problem and effectiveness of DBSCAN for poisoning attacks.

Study compares RL and DT-based control for hedging European call options.

problem Optimizing hedging strategies for European call options with transaction costs.
method Reinforcement Learning vs. Deep Trajectory-based Stochastic Control.
result RL and DT-based methods perform differently under stepwise mean-variance hedging.

Three new oracle-efficient algorithms for private synthetic data release.

problem Constructing private synthetic data that preserves statistical query answers.
method Oracle-efficient algorithms using optimization oracles for differential privacy.
result Better accuracy in large workload and high privacy regime compared to state-of-the-art.

HLTF generates chemically valid 3D molecules with improved topology control.

problem Generating chemically valid 3D molecules is challenging due to bond topology errors.
method HLTF uses a latent multi-scale plan for global context and a constraint-aware sampler to suppress topology-driven failures.
result HLTF achieves high validity and uniqueness on QM9 and GEOM-DRUGS datasets.

Researchers developed a differentially private method for computing Wasserstein distances.

problem Computing divergences between distributions while preserving privacy.
method They focused on the Sliced Wasserstein Distance and added Gaussian perturbations to make it differentially private.
result They introduced a new differentially private distance, the Smoothed Sliced Wasserstein Distance, which performs well in generative models and domain adaptation.

Develops user-controlled privacy collaboration between users and data providers.

problem Ensuring data privacy while maintaining utility for users and providers.
method Collaborative learning of a sensitization function to control data sharing and privacy.
result Maintains data utility while fully protecting private information through user-controlled privacy settings.

Modeling influenza spread using feature engineering and international flow deconvolution.

problem Predicting and mitigating influenza spread through feature extraction and international flow analysis.
method Discrete Fourier Transform, matrix completion, SVM, autoencoders, PCA, deconvolution of international flow.
result Significant environmental and economic features are crucial to influenza mortality.

Study maps interdependence of SDGs, finds complex, dynamic linkages.

problem Identify which SDGs promote progress and how quickly.
method Used a balanced panel of 114 countries from 2000 to 2024, applying two estimators to recover directed interaction network and measure dynamic linkages.
result 84 goal linkages survive false-discovery control, showing both synergies and trade-offs, with no single goal acting as a universal accelerator.