C3L clusters data with user-controlled leakage, improving semi-supervised models.
problem Finding clusters in partially categorized data sets.
method Semi-supervised Gaussian mixture model with user-defined leakage level.
result C3L finds high-quality clustering models with controlled inconsistency.
bioLeak addresses data leakage in biomedical machine learning studies.
problem Data leakage causes optimistic bias in machine learning models for biomedical studies.
method bioLeak provides leakage-aware resampling workflows and model audits in R.
result The package supports various machine learning tasks and can detect leakage mechanisms.
New method detects and adjusts for temporal leakage in LLM backtests.
problem Standard backtest leakage detection is ineffective for modern models.
method Developed new methods to measure and adjust for temporal leakage.
result Demonstrated that models legitimately know more about times near their cutoffs, leading to structural leakage.
New approach to adaptive data analysis via maximal leakage, improving statistical guarantees.
problem False findings in published research due to disconnection between theory and practice in data analysis.
method Introduces Maximal Leakage, an information-theoretic measure, to compare adaptivity with non-adaptivity.
result Generalizes statistical guarantees for adaptive contexts, improving or replicating existing results.
New framework quantifies and reduces concept-based models' leakage.
problem Information leakage in concept-based models reduces interpretability.
method Information-theoretic framework with CTL and ICL measures.
result Measures predict model behaviour and identify leakage causes.
LLM forecasting benchmarks suffer from information leakage, which confounds model performance.
problem LLM forecasting benchmarks suffer from information leakage.
method A retrieval-augmented LLM forecaster observes only decision-time information.
result The full pipeline obtains a median monthly Spearman rank IC of +0.154.
The paper examines how trading strategies lose value due to stock turnover.
problem Leakage of rank-dependent trading strategies due to stock turnover.
method Theoretical analysis and empirical estimation of leakage in discrete time.
result A new method to estimate leakage in trading strategies is introduced.
This paper bounds min-entropy leakage for Blowfish privacy using graph symmetries.
problem Bounding min-entropy leakage for Blowfish privacy mechanisms.
method Organizing analysis over symmetrical partitions corresponding to orbits of graph automorphism groups.
result Demonstrates a construction meeting the bound with asymptotic equality, showing tightness.
Deep Leakage from Gradients can reveal private training data from shared gradients.
problem The safety of gradients in multi-node machine learning systems is questionable.
method Empirical validation of Deep Leakage from Gradient attack on computer vision and natural language processing tasks.
result The attack is more effective than previous methods, achieving pixel-wise accuracy for images and token-wise matching for texts.
Real-time fuel leakage detection framework MOCPD improves accuracy.
problem Early detection of fuel leakage to prevent hazards and losses.
method Memory-based Online Change Point Detection (MOCPD) framework.
result MOCPD outperforms baseline methods in detection accuracy.
Sage platform protects ML models trained on sensitive data from leakage.
problem Protecting sensitive data in machine learning models exposed to untrusted domains.
method Develops block composition for privacy accounting and privacy-adaptive training to manage privacy budget and utility tradeoff.
result Enables continuous training of models on sensitive data streams while maintaining global DP guarantees.
Unified analysis of privacy leakage in correlated data considering prior knowledge.
problem Understanding the impact of prior knowledge and data correlation on privacy leakage.
method Proposed prior differential privacy (PDP) and analyzed using WHG and multivariate Gaussian models.
result Derived closed-form expression for continuous data and chain rule for discrete data.
New model quantifies how much machine learning models can reveal about individual data usage.
problem Measuring and reducing the leakage of membership information from machine learning models.
method Using information theory, conditional mutual information leakage, and Kullback-Leibler divergence to quantify and bound the leakage.
result The amount of membership information leakage is reduced by adding Gaussian (ε,δ)-differentially-private additive noises. The paper analyzes generalization of noisy, iterative algorithms using maximal leakage.
problem Analyzing the generalization behavior of noisy, iterative learning algorithms.
method Information-theoretic framework with maximal leakage metric.
result Explicit upper bounds on maximal leakage for various scenarios.
Differential privacy protects data privacy by adding noise to data.
problem Leakage of sensitive data through common methods like encryption and endpoint protection.
method Randomized response technique to add noise to data collection.
result Differential privacy ensures strong privacy with better utility.
This paper evaluates and compares gradient leakage attacks in federated learning.
problem Gradient leakage attacks compromise client privacy in federated learning.
method Formal and experimental analysis of gradient leakage attacks, evaluation of attack effectiveness and cost.
result Gradient leakage attacks can reconstruct private local training data from shared parameter updates.
Reduces data leakage in distributed deep learning models.
problem Prevents reconstruction of sensitive raw data patterns during client communications.
method Reduces distance correlation between raw data and learned representations.
result Resilient to reconstruction attacks while maintaining model accuracy.
Paper tackles treatment leakage in text-based causal inference, proposing methods to mitigate bias.
problem Treatment leakage in text-as-confounder applications introduces bias in causal estimates.
method Formal definitions, four text distillation methods (passage removal, classification, salient feature removal, nullspace projection).
result Moderate distillation optimally balances bias reduction against confounder retention.
Researchers create an exact entangling gate using braiding and measurement of Fibonacci anyons.
problem No known leakage-free entangling gate using braiding of Fibonacci anyons.
method Supplement braiding with measurement operations to produce an exact controlled rotation gate.
result Exact entangling gate on two qubits created using Fibonacci anyons and measurement.
The paper analyzes privacy leakage in federated learning using linear algebra and optimization theory.
problem Privacy leakage in federated learning despite its promise for data privacy.
method Theoretical analysis from linear algebra and optimization theory perspectives.
result Derives sufficient conditions to prevent data reconstruction attacks and establishes an upper bound on privacy leakage.
ForesightFlow detects informed trading on prediction markets using an information leakage score.
problem Detecting informed trading on decentralized prediction markets.
method Developed an Information Leakage Score (ILS) framework to quantify the fraction of terminal information move priced in before public news events.
result The score connects label generation to proper-scoring-rule literature and reveals systematic biases in insider trading documentation.
Synth-MIA assesses privacy leakage in synthetic tabular data models.
problem Challenges in evaluating privacy leakage in synthetic tabular data.
method Unified threat framework deploying multiple attacks.
result Higher synthetic data quality correlates with greater privacy leakage.
New defense method protects client data privacy in federated learning.
problem Gradient leakage attacks in federated learning.
method Learning to obscure data to generate synthetic samples.
result Synthetic samples preserve predictive features and protect privacy.
Deep RL policies can leak private information from trained policies.
problem Privacy leakage in deep reinforcement learning models.
method Environment dynamics search via genetic algorithm and candidate inference based on shadow policies.
result 95.83% average recovery rate of floor plans from trained Grid World navigation DRL agents.
Analysis of model updates reveals sensitive data leaks.
problem Information leakage during model updates.
method Differential analysis of language model snapshots.
result New metrics (differential score, differential rank) reveal sensitive data leaks.
Paper mitigates information leakage in image representations using maximum entropy.
problem Mitigating unintended leakage of user information from image representations.
method Formulates an adversarial non-zero sum game to find an embedding function that maximizes task-dependent discriminative information while minimizing entropy of sensitive attributes.
result Proposed approach learns image representations with high task performance and reduced leakage of sensitive information.
Benchmark detects decision-time leakage in financial backtests.
problem Detecting decision-time leakage in financial machine-learning backtests.
method Toggles one evaluation convention at a time around a clean t+1-open reference, holding other factors fixed. result Inflation is highly selective, affecting specific features and execution methods.
Paper explores how poisoning data can increase privacy risks in machine learning models.
problem Increasing privacy risks of benign training samples through data poisoning attacks.
method Proposes generic and optimization-based attacks to amplify membership exposure.
result Demonstrates substantial increase in membership inference precision with minimal model performance degradation.
Framework prevents data leakage in mobile cloud DNNs.
problem Data leakage from cloud DNNs poses privacy risks.
method Privacy-preserving reinforcement learning framework.
result Framework successfully defends against various privacy attacks.
Multi-party machine learning leaks global dataset properties even with black-box access.
problem Leakage of global dataset properties in multi-party machine learning.
method Demonstrated leakage of sensitive attribute distributions in pooled data.
result A curious party can infer sensitive attribute distributions in other parties' data with high accuracy.
Benchmark evaluates LLM trading agents by masking identifiers to prevent memory leaks.
problem Evaluate LLM trading agents without relying on market memory or noise.
method Data-side masking protocol, Barra-style performance attribution framework.
result LLM agents' returns are largely explained by market and style exposure, not stock selection.
Deep learning reconstructs pressure fields and classifies leakage rates in CCS storage sites.
problem Monitoring CO2 leakage in CCS storage sites.
method Variational auto-encoder tailored for pressure field reconstruction and leakage rate classification.
result Uncertainty estimates of predictions illustrated on synthetic data.
PASE method protects machine learning models from membership inference attacks without significant accuracy loss.
problem Membership inference attacks on machine learning models.
method Switching ensembles approach to mitigate privacy leakage.
result PASE method provides effective privacy protection with minimal accuracy penalty.
New method recalibrates VaR for option books, reducing forecast errors.
problem Inaccurate VaR forecasts due to missing operational choices.
method Marking-aware sequential VaR recalibration targeting normalized book-level loss.
result Sequential VaR recalibration improves VaR performance across different markets and options.
fastml guards against data leakage in automated machine learning.
problem Data leakage during preprocessing before resampling inflates apparent performance.
method fastml uses guarded resampling to re-estimate preprocessing inside each resample.
result Guarded resampling reduces apparent performance compared to global preprocessing.
New insights link diverse statistical problems via secret leakage planted clique.
problem Statistical-computational gaps in inference problems.
method Secret leakage planted clique as a new hardness assumption for reductions.
result Establishes tight statistical-computational tradeoffs for various problems.
Log-Loss scores expose membership privacy breaches.
problem Privacy leakage from statistical aggregates like Log-Loss scores.
method Proved that Log-Loss scores enable full accuracy membership inference in a single query.
result Complete membership privacy breach is possible with Log-Loss scores.
RB-Modulation trains free diffusion models without external adapters.
problem Training-free personalization of diffusion models with style and content control.
method Stochastic optimal control with a style descriptor and cross-attention aggregation.
result Precise content and style extraction and control without external adapters.
Machine learning confound removal biases results, leading to misleading predictions.
problem Common confound removal methods in machine learning lead to misleading predictions.
method Featurewise removal of confound variance by linear regression before applying ML.
result This common deconfounding approach can leak information, amplifying null or moderate effects.
Improved canary crafting for one-run privacy auditing reduces leakage estimates.
problem Detecting canaries in one-run privacy auditing to estimate leakage effectively.
method Optimizes canaries for detectability and diversity, using a greedy initialization and bilevel optimization.
result Achieves stronger leakage estimates at lower computational cost.
New bounds on machine learning data leakage identified.
problem Machine Learning models can leak sensitive information.
method Formalized attack setups, derived universal bounds, studied mutual information.
result Connected attack success rate to generalization gap and mutual information.
New method detects information leakage using approximate Bayes predictor.
problem Unintentional exposure of sensitive information via observable data.
method Statistical learning theory and information theory framework, approximating Bayes predictor's log-loss and accuracy.
result MI can be accurately estimated to detect ILs, outperforming state-of-the-art baselines.
Develops a framework to obfuscate sensitive attributes in machine learning models.
problem Minimizing information leakage of sensitive attributes in crowdsourced data.
method Proposes a minimax optimization formulation and proves an information-theoretic lower bound.
result Adversarial learning achieves the best trade-off between attribute obfuscation and accuracy.
Optimal defenses protect FL models from gradient reconstruction attacks.
problem Gradient reconstruction attacks compromise FL models by recovering original data from shared gradients.
method Derive a theoretical lower bound of reconstruction error, customize noise and pruning defenses, and achieve optimal trade-off between leakage and utility.
result Our methods outperform Gradient Noise and Pruning in protecting training data and maintaining model utility.
TD learning with neural networks can lead to worse solutions than Monte-Carlo methods, especially in discontinuous value functions.
problem TD learning with neural networks can propagate approximation errors, leading to worse solutions than Monte-Carlo methods.
method Investigated the issue of approximation errors in areas of sharp discontinuities of the value function being further propagated by bootstrap updates.
result Empirical and analytical evidence shows that leakage propagation occurs in TD learning with function approximation, especially in sharp discontinuities.
Paper explores GAN generalization via privacy protection, proving bounds on generalization gap and reducing leakage.
problem Understanding and bounding the generalization gap of GANs.
method Theoretical proof using differential privacy, reinterpreting Bayesian GAN, membership attacks.
result Proven bounds on generalization gap and reduced information leakage.
Improved forecasting in daily time series competition using a correlator method.
problem Forecasting daily time series with data leakage issues.
method Ensemble of five statistical forecasting methods and a correlator method.
result The correlator method was responsible for most of the gains over naive forecasting.
DEDACT breaks down feature importance into direct and associative components.
problem Lack of clear distinction between direct and associative feature importance.
method DEDACT framework to decompose direct and associative importance measures.
result Provides insight into sources of prediction-relevant information and feature pathways.