A framework for anonymized risk sharing without revealing identities or preferences.
problem Risk sharing without revealing individual identities or preferences.
method Axiomatic framework with four key axioms: actuarial fairness, risk fairness, risk anonymity, and operational anonymity.
result The conditional mean risk sharing rule is uniquely characterized by these axioms.
First differentially private mechanism for anonymized histograms.
problem Protecting sensitive data in anonymized histograms.
method Proposes a differentially private mechanism for releasing anonymized histograms.
result Achieves near-optimal privacy utility trade-off in terms of number of items and privacy parameter.
Anonymization reduces economic signal extraction from financial texts.
problem Reducing meaningful economic signals from financial texts due to anonymization.
method Analyzed the impact of anonymization on textual understanding and economic signal extraction.
result Information loss due to anonymization is severe and pervasive, outweighing its benefits in certain financial applications.
Recommender systems are widely used to predict personalized preferences of goods or services using users' past activities, such as item ratings or purchase histories. If collections of such personal activities were made publicly available, they could be used to personalize a diverse range of services, including targete…
Anonymizes speech data to protect privacy.
problem Protecting personal speech data from misuse.
method Extracts features, uses x-vectors, and neural models to synthesize anonymized speech.
result Effective in concealing speaker identities without compromising speech quality.
New MAB problem with delayed, anonymous feedback analyzed.
problem Delayed, anonymous feedback in stochastic bandits.
method Phase-based extensions of UCB algorithm for SDCAF.
result Sub-linear regret guarantees for proposed algorithms.
We study a variant of the stochastic K-armed bandit problem, which we call "bandits with delayed, aggregated anonymous feedback". In this problem, when the player pulls an arm, a reward is generated, however it is not immediately observed. Instead, at the end of each round the player observes only the sum of a number…
Graph matching in noisy environments with Markovian errors.
problem Graph matching under time-dependent Markovian noise.
method Introduced edgelighter error model and analyzed graph matching thresholds.
result Graph matching thresholds and mixing times are of order Θ(n2logn) for Erdős-Rényi graphs, and O(nαlogn) for Stochastic Block Model graphs. Anonymizes sensor data to protect user privacy.
problem Protecting user privacy from motion sensor data.
method Information-theoretic approach and multi-objective loss function for deep autoencoders.
result Anonymized sensor data preserves activity recognition accuracy above 92% while keeping user identification accuracy below 7%.
Paper investigates preserving anomalous subgroups in anonymized datasets.
problem Preserving anomalous subgroups in machine learning transformed data.
method Trained a binary classifier to discover anomalous subgroups, then used variational autoencoder (VAE) to anonymize data.
result Synthesized datasets preserved high subgroup differentiation as in original data.
The study examines how brokers' identity affects their trading strategies on the Toronto Stock Exchange.
problem Impact of anonymous trading on brokers' optimal execution strategies.
method Formulated a stochastic differential game and mean-field game to analyze the optimal execution problem of anonymous and identity-revealed trading.
result Obtained a closed-form solution for the optimal strategy under Almgren-Chris price impact framework.
Adaptive MAB algorithms handle composite, anonymous feedback without reward interval knowledge.
problem Multi-armed bandit with composite and anonymous feedback, especially without reward interval size knowledge.
method Proposed adaptive algorithms for stochastic and adversarial cases, without reward interval knowledge.
result First algorithm for adversarial case handling non-oblivious adversary and unknown reward interval size.
Generative model generates synthetic medical images for data augmentation and anonymization.
problem Imbalanced medical imaging data sets, especially for rare pathologies.
method Generative adversarial network (GAN) trained on two public brain MRI datasets.
result Synthetic images improve tumor segmentation performance and serve as an anonymization tool.
Text-based analysis methods allow to reveal privacy relevant author attributes such as gender, age and identify of the text's author. Such methods can compromise the privacy of an anonymous author even when the author tries to remove privacy sensitive content. In this paper, we propose an automatic method, called Adver…
A blindfolded LLM trading framework validates market signals without ticker memorization.
problem Ensuring LLMs trade based on genuine market understanding, not memorized data.
method Anonymize tickers and company names, verify signals through reasoning embeddings, and use PPO-DSR policy.
result Achieved Sharpe ratio of 1.40 +/- 0.22 across 20 seeds, robust in volatile markets.
Preserving the privacy of individuals by protecting their sensitive attributes is an important consideration during microdata release. However, it is equally important to preserve the quality or utility of the data for at least some targeted workloads. We propose a novel framework for privacy preservation based on the …
TIPRDC anonymizes data features to protect privacy while retaining useful information.
problem Privacy concerns from crowdsourced data hinder deep learning applications.
method Hybrid training method combining adversarial and mutual information estimation.
result Feature extractor hides private information while preserving original data features.
This work synthesizes realistic data from neural excitation patterns to anonymize private data.
problem Lack of usable training data due to privacy regulations.
method Synthesize realistic data by exciting trained deep neural network neurons.
result Synthesized data can generalize well and anonymize participants' identities.
The paper proposes a privacy-preserving algorithm for decentralized learning using public-key cryptography.
problem Preserving data privacy in decentralized learning.
method Public-key cryptography and anonymization techniques.
result A scheme that provides anonymity in decentralized learning with theoretical guarantees.
The paper analyzes how clustering sensitive data can improve model generalization without revealing individual information.
problem Ensuring user data privacy in personalized recommendation systems.
method Look-alike clustering to replace sensitive features with cluster averages, analyzed using Convex Gaussian Minimax Theorem.
result Training models using anonymous cluster centers can improve generalization error, especially in high-dimensional settings.
New algorithm reduces privacy cost of LDP to central privacy model.
problem Collecting sensitive statistics from multiple users over time.
method Privacy amplification technique using permutation-invariant algorithms.
result Privacy cost of LDP can be much lower in central model.
WAFFLe anonymizes federated learning weights to protect data privacy and fairness.
problem Federated learning exposes local models to attacks and underfits heterogeneous clients.
method Combines Indian Buffet Process with shared weight factors.
result Significant improvement in local test performance and fairness.
Statistical methods protecting sensitive information or the identity of the data owner have become critical to ensure privacy of individuals as well as of organizations. This paper investigates anonymization methods based on representation learning and deep neural networks, and motivated by novel information theoretica…
UNMIX identifies hidden buyers in darknet markets by clustering anonymized IDs.
problem Identifying hidden buyers in darknet markets where IDs are anonymized.
method UNMIX, a hidden buyer identification model using Dirichlet Hawkes Process.
result UNMIX successfully groups transactions from one hidden buyer into one cluster.
Anonymizing company names in financial news improves trading performance, contrary to initial expectations.
problem Look-ahead and distraction biases in sentiment analysis of financial news.
method Investigated trading strategies based on original and anonymized headlines, comparing performance.
result Anonymized headlines outperform original in-sample, suggesting distraction effect is stronger.
GraLSP improves graph neural networks by incorporating local structural patterns.
problem GNNs struggle with identifying common structural patterns in graphs.
method GraLSP uses random anonymous walks to capture local graph structures and incorporates these into feature aggregation mechanisms.
result GraLSP outperforms other models in various prediction tasks on multiple datasets.
Study non-oblivious adversarial bandits with delayed feedback and propose algorithms with improved regret bounds.
problem Adversarial bandit problem with delayed, composite anonymous feedback.
method Propose wrapper algorithm for non-oblivious delay setting, achieving o(T) policy regret. result Achieve o(T) policy regret for many adversarial bandit problems with bounded memory loss sequences. We analyze a proprietary dataset of trades by a single asset manager, comparing their price impact with that of the trades of the rest of the market. In the context of a linear propagator model we find no significant difference between the two, suggesting that both the magnitude and time dependence of impact are univer…
A new method learns from pairwise comparisons to predict sensitive data without making strong assumptions.
problem Predicting sensitive data like annual income from unlabeled data with unknown target correspondence.
method Utilizes pairwise comparison data to learn a regression model without strong assumptions.
result The learned model converges to optimal with optimal parametric rate for uniformly distributed targets.
The paper explores how to measure and optimize ad reach while maintaining user privacy.
problem Measuring ad reach while preserving user privacy in online advertising.
method Introduces k-anonymity and probabilistic discounting for frequency capping. result Privacy introduces a significant performance drop but with manageable costs.
The task of representing entire graphs has seen a surge of prominent results, mainly due to learning convolutional neural networks (CNNs) on graph-structured data. While CNNs demonstrate state-of-the-art performance in graph classification task, such methods are supervised and therefore steer away from the original pro…
Paper proposes a framework to protect user anonymity in emotion recognition.
problem Preserving user anonymity in face-based emotion recognition systems.
method Adversarial learning framework using CNN architecture.
result The proposed approach minimizes identity-specific information and maximizes emotion-dependent information.
Differentially private GANs improve image privacy without significant quality loss.
problem Anonymizing image data sets while maintaining image quality.
method Training GANs with differential privacy on MNIST, analyzing privacy-utility trade-offs and explaining optimization methods.
result An increasing privacy budget adds little to generated image quality, revealing a saturated training regime.
Study evaluates gender bias in relation extraction systems.
problem Gender bias in relation extraction systems.
method Created WikiGenderBias dataset, evaluated systems for bias, analyzed bias mitigation techniques.
result NRE systems exhibit gender bias in predictions.
FOCA method prevents co-adaptation between feature extractor and classifier.
problem Co-adaptation between feature extractor and classifier degrades neural network performance.
method FOCA method uses randomly-generated, weak classifiers to optimize feature extractor without explicit co-adaptation.
result FOCA features form a point-like distribution within the same class under special conditions.
CAWs learn temporal network dynamics without node identities or edge attributes.
problem Learning temporal network dynamics without node identities or edge attributes.
method Causal Anonymous Walks (CAWs) using temporal random walks and hitting counts.
result CAW-N outperforms previous methods in predicting links over 6 real temporal networks.
Efficient sequential matching of supply and demand is a problem of interest in many online to offline services. For instance, Uber, Lyft, Grab for matching taxis to customers; Ubereats, Deliveroo, FoodPanda etc for matching restaurants to customers. In these online to offline service problems, individuals who are respo…
Framework for AI customer support that protects privacy and reduces costs.
problem Privacy risks and compliance challenges in AI customer support.
method Zero-shot learning with large language models, real-time data anonymization, retrieval-augmented generation, robust post-processing.
result Reduces privacy risks and compliance costs while maintaining accuracy.
Dynamic time warping applied to temporal graphs for pattern recognition.
problem Comparing temporal graphs using a proximity measure.
method Dynamic Time Warping (DTGW) on temporal graphs.
result DTGW is a flexible measure for temporal graph comparison, NP-hard in general but with polynomial-time solvable special cases.
We confirm the square-root law of market impact on Apple Inc. using a large dataset.
problem Testing the square-root law of market impact on a single U.S. large-cap equity.
method Using a full market-by-order feed, we reconstruct metaorders and calibrate impact using the square-root formula.
result The square-root law is confirmed with a prefactor of 0.34, consistent with worldwide data.
SMOTE-DP enhances synthetic data privacy without sacrificing utility.
problem Balancing privacy and utility in synthetic data generation.
method Integrating SMOTE with differential privacy mechanisms.
result SMOTE-DP produces synthetic data that is both private and useful.
Most users of online services have unique behavioral or usage patterns. These behavioral patterns can be exploited to identify and track users by using only the observed patterns in the behavior. We study the task of identifying users from statistics of their behavioral patterns. Specifically, we focus on the setting i…
Federated Learning (FL) systems are gaining popularity as a solution to training Machine Learning (ML) models from large-scale user data collected on personal devices (e.g., smartphones) without their raw data leaving the device. At the core of FL is a network of anonymous user devices sharing training information (mod…
The study identifies and analyzes cryptocurrency scams on social media.
problem Cryptocurrency fraud, especially 'pump and dump' scams, on social media platforms.
method Combining social media data to identify and predict scams, analyzing bot activity.
result Significant increase in bot activity during pump attempts.
Challenge identifies best-performing stocks over 6 months using financial predictors.
problem Identifying the best performing stocks over a 6-month period.
method Analyzed financial predictors and semi-annual returns; used various models including neural networks and boosting algorithms.
result Top six participants used diverse approaches, showcasing varied solutions.
Bank behaviour is important for pricing XVA because it links different counterparties and thus breaks the usual XVA pricing assumption of counterparty independence. Consider a typical case of a bank hedging a client trade via a CCP. On client default the hedge (effects) will be removed (rebalanced). On the other hand, …
Tourism is one of the most important economic activities in the world: for many countries it represents the single largest product in their export basket. However, it is a product difficult to chart: "exporters" of tourism do not ship it abroad, but they welcome importers inside the country. Current research uses socia…
MOAI evaluates indoor airflow's impact on COVID-19 transmission.
problem Understanding indoor airflow's role in COVID-19 transmission.
method Developed a privacy-preserving app and model to evaluate risk exposure.
result Quantified factors contributing to higher or lower contamination in settings.