Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

0111 · Jun 201619922001200920182026
15 results for De-identification

Objective: Patient notes in electronic health records (EHRs) may contain critical information for medical investigations. However, the vast majority of medical investigators can only access de-identified notes, in order to protect the confidentiality of patients. In the United States, the Health Insurance Portability a…

2016-06-10abs ↗pdf ↗

Patient notes contain a wealth of information of potentially great interest to medical investigators. However, to protect patients' privacy, Protected Health Information (PHI) must be removed from the patient notes before they can be legally released, a process known as patient note de-identification. The main objectiv…

2016-10-30abs ↗pdf ↗

Review of automatic de-identification systems for EHR, highlighting challenges beyond accuracy.

problem Challenges in surrogate generation and patient privacy in de-identification of EHR.
method Comprehensive review of 18 recently published systems, focusing on accuracy and challenges.
result Despite accuracy improvements, challenges remain in surrogate generation and patient privacy.

Automatic video modification to hide faces while maintaining pose, illumination, and expression.

problem Face de-identification in video to protect identities.
method A novel feed-forward encoder-decoder network architecture conditioned on facial image high-level representation.
result Fully automatic video modification at high frame rates with minimal distortion.

Proposes k-Same-Siamese-GAN for de-identifying facial images efficiently and preserving privacy.

problem Ensuring privacy of facial images while preserving useful data.
method Uses k-Same-Anonymity, GAN, hyperparameter tuning, and mixed precision training.
result Efficient de-identification of high-resolution facial images with privacy guarantees.

Recent approaches based on artificial neural networks (ANNs) have shown promising results for named-entity recognition (NER). In order to achieve high performances, ANNs need to be trained on a large labeled dataset. However, labels might be difficult to obtain for the dataset on which the user wants to perform NER: la…

2017-05-17abs ↗pdf ↗

Scrubbing PHI data from medical records is now efficient and scalable with SpaCy.

problem Efficiency and scalability of de-identification techniques for PHI data.
method Evaluated numerous deep learning techniques including SpaCy for performance and efficiency.
result SpaCy model is both well performing and extremely efficient for PHI data scrubbing.

Releasing full data records is one of the most challenging problems in data privacy. On the one hand, many of the popular techniques such as data de-identification are problematic because of their dependence on the background knowledge of adversaries. On the other hand, rigorous methods such as the exponential mechanis…

2017-08-26abs ↗pdf ↗

Clinical models trained on EHRs degrade in performance over time due to data drift.

problem Model performance degradation over time in clinical settings.
method Accessed year of care for each record in MIMIC, aggregated features into clinical concepts, and tested mitigation strategies.
result State-of-the-art models show significant performance drops when tested on future data compared to historical data.