Synthetic Data Generation for Privacy / Tạo Dữ liệu Tổng hợp cho Quyền riêng tư
📖 Synthetic Data Generation for Privacy
Definition (English):
Synthetic data generators (GANs, VAEs, diffusion models, LLMs) create artificial datasets that preserve statistical properties of real data without containing actual records. Metrics: statistical distance, utility preservation, privacy risk (membership inference, attribute inference). Synthetic data enables data sharing for ML without direct privacy exposure. Regulatory recognition: GDPR considers synthetic data as anonymized if re-identification risk is negligible.
📖 Tạo Dữ liệu Tổng hợp cho Quyền riêng tư
Định nghĩa (Tiếng Việt):
Trình tạo dữ liệu tổng hợp (GAN, VAE, diffusion models, LLM) tạo dataset nhân tạo bảo toàn thuộc tính thống kê của dữ liệu thực mà không chứa bản ghi thực tế. Chỉ số: khoảng cách thống kê, bảo toàn hiệu dụng, rủi ro quyền riêng tư (membership inference, attribute inference). Dữ liệu tổng hợp cho phép chia sẻ dữ liệu cho ML mà không phơi bày quyền riêng tư trực tiếp. Nhận định quy định: GDPR xem dữ liệu tổng hợp là ẩn danh nếu rủi ro tái nhận dạng là không đáng kể.
📂 Phân loại: Privacy