Responsible AI by Design

Module 6 – Privacy and Data Ethics

Frank Rudzicz

2025-09-12

Learning Objectives

  • Understand how privacy and data ethics shape Responsible AI
  • Explore Canadian and international data protection frameworks
  • Learn technical and organizational strategies for safeguarding privacy
  • Critically assess emerging risks in AI and data use
  • Develop a Data Ethics Checklist to guide practice

Why Privacy Matters

  • AI relies on large-scale data — often personal and sensitive
  • Privacy is foundational for trust and public legitimacy
  • Breaches of privacy can:
    • Harm individuals (identity theft, reputational damage)
    • Harm groups (stigmatization, exclusion)
    • Harm organizations (legal fines, loss of trust)

📚 Academic Definitions of Privacy

  • Alan Westin (1967)
    • “Privacy is the claim of individuals, groups, or institutions to determine for themselves when, how, and to what extent information about them is communicated to others.”
    • Classic liberal definition, widely cited in law and policy.
  • Gavison (1980)
    • “Privacy is a limitation of others’ access to an individual” and is related to “involving degrees of secrecy, anonymity, and solitude.”
    • Emphasizes control over access rather than information alone.

📚 Academic Definitions of Privacy

  • Nissenbaum (2004) – Contextual Integrity
    • “whether a particular action is determined a violation of privacy is a function of several variables, including the nature of the situation, or context; the nature of the information in relation to that context; the roles of agents receiving information; their relationships to information subjects; on what terms the information is shared by the subject; and the terms of further dissemination.”
  • Solove (2006)
    • “privacy is best understood as a family resemblance concept.”
    • Suggests privacy is not a single essence, but a cluster of related concerns.

🏈 HIPPA

  • Expert Determination
    • “A person with appropriate knowledge of and experience with generally accepted statistical and scientific principles and methods [de-identifies] information” to a particular level of risk.
  • Safe Harbor (sic)
    • Removal of names, geographic divisions < a ‘state’, dates (except year), telephone numbers, vehicle identifiers, fax numbers, email addresses, SINs, medical records, full-face photographs, …

Activity 6.1: Data leakage

  • Reflect:
    • ✍️ Which definition resonates most with your business/sector?
    • 💧 What personal data did you generate today without realizing?
    • 🤔 Is there any data you would not want to share with the public? With a private entity?

Foundations of Data Privacy

  • Personal Data: Information that identifies or could identify an individual
  • Sensitive Data: Health status, political beliefs, sexual orientation
  • Inferred Data: Predictions or classifications about a person
    • If an extremely accurate model predicts that you like 🍦ice cream, is that prediction now part of your sensitive data?

⚖️ Core Privacy Rights (1/2)

  • Access 🔍
    • Individuals can obtain a copy of their personal data held by an organization.
    • Ensures transparency and builds trust.
  • Correction ✏️
    • Right to request rectification of inaccurate or incomplete data.
    • Critical for fairness and data quality.

⚖️ Core Privacy Rights (2/2)

  • Erasure (“Right to be Forgotten”) 🗑️
    • Individuals may demand deletion when data is no longer necessary, consent is withdrawn, or processing is unlawful.
    • Limits long-term risks of data misuse.
  • Portability 🔄
    • Right to receive data in a structured, machine-readable format and transfer it to another controller.
    • Promotes competition and user empowerment.

Note

These rights are enshrined in GDPR, echoed in Canada’s proposed CPPA, and increasingly reflected in global AI/data ethics frameworks.

⚠️ Core Privacy Risks (1/3)

  • Re-identification 🕵️
    • Even “anonymized” data can often be linked back to individuals.
    • Example: Netflix Prize dataset re-identified by cross-referencing with IMDb ratings (Narayanan and Shmatikov 2007).
    • Risk increases when multiple datasets are combined (data linkage).

From 🔗HHS

⚠️ Core Privacy Risks (1/3)

Warning

Despite your best efforts in anonymization (see here), anonymization is not robust to linkage (i.e., if you have ancillary data)

⚠️ Core Privacy Risks (2/3)

  • Inference Risks 🔮
    • Sensitive attributes (e.g., health status, political beliefs) can be inferred from non-sensitive data.
    • Common in machine learning models trained on behavioural data.
  • Secondary Use 🔄
    • Data collected for one purpose used for another without consent.
    • Violates principles of purpose limitation and informed consent.

⚠️ Core Privacy Risks (3/3)

  • Data Breaches & Unauthorized Access 🔓
    • Insider threats or weak security controls can expose personal data.
    • Regulatory consequences under GDPR, PHIA, and proposed CPPA.
  • Opacity of AI Models 🧩
    • Black-box systems may use personal data in ways that are invisible to users.
      • Have they been trained in a way that merely compresses your data?
    • Raises concerns about fairness, accountability, and trust.

Warning

Privacy risks are dynamic and evolving:
- Re-identification and inference show that technical anonymization is not enough.
- Strong governance, continuous monitoring, and technical safeguards (e.g., Differential Privacy) are essential.

Hacker’s delight

  • ‘Hackers’ are not necessarily looking for your data specifically.
    • They want to find a relatively small set of people who may be targeted, from a larger set
  • This is similar to the concept of ‘multiple comparisons’.
    • The chance that they can get your information may be very small, but the chance that they can get someone is high.
  • This does not need to be malevolent – human error is also prevalent.

Case Study: Cambridge Analytica

  • What: Data from millions of Facebook users harvested, including data of users’ friends, without consent. This enabled psychographic profiling for political microtargeting.
  • Key Privacy Breach: the mosaic effect (linking cross-source social data to re-identify and build behavioural profiles).
    • This combined with algorithmic inference allows nearly de-identified data to become traceable.
  • See Hinds, Williams, and Joinson (2020).

Case Study: Clearview AI

  • Collected 3B+ images from the web (incl. social media) without consent to power facial recognition
  • BC, Alberta, and Québec jointly investigated Clearview.
    • Clearview’s data collection must comply with provincial privacy laws, even if the images were publicly accessible online.
    • Clearview 🔗 was used by Halifax Regional Police
  • Additionally, Clearview’s past class-action clients list was stolen in a data breach, further raising ethical and security concerns.

From 🔗The Conversation

Jung and Kwon (2024) explored governance gaps. Shan et al. (2020) proposed adding pixel-level cloaks but Radiya-Dixit et al. (2022) suggest such ‘data-poisoning’ has limits.

Case Study: Google & NHS

  • Context 🏥: In 2015, Google’s DeepMind partnered with the Royal Free London NHS Foundation Trust.
    • Aim: develop Streams, an app to help clinicians detect acute kidney injury (AKI).
  • The Issue ⚠️: DeepMind was given access to 1.6 million patient records.
    • Data included sensitive information (HIV status, mental health, drug use).
    • Patients were not informed; no clear legal basis for such wide data sharing.
  • Fallout: UK Information Commissioner’s Office ruled (2017) that the data sharing violated the UK Data Protection Act.
    • Patients were not properly informed; “reasonable expectations” of privacy were breached.

Note

Powles and Hodson (2017)

Case Study: Strava Heat Map

  • Background: Strava (🏃) released a global activity heatmap (2018), visualizing 3 trillion GPS data points from runners & cyclists, intended as feature.
  • The Problem ⚠️: Heatmap unintentionally exposed sensitive military locations (🔗BBC, 🔗Guardian).
  • Lesson 📚: Aggregate ≠ anonymous when context-sensitive data is involved.
    • Even when individual identities are hidden, patterns can endanger groups (e.g., soldiers, aid workers).

Data Protection Regulations

🇪🇺 Lawful Bases for Processing

GDPR (Article 6)

  • Consent
    Must be freely given, specific, informed, unambiguous; withdrawal must be as easy as giving.
  • Contract 📑
    Processing necessary to perform or prepare a contract with the data subject.
  • Legal obligation ⚖️
    Required by EU or Member State law (e.g., tax reporting).
  • Vital interests ❤️
    Protecting life or health when consent is not possible.
  • Public task 🏛️
    Processing necessary for official authority or public interest functions.
  • Legitimate interests 💼
    Flexible basis for controllers unless overridden by fundamental rights.

🇨🇦 CPPA (Canada, 2022 draft)

  • PIPEDA is still the current baseline
  • New Consumer Privacy Protection Act (CPPA) may be in limbo
  • Stronger enforcement: 🔗Privacy Commissioner empowered to issue orders
  • New tribunal for fines up to max(5% of global revenue, $25M)
  • Requires transparency in automated decision-making

Note

Implication for AI: Explainability and consent would be mandatory.

🇨🇦 CPPA (Canada, 2022 draft)

  • Consent as default
    Organizations must obtain valid, meaningful consent (clear, accessible).
  • Exceptions (no consent needed; documented rationale):
    • Business operations (e.g., fraud prevention, IT security, product safety).
    • Legitimate business interests (narrower than GDPR’s; must be documented).
    • De-identified information: may be used internally without consent, but not for influencing decisions about individuals.
    • Research & development: if safeguards are in place.

Note

Key Distinctions
- GDPR lists six equal bases; CPPA sets consent as the rule, with limited exceptions.
- CPPA requires plain-language explanations of why data is collected/used.
- Both frameworks stress accountability, but CPPA is more business-operations focused, reflecting Canadian regulatory pragmatism.

Cross-border Data Transfers

  • Background 🌍
    • Many businesses rely on cloud services & processors based in the U.S.
    • EU–U.S. “Privacy Shield” was the framework for lawful transfers.
  • 🔗Schrems II (CJEU, 2020) ⚖️
    • Court struck down Privacy Shield as inadequate.
    • Reason: U.S. surveillance laws (e.g., FISA 702, Executive Order 12333) allow disproportionate access to EU personal data without effective redress.
    • Standard Contractual Clauses (SCCs) remain valid, but only given “supplementary safeguards” (e.g., encryption, pseudonymization).

Implications

  • For EU Businesses
    • Extra due diligence required when using U.S.-based processors.
    • Must conduct Transfer Impact Assessments (TIAs).
    • Supplementary measures often needed (e.g., end-to-end encryption, data minimization).
  • For Canada & CPPA 🇨🇦
    • No adequacy decision with the EU (unlike Japan or UK).
    • Organizations transferring data from EU → Canada must use SCCs + safeguards.
    • Proposed CPPA incl. transparency requirements around international transfers

Warning

Key Takeaway:
Schrems II shifted the burden onto organizations to prove equivalent protection abroad.
Cross-border transfers are now legally and technically complex, with real compliance risk.

Data Protection Impact Assessments (DPIA/PIA)

  • Purpose
    • Early warning system 🛑: Anticipates privacy risks before a system or project goes live.
    • Risk-based approach ⚖️: Focus on high-risk uses of personal data (large-scale monitoring, sensitive data, automated decisions).
    • Accountability tool 📒: Documents compliance decisions, safeguards, and residual risks.

DPIA Process

  1. Describe processing
    • Data flows, purposes, stakeholders.
  2. Assess necessity & proportionality
    • Is all data needed? Is there a less intrusive option?
  3. Identify risks
    • Re-identification, bias, security vulnerabilities, misuse.
  4. Mitigation measures
    • Technical (encryption, differential privacy), organizational (training, access limits).
  5. Review & update
    • Living document — updated as system evolves.

Tip

  • 👀 Read the sample 🇪🇺GDPR DPIA template 🔗here
  • 👀 Read the sample 🇨🇦Canadian PIA template 🔗here

Theoretical Foundation

  • Privacy by Design (Cavoukian 2009) 🌱
    • 7 principles: proactive, default, embedded, positive-sum, end-to-end, visible, user-centric.
    • DPIAs operationalize these principles by embedding privacy safeguards into design rather than bolting them on after deployment.

Tip

Key Takeaway:
A DPIA/PIA is not just a compliance checkbox — it is a strategic governance tool that builds trust, demonstrates accountability, and reduces long-term legal/ethical risk.

Swiss cheese

There are (nevertheless porous) layers which can be used to protect our data

  1. Avoiding ‘dark patterns’
  2. \(k\)-anonymity
  3. Obfuscation
  4. Differential privacy
  5. Federated learning

2. \(k\)-anonymity

  • Definition 📊: A dataset satisfies \(k\)-anonymity if each record is indistinguishable from at least \(k–1\) others wrt certain quasi-identifiers (e.g., age, sex, ZIP).
  • How it works: Generalize or suppress values until every individual “blends in” with at least \(k-1\) others (e.g., exact age → anges (20–29, 30–39)).
  • Benefits ✅: Reduces risk of re-identification through dataset linkage. Basis for regulatory guidance (e.g., 🔗 HIPAA Safe Harbor).
Age (Years) Sex ZIP Code Diagnosis
16 Male 00002 Diabetes
20 Female 00000 Influenza
34 Male 10000 Broken Arm
93 Female 10003 Acid Reflux
Age (Years) Sex ZIP Code Diagnosis
\(< 30\) 00000* Diabetes
\(< 30\) 00000* Influenza
\(\geq 30\) 10000* Broken Arm
\(\geq 30\) 10000* Acid Reflux

Note

Source: U.S. HHS, Guidance Regarding Methods for De-identification of Protected Health Information (Table 5).
See: HHS HIPAA De-identification Guidance and El Emam et al. (2009).

Warning

Still susceptible to linking.

3. Obfuscation: Concepts & Ethics

  • Data obfuscation 🔒: Masking, blurring, pseudonymization of sensitive features in data (faces, names, identifiers).
    • E.g., pixelating or skeletonizing humans in video.
  • Code obfuscation 💻: Altering software code to prevent reverse-engineering (used in IP protection & security).
  • Ethical tensions ⚖️: Protects privacy but reduces transparency, interpretability, and sometimes accountability.
    • Trade-off between protecting individuals vs. enabling oversight (e.g., in healthcare, policing).
  • Theoretical anchor 📚: 🔗Brunton & Nissenbaum (2015): Obfuscation as privacy protest—a tactical response by individuals against surveillance and data extraction.

3. Obfuscation example

3. Obfuscation: Developments

  • Video obfuscation methods 🎥: Skeletonization & pose-based anonymization: replacing bodies with stick figures while preserving action semantics (Z. Wang et al. 2022).
  • Text obfuscation ✍️: Adversarial stylometry (H. Wang 2023): altering linguistic features to conceal authorship while maintaining meaning.
    • Tools like 🔗Mutant-X and 🔗Anonymouth modify syntax/word choice to defend against deanonymization.
  • Cross-cutting risks 🚨: Obfuscation can undermine trust if seen as deceptive.
    • May not be robust against advanced re-identification (e.g., deep learning models trained on obfuscated data).
    • SotA word-replacement tools are ~99% accurate (Zhang and Jiang 2024).

4. Differential Privacy

  • 💡Solution: Don’t obfuscate the data, obfuscate the model
  • Differential Privacy (DP) ensures that the probability of any output is nearly the same whether or not an individual’s data is included in the dataset.
    • This guarantees indistinguishability: attackers cannot confidently tell if a specific person’s data was used.
  • It accomplishes this through adding noise to the model

(Dwork and Roth 2014)

From 🔗here

4. Differential Privacy

  • Noise injection: Random noise (from Laplace or Gaussian distributions) is added to counts, sums, or model updates.
  • Balance:
    • Smaller \(\epsilon\) → 🔒stronger privacy, 📉less accuracy.
    • Larger \(\epsilon\) → 🔓weaker privacy, 📈higher utility.
  • Guarantee: No attacker can infer whether a specific individual contributed, within the bounds of \(\epsilon\).

4. Differential Privacy

  • Apple 🍏🪱: Uses DP in iOS for keyboard predictions, emoji usage, and Maps traffic patterns.
  • U.S. Census (2020) 🏈: First national census to implement DP at scale. Protected sensitive sub-population counts but raised debates over accuracy in small communities.
  • Canadian Research 🇨🇦: Applied DP to health datasets (Nova Scotia, Ontario) to enable epidemiological studies without exposing patients.
    • Aligns with values of data minimization & consent embedded in PIPEDA, proposed CPPA, and Nova Scotia’s PHIA.
    • But increasing privacy can decrease fairness! (Dadsetan et al. 2024)

5. Federated Learning

  • Definition: A machine learning paradigm where models are trained across decentralized devices or servers holding local data samples, without transferring raw data to a central server.
  • Core benefit: Keeps personal or sensitive data on-device or in-institution, reducing privacy risks.
  • Origin: Popularized by Google (2017) for mobile keyboard prediction (Kairouz et al. 2021).

5. How FL Works

  1. Local training
    • Each device/institution trains a model update using its own data.
  2. Aggregation
    • Updates (gradients, parameters) are sent to a central server.
    • The server aggregates updates into a global model.
  3. Privacy layers
    • Differential Privacy: noise added to updates.
    • Secure aggregation: cryptographic protocols ensure only aggregated results are visible.

From 🔗here

5. FL Advantages

  • Data never leaves local environment → reduced exposure risk.
  • Scalable across millions of devices (phones, hospitals, schools).
  • Adaptive → models can continuously improve while respecting privacy.

5. FL Challenges

  • Communication cost: transmitting model updates at scale.
  • Heterogeneity: devices/institutions may have different data distributions.
  • Security risks: poisoned updates (adversarial participants) can corrupt the model.
  • Transparency: harder to audit when data remains siloed.

5. Applications of FL

  • Healthcare: Hospitals can collaboratively train models for diagnostics without sharing patient data (e.g., radiology imaging).
  • Finance: Banks can use FL to detect fraud across institutions without centralizing sensitive transaction data.
  • Complements DP: together, they form a privacy-preserving AI toolkit.
    • 🛠️ Florist: A platform to launch and monitor FL jobs
    • 🛠️ FL4Health: A modular library to facilitate FL in healthcare.

Tip

Key takeaway: FL “moves the model to the data rather than the data to the model”.
It is a cornerstone of privacy-preserving AI, especially powerful when paired with Differential Privacy.

5. To be continued…

(Zeng and Rudzicz 2025)

Emerging Challenges

  • Generative AI 🧠: Models trained on scraped data without consent → privacy + copyright concerns.
    • E.g., lawsuits over data used to train large language models. (Lemley 2024)
  • Synthetic data 🔄: Used to augment or replace real data; protects privacy in principle.
    • Risk: can still replicate underlying bias. 🔗OECD
  • Data brokers 💰: Opaque industry trading personal data (location, health, consumer).
    • Challenges: lack of transparency, weak accountability. 🔗EPIC

Tools and Practices

Drafting a Data Ethics Checklist

  • Checklist elements:
    • What personal/sensitive data is used?
    • Is consent informed, revocable?
    • How is data minimized, anonymized, or protected?
    • Is a PIA/DPIA completed?
    • How are Indigenous/sovereignty rights respected?
    • What redress is available if harm occurs?

Warning

Again, don’t consider a ‘checklist’ to be a “one-time thing”

Activity 6.2: Checklist

  • Task: Develop a Data Ethics Checklist for your organization’s AI project.
  • Requirements (0.5-1.5 pages):
    • Consent (How will you obtain and document it?)
    • Anonymization (How will you protect identities?)
    • Stewardship (Who is responsible for data governance?)
    • Retention (How long is data kept, and why?)
    • Transparency (What will users and stakeholders be told?)
  • Apply your checklist to one AI system (real or hypothetical) in your sector.
    • E.g.,: a hiring algorithm, healthcare triage tool, or customer chatbot.
    • Show briefly how each item on your checklist would be met.

Tip

Keep it practical — imagine you are advising your organization on how to use data responsibly in this project.

Wrap-Up

  • Privacy is one foundational pillar for ethical AI adoption
  • Canadian context: CPPA, provincial laws
  • Technical tools: obfuscation, DP, FL
  • Organizational tools: DPIAs,
  • 👀 Emerging challenges demand ongoing vigilance

References

Abdalla, Mohamed, Moustafa Abdalla, Frank Rudzicz, and Graeme Hirst. 2020. “Using Word Embeddings to Improve the Privacy of Clinical Notes.” Journal of the American Medical Informatics Association : JAMIA 27 (6): 901–7. https://doi.org/10.1093/jamia/ocaa038.
Cavoukian, Ann. 2009. “Privacy by Design.” https://www.privacybydesign.ca.
Dadsetan, Ali, Dorsa Soleymani, Xijie Zeng, and Frank Rudzicz. 2024. “Can Large Language Models Be Privacy Preserving and Fair Medical Coders?” arXiv. https://doi.org/10.48550/arXiv.2412.05533.
Dwork, Cynthia, and Aaron Roth. 2014. https://doi.org/10.1561/0400000042.
El Emam, Khaled, Fida Kamal Dankar, Romeo Issa, Elizabeth Jonker, Daniel Amyot, Elise Cogo, Jean-Pierre Corriveau, et al. 2009. “A Globally Optimal k-Anonymity Method for the De-Identification of Health Data.” Journal of the American Medical Informatics Association 16 (5): 670–82. https://doi.org/10.1197/jamia.M3144.
Gavison, Ruth. 1980. “Privacy and the Limits of Law.” The Yale Law Journal 89 (3): 421. https://doi.org/10.2307/795891.
Gebru, Timnit, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. 2021. “Datasheets for Datasets.” arXiv. https://doi.org/10.48550/arXiv.1803.09010.
Hinds, Joanne, Emma J. Williams, and Adam N. Joinson. 2020. It Wouldn’t Happen to Me’: Privacy Concerns and Perspectives Following the Cambridge Analytica Scandal.” International Journal of Human-Computer Studies 143 (November): 102498. https://doi.org/10.1016/j.ijhcs.2020.102498.
Jung, Won Kyung, and Hun Yeong Kwon. 2024. “Privacy and Data Protection Regulations for AI Using Publicly Available Data: Clearview AI Case.” In Proceedings of the 17th International Conference on Theory and Practice of Electronic Governance, 48–55. Pretoria South Africa: ACM. https://doi.org/10.1145/3680127.3680200.
Kairouz, Peter, H. Brendan McMahan, Brendan Avent, Aurélien Bellet, Mehdi Bennis, Arjun Nitin Bhagoji, Kallista Bonawitz, et al. 2021. “Advances and Open Problems in Federated Learning.” arXiv. https://doi.org/10.48550/arXiv.1912.04977.
Lemley, Mark A. 2024. “How Generarative AI Turns Copyright Upside Down.” SCIENCE & TECHNOLOGY LAW REVIEW XXV. https://law.stanford.edu/wp-content/uploads/2024/09/2024-09-30_How-Gerative-AI-Turns-Copyright-Upside-Down.pdf.
Mitchell, Margaret, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2019. “Model Cards for Model Reporting.” In Proceedings of the Conference on Fairness, Accountability, and Transparency, 220–29. https://doi.org/10.1145/3287560.3287596.
Narayanan, Arvind, and Vitaly Shmatikov. 2007. “How To Break Anonymity of the Netflix Prize Dataset.” arXiv. https://doi.org/10.48550/arXiv.cs/0610105.
Nissenbaum, Helen. 2004. “Privacy as Contextual Integrity.” Washington Law Review 79.
Powles, Julia, and Hal Hodson. 2017. “Google DeepMind and Healthcare in an Age of Algorithms.” Health and Technology 7 (4): 351–67. https://doi.org/10.1007/s12553-017-0179-1.
Radiya-Dixit, Evani, Sanghyun Hong, Nicholas Carlini, and Florian Tramèr. 2022. “Data Poisoning Won’t Save You From Facial Recognition.” arXiv. https://doi.org/10.48550/arXiv.2106.14851.
Shan, Shawn, Emily Wenger, Jiayun Zhang, Huiying Li, Haitao Zheng, and Ben Y Zhao. 2020. “Fawkes: Protecting Privacy Against Unauthorized Deep Learning Models.” In Proceedings of the 29th USENIX Security Symposium. https://www.usenix.org/system/files/sec20-shan.pdf.
Solove, Daniel J. 2006. “A Taxonomy of Privacy.” University of Pennsylvania Law Review 154 (3): 477. https://doi.org/10.2307/40041279.
Sun, Qianru, Ayush Tewari, Weipeng Xu, Mario Fritz, Christian Theobalt, and Bernt Schiele. 2018. “A Hybrid Model for Identity Obfuscation by Face Replacement.” arXiv. https://doi.org/10.48550/arXiv.1804.04779.
Wang, Haining. 2023. “Defending Against Authorship Identification Attacks.” arXiv. https://doi.org/10.48550/arXiv.2310.01568.
Wang, Zhengjie, Mingjing Ma, Xiaoxue Feng, Xue Li, Fei Liu, Yinjing Guo, and Da Chen. 2022. “Skeleton-Based Human Pose Recognition Using Channel State Information: A Survey.” Sensors 22 (22): 8738. https://doi.org/10.3390/s22228738.
Zeng, Xijie, and Frank Rudzicz. 2025. “How to Recover Long Audio Sequences Through Gradient Inversion Attack With Dynamic Segment-Based Reconstruction.” In Interspeech, 5118–22. Rotterdam, The Netherlands. https://doi.org/10.21437/Interspeech.2025-244.
Zhang, Kai, and Xiaoqian Jiang. 2024. “Sensitive Data Detection with High-Throughput Machine Learning Models in Electrical Health Records.” In AMIA Annu Symp Proc, 814–23. https://pmc.ncbi.nlm.nih.gov/articles/PMC10785837/.