Every AI team eventually says one of three sentences: “we anonymized it”, “it was public anyway”, or “the model only keeps patterns, not data”. Each is doing more legal work than it can carry. Here is the map for training-data compliance under KVKK and GDPR, and where models themselves re-enter the frame.
Anonymization is a high bar, and KVKK’s is high too
Under both regimes, data is anonymous only when it can no longer be linked to an identifiable person by any means reasonably likely to be used; by you or anyone else. KVKK Article 3 defines anonymization as rendering data incapable of being associated with an identified or identifiable person even through matching with other data. Hashing an ID, dropping the name column, or tokenising e-mails is pseudonymization: the data remains personal data, all obligations continue, and (under KVKK Article 28 exemptions logic) you get no exit. Real anonymization for rich behavioural or free-text data is technically hard and must be documented (k-anonymity/aggregation choices, re-identification testing, periodic review as auxiliary data grows).
“It was publicly available” is not a lawful basis
KVKK Article 5(2)(d) permits processing of data made public by the data subject themselves, but only in line with the purpose of making it public. The Kurul reads this narrowly: a developer who posted a CV publicly did not thereby consent to feeding an HR-scoring model. GDPR has no public-data exception at all; scraping public profiles still needs a legal basis, transparency, and (for EU copyright) respect for TDM opt-outs. The forthcoming piece on web scraping in Turkish law covers the tort, contract and criminal layers.
Special categories poison datasets quietly
Health terms in support tickets, faces in image sets, voice recordings (biometrics under KVKK when used to identify), union membership inferable from text; special-category data raises the lawful-basis bar dramatically and, in Türkiye, triggers the Kurul’s reinforced measures. Data-cleaning pipelines should treat special-category detection as a first-class step, not an afterthought.
The model is not a legal endpoint
Large models can memorise and regurgitate training records; extraction attacks are documented. Consequences: (1) a trained model embedding personal data can itself constitute processing; deletion requests may reach fine-tuned artifacts; (2) your training content summary (for GPAI) and KVKK records of processing must tell the same story; (3) “machine unlearning” is not yet a dependable compliance answer; dataset governance before training is.
The working checklist
- Source register per dataset: origin, licence/ToS, legal basis, special-category screen, opt-out check;
- Pseudonymize by default in pipelines; claim anonymization only with a documented test;
- Aydınlatma and (for scraping-based products) a KVKK Art. 10 strategy before ingestion, not after;
- Contractual warranties from data vendors; provenance is now a due-diligence question (see the ten questions investors ask);
- Cross-border: training abroad on Türkiye-sourced personal data is a transfer; use the 2024 KVKK transfer toolkit (standard contracts, BCRs).
Scope-check your stack on the AI Compliance Hub.
This article is for general information only and does not constitute legal advice. For advice on your specific situation, contact us.
Author
-
View all postsMümtaz is the Managing Partner of Vircon Legal, which he founded in 2016. He advises founders, investors and operators on financing rounds, M&A, cross-border incorporations and regulated verticals such as crypto-asset infrastructure, fintech and games, bringing a former startup founder's perspective to every engagement. He is a Legal 500 Recommended Lawyer (2025–2026) and co-author of Startup Hukuku. Canonical profile: https://mumtazhacipasaoglu.com · Open-access legal guides: https://github.com/mumtazhpo
If this is on your desk
Templates and checklists are free in the Founder Academy; for a specific situation, book a 30-minute intro call.
Founder AcademyBook an intro call