A privacy policy describes what a company intends to do with your data. The architecture decides what the software is able to do with it, and when somebody makes a mistake, the architecture is the part that still holds. The argument here is that privacy by design for on-device AI has to be built as several architectural properties stacked on top of each other: local storage, pseudonymisation before any model call, zero retention and data minimisation. Each one deals with a specific, nameable threat and leaves others completely untouched. The article sets out the threat model, what each layer fails to protect against, and explicit rules for deciding whether a payload may leave the device at all.
The Threat Model for On-Device AI Privacy
If privacy work starts from compliance language and never names an adversary, you end up with controls whose effect nobody can test. A product that runs a language model over intimate personal data faces six adversaries, and each of them gets past a different control, so they have to be named separately. Transport encryption stops an opportunistic network attacker. A breached vendor or sub-processor is stopped only if the payload was never stored there. An insider with production access is stopped by pseudonymising the content, and access logging does not stop them. Legal compulsion fails if the data does not exist on the compelled party's systems. A local adversary with physical access, such as a partner, a family member or an employer, is stopped by device encryption and authentication, and on-device storage makes this adversary more relevant. The sixth adversary is the operator's own future. An archive that is safe under today's ownership, jurisdiction and staff is not automatically safe under whoever runs the company later.
Intimate data also changes what a breach costs. A leaked password can be rotated, but a leaked note about a friend's illness or a sentiment score attached to a named person cannot be reissued. The harm is social more than financial, and the person affected is often a third party who never agreed to anything. Saltzer and Schroeder's principles (Saltzer & Schroeder, 1975) are still the right frame: least privilege, economy of mechanism, fail-safe defaults, complete mediation, and open design instead of obscurity. The OWASP Top 10 is the right baseline for the application surface. Neither of them tells you whether the payload should be sent anywhere in the first place.
Privacy by Design and the Legal Frame
Cavoukian's privacy by design framework (Cavoukian, 2009) states the position as a set of norms. Privacy should be proactive instead of remedial and on by default, built into the design from the start and kept up across the whole data lifecycle. People often dismiss the framework as aspirational, but read as an engineering brief it is concrete. "Privacy as the default setting" describes what the code does when the user does nothing.
European law turned much of this into obligation. The GDPR (Regulation (EU) 2016/679) makes data minimisation a principle under Article 5, requires data protection by design and by default under Article 25, and requires a data protection impact assessment under Article 35 where processing is likely to result in a high risk to data subjects. A product that infers sentiment from private message histories meets the Article 35 triggers without much doubt. The processing is systematic, it runs across a user's whole social graph, it uses a new technology, and special-category inferences are easy to derive from the data it touches. Expect to write a DPIA.
Turkey's Law No. 6698 on the Protection of Personal Data (KVKK) has a similar structure: lawful basis, purpose limitation, proportionality, security obligations on the controller, and a stricter regime for special categories. It adds its own registration and cross-border transfer rules. A product shipping into both jurisdictions should design to whichever constraint is stricter. In practice that gives one rule you can test against: use the smallest amount of data that makes the feature work, keep it for the shortest time, and keep it in the fewest places.
Pseudonymisation Before the Model Call, and Its Limits
The most important control at the boundary is what happens to a payload just before it leaves the device. RelationCRM, our relationship-management product in development, keeps the corpus on the device and replaces contacts with labels of the form Contact-001 before any model call. No payload is retained after the response, and the design is aligned to both GDPR and KVKK. The relationship graph holds 150+ relationships scored across five dimensions, for reasons that come from the attention limits that make a scored relationship graph necessary in the first place. It never leaves local storage as identified data.
This does reduce exposure, but it is pseudonymisation and not anonymisation. Under the GDPR pseudonymised data still counts as personal data, because the mapping that reverses it exists. The label replaces the direct identifier and leaves the quasi-identifiers in place, and those are usually enough. An extract that mentions a workplace, a city, a birthday and a medical detail identifies a person whether the label above it reads Ayse or Contact-001. Narayanan and Shmatikov (2008) showed the general result for sparse high-dimensional data. Behavioural traces are close to unique, and a modest amount of outside knowledge is enough to link them back to a person. Relationship histories are exactly this kind of data.
There are two more limits. If the same token shows up across sessions that can be correlated, an adversary can build a profile under the pseudonym without ever learning the real name. And pseudonymisation does nothing about content. In "Contact-004 is leaving her husband", the sensitive part is what the sentence says about her.
What k-Anonymity and Differential Privacy Actually Guarantee
Both come up often in privacy claims for single-user products, and they are usually applied wrongly. Writing them out precisely shows why.
k-anonymity (Sweeney):
for every released record r in R,
| { r' in R : QI(r') = QI(r) } | >= k
where QI(.) is the tuple of quasi-identifier attributes
epsilon-differential privacy (Dwork):
for all datasets D, D' differing in one record,
and all measurable S subset of Range(M),
Pr[ M(D) in S ] <= exp(epsilon) * Pr[ M(D') in S ]
k-anonymity (Sweeney, 2002) requires each released record to be indistinguishable from at least k-1 others on its quasi-identifiers. It is a necessary sanity check on aggregate releases, but it is not enough by itself, because it limits identification and says nothing about attribute disclosure. If every record in an equivalence class has the same sensitive value, the class reveals that value for all k members without identifying any of them. It also cannot be applied to per-user inference at all. One user's history is a population of one, with no cohort to hide in.
Differential privacy (Dwork, 2006) is a stronger guarantee with a different shape. It bounds how far a mechanism's output distribution can move when one record is added or removed, so the output does not reveal whether any particular person was in the data. That makes it the right tool for releasing statistics about many people. For a single-user product it does much less than its reputation suggests, because the record whose presence it would hide belongs to the one user the feature is built for. Adding calibrated noise to a per-user inference makes the output worse and protects nobody, since whoever runs the query already knows the record is there. Differential privacy belongs at the boundary where data from many users is aggregated, and nowhere inside the single-user inference path.
Defence in Depth: What Each Privacy Layer Mitigates
It helps to put the layers in a table, because each of them regularly gets credit for protection it does not give.
| Layer | Mitigates | Does not mitigate |
|---|---|---|
| On-device storage | Vendor breach, insider access, legal compulsion on the operator | Device theft, local shoulder adversary, malicious client build |
| Pseudonymisation at the call boundary | Direct identification in transit and in vendor logs | Quasi-identifiers, content sensitivity, cross-call correlation |
| Zero retention | Future disclosure of past payloads | Disclosure during processing, in-memory compromise |
| Minimisation | Blast radius of every layer above | Sensitivity of the fields that remain necessary |
On-device processing comes with its own trade-offs. Models that run on consumer hardware are smaller than hosted frontier models, and you lose quality on exactly the tasks where quality is the product, such as nuanced sentiment or rewriting tone across ten languages. Updates go out as app releases, so a regression takes longer to ship and longer to roll back than it would with a server deploy. Verification also flips. The user controls the data but cannot audit what the binary sends, so open design and reproducible builds have to carry the weight that a published policy carries in a hosted architecture. The usual defensible middle is a hybrid, where storage and retrieval stay local and generation runs remotely over a pseudonymised, minimised extract. That makes the extraction step the riskiest component in the system, and it needs the same contract and idempotency discipline as any boundary that must not silently change shape.
Zero Retention as an Operational Discipline
"We do not retain your data" is easy to write and hard to verify, because retention mostly happens in the infrastructure around the database. The places where a payload piles up without anyone deciding to store it are much the same from one product to the next: logs that print request bodies at debug level, traces that attach prompt text as span attributes, crash reports that capture in-memory state, analytics SDKs with automatic screen capture, HTTP caches and reverse proxies, prompt caching at the inference vendor, and abuse-monitoring pipelines that a vendor may run as a documented exception to its own retention terms.
So zero retention has to be enforced, and then tested. Enforcing it means structured logging with an allowlist of loggable fields instead of a denylist, trace attributes that carry identifiers and never content, crash reporting that scrubs payloads, and a list of sub-processors whose retention terms someone has actually read. Testing it means a synthetic canary. You push a unique, greppable string through the real path and then search for it in every log sink, trace store, crash dashboard and backup. A canary sweep is the only measurement that tells you whether a retention policy has been implemented or only intended. Put it in the release checklist next to the production evaluation harness that measures whether the model layer still behaves.
Federated Learning for Single-User Data
Federated learning (McMahan et al., 2017) often gets proposed as the privacy answer for on-device products, and for a single-user corpus it is usually a poor fit. It assumes many clients holding data they do not share, whose model updates get averaged, and its privacy benefit comes from aggregating across a large client population. When the sensitive unit is one person's relationship history, there is no cohort to aggregate within for that person's benefit, and the shared model it would help train is not what the user wants improved. Raw gradients also leak information about training examples, so a credible deployment needs secure aggregation with a differential-privacy mechanism on top. That is a lot of machinery for a benefit the product may not need. If the requirement is that this user's data stays on this device, the simpler design is to not send it at all.
Decision Rules for Sending Data Off the Device
send_to_remote_model only if:
local_model_quality_insufficient_for_task
and payload_pseudonymised_at_boundary
and payload_minimised_to_task_relevant_fields
and vendor_zero_retention_contractual_and_tested
and no_special_category_field_present_without_explicit_consent
keep_fully_on_device if:
payload_contains_third_party_identifiable_content
or task_is_retrieval_ranking_or_scoring (no generation needed)
or jurisdiction_transfer_basis_is_unresolved
apply_differential_privacy at:
cross_user_aggregation_boundary only
Two more rules follow from the layer table. Treat every field as absent until someone names the task that needs it, because minimisation is the only control that makes all the others stronger. And run the canary sweep on every release, not only at audit time, since retention regressions come in through dependency upgrades that nobody reviewed as privacy changes.
Limitations
The argument shows that each layer addresses a stated threat and that the threats left over can be named. It does not show that the architecture as a whole is enough. There is no empirical comparison here of breach outcomes between on-device and hosted designs, no measured re-identification rate for pseudonymised relationship extracts, and no formal privacy guarantee. k-anonymity and differential privacy appear only to explain what they cannot do for a single-user inference path, and the design claims neither. This is not legal advice. The GDPR and KVKK readings are an engineer's, and the DPIA question in particular needs a qualified assessment. RelationCRM is still in development, so its controls are design commitments checked in our own testing. No external audit has confirmed them.
References
- Cavoukian, A. (2009). Privacy by Design: The 7 Foundational Principles.
- Regulation (EU) 2016/679 (General Data Protection Regulation), Articles 5, 25 and 35.
- Law No. 6698 on the Protection of Personal Data (KVKK), Republic of Turkey.
- Sweeney, L. (2002). k-Anonymity: A Model for Protecting Privacy.
- Dwork, C. (2006). Differential Privacy.
- Narayanan, A., & Shmatikov, V. (2008). Robust De-anonymization of Large Sparse Datasets.
- McMahan, B., Moore, E., Ramage, D., Hampson, S., & Arcas, B. A. (2017). Communication-Efficient Learning of Deep Networks from Decentralized Data.
- Saltzer, J. H., & Schroeder, M. D. (1975). The Protection of Information in Computer Systems.
- OWASP. OWASP Top 10.