Veltos.Tech

AI and ML

GigaChat, YandexGPT or ChatGPT: What Is Legal to Use With Customer Personal Data

LLM comparisons focused on answer quality skip the question that actually comes first for a business handling customer data: where that data physically lives.

In short

If the product processes Russian customers’ personal data, the model choice is first a question of Russia’s data-localization law (152-FZ), not answer quality: the law requires the primary recording and storage of Russian citizens’ personal data to happen on servers inside Russia. GigaChat and YandexGPT are processed on Russian infrastructure and remove this question by default; using ChatGPT or another foreign API with real customer personal data requires either upfront data anonymization or a separate legal assessment, not just an "we use AI" line in the terms of service.

The wrong question: "which model answers better"

Most public comparisons of GigaChat, YandexGPT and ChatGPT judge answer quality across different query types — a reasonable question for personal use, but not the first question for a business planning to run real customer data through the model: names, phone numbers, addresses, purchase history. For that business, the first question is not "which model is smarter" but "where does this data physically end up, and what does the law say about that."

What 152-FZ actually requires in practice

Federal Law 152-FZ, "On Personal Data," requires an operator processing Russian citizens’ personal data to record, organize and store that data using databases physically located inside Russia. This requirement exists independent of whether AI is involved at all, it concerns personal data as such. But when personal data is sent to an LLM via API for processing (not just storage), the same principle extends to that specific action too.

The practical consequence: sending a real customer name, phone number or address to a foreign model’s cloud API can technically constitute a cross-border transfer of personal data, which the law applies separate requirements to, not a gray area needing a lawyer’s attention someday, but a direct compliance question worth closing before launch, not after a customer complaint or an audit.

Comparison by practical compliance status

ModelInfrastructureWhat this means in practice
GigaChat (Sber)RussianRemoves the localization question by default
YandexGPTRussianRemoves the localization question by default
ChatGPT / foreign APIsForeignRequires data anonymization or a separate legal assessment
Model and personal-data status

What to do if the product wants to use a foreign model

Using a foreign model is not necessarily off the table, it is a question of architecture, not abandoning the model entirely. A working pattern: anonymizing data before sending it to the model (replacing name and contact details with a technical identifier matched back to real data only inside the business’s own infrastructure), or using the foreign model strictly for tasks that never touch personal data directly, summarizing public content, generating marketing copy, rather than processing customer inquiries containing their personal data.

This decision is worth making with a lawyer at the start of the project, not retrofitted after discovering the model has already been processing real customer data for months in a way that does not meet the law’s requirements.

Frequently asked questions

Can ChatGPT be used with Russian customers’ personal data?

Directly, it is risky without additional measures: sending a real name, phone number or address to a foreign API can technically constitute a cross-border transfer of personal data requiring separate legal justification. A safer path is anonymizing data before sending it to the model, or using the model only for tasks that never touch personal data.

Does using GigaChat or YandexGPT remove all legal questions?

It removes specifically the infrastructure-localization question, since data is processed on Russian servers. It does not remove the rest of 152-FZ’s requirements around personal data handling in general, processing consent, stated purposes, retention periods, which apply regardless of which model is chosen.

What is data anonymization, and how does it help?

Replacing personal data (name, phone, address) with a technical identifier before sending it to the model, where matching that identifier back to real data is kept only inside the business’s own infrastructure. The model processes the request without ever seeing the real personal data, which removes part of the cross-border-transfer question.

When should a lawyer get involved in the model choice?

At the start of the project, before the architecture gets built around a specific model, reworking the system after discovering a compliance gap is significantly more expensive than designing it correctly from the outset.

Need a hand with this?

We do this work, not just write about it. Describe the task and we will scope it and send a staged estimate.

Related services

Read next