Blog

Are AI Assistants Safe? Data Privacy for Owners

July 29, 2026 · Privacy & self-hosting · 10 min

By , founder of Raegan

A laptop on a quiet desk at dusk, suggesting an AI assistant used with data kept under the owner's control.

AI assistants can be safe, but most consumer ones are not safe by default. Safety comes down to three things you can check: whether your inputs are used to train shared models, where your data physically lives, and who at the vendor can read it. A private AI assistant is the version built to answer "no" to the first risk and "you control it" to the rest. The technology is not the danger. The data contract underneath it is.

TL;DR: Whether an AI assistant is safe depends on its data terms, not its features. In Cisco's 2025 Data Privacy Benchmark Study, 64% of organizations worried about leaking sensitive information through generative AI, yet nearly half admitted entering employee or non-public company data into it anyway. A private AI assistant closes that gap structurally: no training on your inputs, data residency you can verify, a self-hosted option, and an approval gate before anything customer-facing goes out.

This is not a fear piece. AI assistants save real hours, and the evidence that they leak data is not a reason to avoid them. It is a reason to choose the right kind. The owners who get burned are not the ones who used AI. They are the ones who pasted a contract, a client list, or a payroll file into a free public tool with broad data-use terms and assumed it was private. Below is what "safe" actually requires, what the real risks are, and a checklist you can use before connecting any assistant to your inbox.

Are AI assistants safe to use for business?

AI assistants are safe to use for business when three conditions hold: your data is not used to train shared models, it stays in infrastructure you can locate, and a human approves anything that leaves your control. Most free consumer tools fail at least the first condition. The risk is not the AI doing the work. It is the standard terms that let your inputs become someone else's training data or sit in a shared cloud pool.

The numbers show the gap clearly. Cisco's 2025 Data Privacy Benchmark Study, which surveyed 2,600 privacy and security professionals across 12 countries, found that 64% of organizations worry about inadvertently sharing sensitive information publicly or with competitors through generative AI. Yet nearly half admit to entering personal employee or non-public company data into those same tools. Concern is high and behavior contradicts it, which is exactly the condition under which accidents happen.

So "are AI assistants safe" is the wrong question. The useful question is "is this assistant private," and that has a definite, checkable answer. A private AI assistant is the category built to make the answer yes.

What are the real data privacy risks of AI assistants?

The real risks are four: your data trains a model you do not control, it is retained on shared servers, it surfaces in a breach, and it leaks through tools no one approved. These are concrete and measurable, not hypothetical. Understanding each one is how you tell a marketing claim of "secure" from an assistant that is actually safe to put your business inside.

Training on your data. Many public tools reserve the right to use your inputs and outputs to improve their models. Once a contract or client email is absorbed into training data, you cannot pull it back. Stanford's 2025 AI Index Report flags this directly, listing unauthorized data access during training and the persistence of personal data long after it should be deleted among the core privacy risks of modern AI systems.

Retention and residency. Even without training, your conversations are stored somewhere. If you cannot say which region they sit in or how long they are kept, you cannot answer for them to a client or a regulator. Cisco found a revealing split: 90% of organizations believe storing data locally is inherently safer, while 91% trust global providers to protect it better. Both instincts are about the same thing, knowing where the data lives.

Breaches. Storage is exposure. IBM's Cost of a Data Breach Report 2025 put the global average breach at $4.44 million, and tied breaches involving high levels of shadow AI to an extra $670,000 in cost, with shadow AI a factor in roughly one in five breaches.

Shadow AI. Most leakage is quiet. It happens when an employee uses an unapproved tool on the side, outside any policy. That is the fastest-growing risk of the four, and the hardest to see.

The concern-behavior gap in business AI use Bar chart comparing organizations that worry about leaking sensitive data through generative AI, 64 percent, against the share that nonetheless enter employee or non-public data into it, nearly 50 percent, per Cisco's 2025 Data Privacy Benchmark Study. Worry is high. Behavior contradicts it. Share of organizations, generative AI use (Cisco, 2025) 64% Worry about leaking data ~48% Enter employee / non-public data Source: Cisco, 2025 Data Privacy Benchmark Study (n=2,600, 12 countries).
Nearly half of organizations feed sensitive data into generative AI even though most fear exactly that leak. Source: Cisco, "2025 Data Privacy Benchmark Study."

How big is the shadow AI problem?

Shadow AI is now a leading and measurable cause of AI-related data loss. It happens when staff use unsanctioned AI tools, so the data goes somewhere the business never vetted. IBM's Cost of a Data Breach Report 2025 found that breaches involving high levels of shadow AI cost about $670,000 more than those without, and that 97% of organizations hit by an AI-related incident lacked proper AI access controls. The exposure is not a tail risk. It is the common case.

What makes shadow AI dangerous for owners specifically is what it tends to expose. IBM found shadow AI incidents disproportionately compromised customer personal information, 65% of the time versus a 53% global average. For a small company, that is the trust you sell. And the volume of sensitive data in play keeps rising: Cyberhaven Labs, measuring real usage across 7 million workers, found 34.8% of the corporate data put into AI tools in 2025 was sensitive, up from 27.4% the year before and 10.7% two years earlier.

The scale of AI risk overall is climbing too. Stanford's 2025 AI Index Report recorded 233 AI-related incidents logged in the AI Incident Database in 2024, a record, while public trust that AI companies will protect personal data fell from 50% in 2023 to 47% in 2024. A ban is one response, but it pushes usage further into the shadows. A private assistant is the version people will actually adopt instead of working around.

What shadow AI adds to a data breach Comparison showing the global average data breach cost of 4.44 million dollars versus 5.11 million dollars when high levels of shadow AI are involved, an increase of about 670 thousand dollars, per IBM's Cost of a Data Breach Report 2025. Shadow AI adds about $670K to a breach Average data breach cost, USD (IBM, 2025) $4.44M Global average ~$5.11M High shadow AI Source: IBM, Cost of a Data Breach Report 2025 (with Ponemon Institute).
Ungoverned, unapproved AI use is not a rounding error. It is roughly a $670,000 swing on the average breach. Source: IBM, "Cost of a Data Breach Report 2025."

What does a genuinely safe AI assistant require?

A genuinely safe AI assistant requires four guarantees: no training on your data, verifiable data residency, a self-hosted or dedicated option, and an approval gate on anything that leaves your business. Each one answers a specific risk above. Together they turn "trust us" into something you can actually inspect. Treat the absence of any one as a reason to keep the tool away from sensitive work.

No training on your inputs. This must be in writing, not implied. It is the difference between your data doing a job and your data becoming a permanent part of someone else's product.

Data residency you can verify. You should be able to name the region your data sits in and how long it is retained. This is also where the regulatory wind is blowing. Gartner predicts that by 2027, 35% of countries will be locked into region-specific AI platforms using in-region data, up from around 5% today, driven by data-sovereignty rules. The pressure pushing nations toward sovereign AI is the same pressure pushing owners toward private assistants.

A self-hosted or dedicated option. The strongest privacy is enforced by physics, not policy. When the assistant runs on a self-hosted AI assistant setup or dedicated infrastructure, your data is not pooled with other customers at all. The trade-off between control and convenience is worth weighing directly, which is the subject of self-hosted vs managed-private AI assistant.

An approval gate on outbound. Privacy is about more than where data sits. It is also about what the assistant is allowed to do unsupervised. An assistant that can send a customer email on its own is a liability. One that drafts and waits for your approval keeps you in control of the one thing that is hardest to undo.

Raegan is built this way. It runs self-hosted on Hermes, an open-source agent from Nous Research, on the customer's own server, so data is never sold, shared, or dropped into a common cloud pool, and nothing customer-facing sends without the owner's approval.

A practical data privacy checklist for owners

Use this checklist before you connect any AI assistant to your inbox, files, or customer data. The aim is plain answers. If a vendor cannot give a clear response to the first three questions quickly, treat the tool as public no matter how it markets itself.

Run an assistant through these seven questions and "is it safe" stops being a feeling and becomes a record you can keep. Safe is not a label a vendor applies. It is a set of answers you collected before you handed over the keys.

Frequently asked questions

Are AI assistants safe to use?

AI assistants are safe when their data terms are, not by default. The decisive factors are whether your inputs train shared models, where your data is stored, and who can read it. Free consumer tools often fail the first test. A private AI assistant with no-training terms, verifiable residency, and an approval gate is the version built to be safe for business use.

Do AI assistants train on my data?

Many public consumer tools reserve the right to use your inputs and outputs to improve their models, often by default unless you opt out. A genuinely private assistant does not train on your data at all, and states so in writing. Always confirm this before connecting an assistant to contracts, client lists, or anything non-public, since the term "private" is used loosely.

What is shadow AI and why does it matter?

Shadow AI is the use of AI tools that no one at the company approved or knows about. It matters because it routes business data into unvetted services. IBM's 2025 report linked high levels of shadow AI to roughly $670,000 in added breach cost and found it disproportionately exposed customer personal information, 65% of incidents versus a 53% average.

Is a self-hosted AI assistant safer than a cloud one?

A self-hosted assistant is generally the strongest privacy option because your data physically stays on infrastructure you control and is never pooled with other customers. A well-run managed-private service can be safe too, with contractual no-training and residency guarantees. The weakest option is a public cloud tool with broad data-use terms, regardless of how it is marketed.

How do I know if an AI assistant is genuinely private?

Ask three questions and require clear answers: is my data used for training, where does it physically live, and who can access it. A genuinely private assistant says no to training, names the location, and lets you keep data on your own or dedicated infrastructure. Vague answers mean treat it as public, whatever the marketing says.

Can I use public AI tools safely if I am careful?

Careful use reduces but does not remove the risk, because the tool's data terms and storage do not change with your good intentions. Cyberhaven found 34.8% of corporate data entered into AI tools in 2025 was sensitive. The structural fix is a private assistant, which removes the leak path rather than relying on every person behaving perfectly every time.

Sources

  1. Cisco. "2025 Data Privacy Benchmark Study," 2025. https://newsroom.cisco.com/c/r/newsroom/en/us/a/y2025/m04/cisco-2025-data-privacy-benchmark-study-privacy-landscape-grows-increasingly-complex-in-the-age-of-ai.html
  2. IBM. "Cost of a Data Breach Report 2025," 2025. https://www.ibm.com/reports/data-breach
  3. Cyberhaven Labs. "2025 AI Adoption and Risk Report," 2025. https://www.cyberhaven.com/press-releases/cyberhaven-report-majority-of-corporate-ai-tools-present-critical-data-security-risks
  4. Stanford HAI. "The 2025 AI Index Report, Responsible AI," 2025. https://hai.stanford.edu/ai-index/2025-ai-index-report/responsible-ai
  5. Cisco. "2024 Data Privacy Benchmark Study," 2024. https://newsroom.cisco.com/c/r/newsroom/en/us/a/y2024/m01/organizations-ban-use-of-generative-ai-over-data-privacy-security-cisco-study.html
  6. Gartner. "Gartner Predicts 35% of Countries Will Be Locked Into Region-Specific AI Platforms by 2027," 2026. https://www.gartner.com/en/newsroom/press-releases/2026-01-29-gartner-predicts-35-percent-of-countries-will-be-locked-into-region-specific-ai-platforms-by-2027

Raegan is a private, self-hosted AI assistant for owners who want the help without giving up their data. Get early access.

Meet Raegan, your private AI chief of staff

She handles email, notes, research, calendar, and a daily briefing - private and self-hosted. Join the early-access waitlist.

Get early access