Research

What public breaches actually show about who controls your data

Aug 19, 2026 · updated Aug 21, 2026 · privacy, security, data ownership

Public incidents show that data risk extends past the company you chose, to its credentials, employees, contractors, dependencies, connected devices, and legal obligations.

Abstract

A review of public breach disclosures, regulator actions, government reports, company notices, and AI privacy documentation shows a recurring pattern: when data is held or processed outside a user's direct control, the security and privacy boundary expands to include more systems, more people, more credentials, more dependencies, and more legal obligations. This does not prove that hosted infrastructure is inherently unsafe. It does show why ownership, explicit permissions, data location, portability, and controlled connectivity are meaningful product properties rather than marketing ones.

Notes

I kept running into the same argument from two directions. One side says the cloud is dangerous. The other says owning your own infrastructure is a fantasy for people who like pain. Both are marketing positions, and neither one is evidence.

So I went and read the primary sources instead: the regulator orders, the securities filings, the company incident notices, the provider privacy documentation. What follows is what those documents actually support, and, just as importantly, what they do not.

Physical ownership does not guarantee data control

Connected products create information that is stored or processed somewhere other than the device itself. The FTC’s case against Ring involved private camera footage, employee and contractor access, and account security. Wyze separately confirmed that a software failure caused some users to be shown thumbnails from cameras that were not theirs.

The connected-vehicle example is the same shape. The FTC’s final order involving GM and OnStar concerns precise geolocation and driving-behavior data. The point is not that connected products are unsafe. It is that owning the physical object does not by itself establish who controls the data pipeline attached to it.

Genetic data is the most durable version of the problem. The UK’s Information Commissioner’s Office fined 23andMe after its 2023 breach, and later called for protections around customer information during the company’s bankruptcy and sale. Unlike a password, genetic information cannot be rotated after exposure.

The security boundary extends past the company you chose

AT&T disclosed that threat actors accessed an AT&T workspace hosted on a third-party cloud platform and exfiltrated customer call and text interaction records. The filing did not say that the underlying cloud provider had been breached.

Live Nation disclosed unauthorized activity in a third-party cloud database environment holding company data, primarily Ticketmaster data. Snowflake has separately stated that attackers accessed customer accounts, while saying it found no evidence that the incidents resulted from a vulnerability, misconfiguration, or breach of the Snowflake platform itself.

CERT-EU’s analysis of a 2026 European Commission incident is the supply-chain version: initial access was attributed to a compromised software dependency that exposed an AWS secret. The accurate description there is a compromised Commission cloud account and credential, not “Amazon was hacked.”

This distinction is the one most often lost in a headline. A cloud-hosted incident, a cloud-account compromise, and a cloud-provider breach are three different events.

Encryption is one layer, not the whole system

LastPass disclosed a sequence of incidents that began in its development environment and later reached cloud backups. Endpoints, employees, credentials, keys, backups, and metadata all stayed relevant even though sensitive vault fields were encrypted under a zero-knowledge design.

The lesson is not that encryption failed as a concept. It is that encryption protects a field, and an incident happens to a system.

Authorized people are part of the threat model

Coinbase disclosed that a threat actor paid multiple contractors or employees in support roles to collect information from internal systems they were already authorized to access. Odido said its 2026 incident began with voice phishing against customer-service staff.

If a person or service can legitimately reach the data, that path belongs inside the threat model. Support workflows, identity verification, and contractor access are security surfaces, not operations trivia.

Centralization concentrates blast radius

HHS reports that Change Healthcare reported approximately 192.7 million individuals impacted by its 2024 cyberattack as of July 31, 2025.

That number should not be used to argue that centralized services are inherently bad. It supports something narrower and more useful: when a widely depended-on service fails, a great many downstream organizations and people fail together. Resilience is partly a question of how much can go wrong at once.

AI conversations belong in this discussion

People increasingly use assistants as a place to think, which means prompts now carry material that used to live in private notes, internal documents, and code repositories.

Practices differ by provider and by account type. OpenAI says it does not train on business and API customer data by default. Google’s Gemini Apps Privacy Hub explains that some collected data can be reviewed by trained service providers, and warns users not to enter confidential information they would not want a reviewer to see.

Security is one axis, legal process is another. Reuters reported that researchers found a publicly accessible DeepSeek database containing backend information, software keys, and chat-related records that appeared to include user prompts. Separately, OpenAI publicly opposed a demand for 20 million consumer ChatGPT conversations in copyright litigation, and Reuters reported a 2026 ruling allowing a search warrant seeking an executive’s chatbot records.

None of that shows providers publishing private conversations. It shows that records held by a third party can become reachable through discovery or legal process, which is a different risk with a different mitigation.

What the evidence supports

  1. Every additional custodian, privileged user, account, dependency, integration, and remote service expands the set of things that must work correctly for data to stay private and available.
  2. Owning hardware is not the same as controlling the data that hardware produces.
  3. A cloud-hosted incident does not automatically mean the cloud provider was breached.
  4. User-controlled infrastructure reduces some third-party dependency, but does not eliminate endpoint compromise, software vulnerabilities, stolen credentials, misconfiguration, or physical risk.
  5. The honest argument is about control and choice, not fear.

What it does not support

It does not support saying that cloud services are unsafe, that large technology companies are untrustworthy as a class, or that infrastructure you own cannot be compromised.

It also does not support flattening an allegation, a settlement, a final regulatory order, a software failure, a company disclosure, a confirmed breach, an attacker claim, and a court order into one undifferentiated pile of “breaches.” Most writing on this topic does exactly that, which is why so little of it is worth citing.

Key findings

  • Third-party custody expands the number of systems and people that can become part of a data incident.
  • A cloud-hosted data incident does not require the underlying cloud provider itself to have been breached.
  • Connected devices can turn cameras, microphones, and vehicles into ongoing sources of sensitive behavioral data.
  • AI conversation records can be sensitive data, and can also become subject to legal process when a provider holds them.
  • User-controlled infrastructure reduces some third-party exposure, but does not remove the need for layered security, recovery, updates, and careful permissions.

Methodology

I worked from primary sources: regulator actions and final orders, securities filings, government cybersecurity bodies, company incident disclosures, and provider privacy documentation. High-quality independent reporting was used only where it added confirmation or legal context. Every example is classified by what kind of evidence it actually is, so that a confirmed breach, a regulatory allegation, a final order, a company disclosure, a software failure, a provider policy, and a court process are never treated as the same thing.

Limits

This is a curated set of public incidents, not a complete census, and it was assembled to test a specific question about custody and control. It does not measure how often these failures happen, it does not compare the security performance of any two named companies, and it does not establish that one infrastructure model is categorically safer than another. A larger or differently selected sample could change the emphasis, though the classification rule would still apply.

Sources

  1. HHS: Change Healthcare cybersecurity incident FAQ
  2. AT&T Form 8-K on the third-party cloud workspace incident
  3. CERT-EU analysis of the European Commission supply-chain incident
  4. FTC case against Ring
  5. Wyze incident update on the thumbnail failure
  6. FTC final order involving GM and OnStar
  7. UK ICO fine against 23andMe
  8. UK ICO on protections during the 23andMe sale
  9. LastPass security incident update
  10. Coinbase Form 8-K on insider-assisted access
  11. Live Nation Form 8-K on the third-party cloud database
  12. Snowflake Form 10-Q on customer account access
  13. OpenAI on business and API data commitments
  14. OpenAI statement on the request for consumer ChatGPT conversations
  15. Google Gemini Apps Privacy Hub
  16. Reuters on the DeepSeek database exposure
  17. Reuters on the 2026 chatbot records warrant ruling
  18. Odido security update on the voice-phishing incident

Take it further

Ask AI about this note

Opens your own assistant with the question filled in. You hit send.

Download this note as PDF