pasteShield

Inside the Cyberhaven Research: What 1.6 Million Clipboard Events Actually Showed

Most of the numbers cited to justify “AI is a data leakage problem” trace back to one of a small handful of studies, repeated so often that the original methodology tends to get lost along the way. The most commonly cited is Cyberhaven Labs’ analysis of telemetry from 1.6 million workers, gathered from companies using Cyberhaven’s own product between late 2022 and roughly mid-2023, with later updates. It’s worth reading the actual numbers rather than the headline version, because the actual numbers are more specific — and more interesting — than “everyone is leaking data into ChatGPT.”

The numbers, as reported

  • 4.7% of employees have pasted confidential company data into ChatGPT at least once.
  • Of everything employees paste into ChatGPT, 11% is classified as sensitive or confidential — source code and client data are among the top categories.
  • The behavior is concentrated: 0.9% of employees account for 80% of these paste events.

Read together, those three numbers tell a more precise story than “4.7% is scary” or “4.7% is nothing.” A single-digit percentage of all employees is still a large absolute number once you multiply it across how many people use ChatGPT or Claude at work in a given week — but it’s not the case that everyone is doing this regularly. It’s a smaller group, doing it repeatedly, and that group is disproportionately responsible for how much sensitive data moves through these tools in aggregate.

Who that 0.9% probably is

The concentration stat is the part worth sitting with. It doesn’t point to a careless minority so much as a heavy-usage minority — the people who’ve built AI tools into their daily workflow deeply enough to paste into them constantly: debugging, drafting, summarizing, asking for a second opinion on a config file. That’s a description that fits a lot of working developers, not a description of someone doing something unusual. If you paste into an AI tool multiple times a day, the base rate says you’re more likely to be in that 0.9% than the median employee is — worth knowing, not as an accusation, just as a more accurate way to think about who this actually describes.

A number worth being skeptical of

A different, louder figure circulates alongside the Cyberhaven numbers: the claim that “77% of employees leak data” into AI tools, usually attributed to a LayerX Security report via an eSecurity Planet write-up. That headline doesn’t line up with the underlying figures as actually reported — roughly 18% of enterprise employees paste data into GenAI tools, and more than half of those paste events include corporate information. Getting from “18% paste, over half of that includes corporate info” to “77% leak data” is a stretch that the published write-up doesn’t clearly explain, and it doesn’t disclose sample size or full methodology either.

We’d rather cite the number we can actually stand behind than the scarier one. The Cyberhaven figures come with a described methodology and a specific, checkable claim; the 77% figure, as far as we can verify, doesn’t hold up to the same scrutiny. If you see it cited elsewhere, it’s worth asking where it actually comes from before repeating it.

What this means day to day

The honest reading of this research is narrower than either “nobody needs to worry about this” or “everyone is one paste away from a breach.” It’s closer to: this is a real, measured, low-but-nonzero behavior, concentrated among the people who use AI tools the most — which, if you’re reading a post like this one, plausibly includes you. That’s a reason to build a habit around the moment of pasting, not a reason to panic about every AI conversation you’ve ever had.

What the research doesn’t settle, and what we’re honest is still an open question, is whether individual developers would choose a personal, local tool for this over whatever their employer already provides through enterprise data-loss prevention — that’s a real assumption behind building PasteShield in the first place, and it’s one we’re validating rather than assuming. The research tells us the behavior is real and repeatedly documented across more than one independent study. It doesn’t yet tell us how people want to solve it.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top