What I never paste into a chat window
I run AI tools every day. Scripts get drafted with them, logs get summarized with them, agents get wired into real infrastructure through MCP to move files, query databases, and touch production systems. None of that requires me to trust a chat window with everything I know. Over time I've settled into a short list of things that never go into a prompt, no matter how convenient it would be, and the reasoning behind each one is boring and mechanical, not paranoid.
The chat window is not a notebook
It's easy to treat a chat window like a scratchpad, because it feels private the same way a text editor feels private. It isn't. When you send a prompt to a cloud model, that text leaves your machine, goes to a server you don't control, and gets processed, logged, and in many cases retained for some period for abuse monitoring, debugging, or safety review. That's true even with providers that promise not to train on your inputs by default, because "not used for training" and "never stored anywhere" are different guarantees. Retention policies vary by vendor and by plan, and they change over time. I don't try to keep a running mental model of who currently retains what for how long, because that model would be out of date within a quarter. Instead I just don't paste the kind of thing that would matter if it leaked.
What actually never goes in
Credentials and secrets. API keys, database passwords, SSH keys, tokens, none of it goes into a prompt, ever, even when I'm asking an agent to debug a connection issue. If a script needs a key, the key lives in an environment variable or a secrets manager, and I ask the model to read the code that references the variable, not the value itself. This isn't about distrust of any specific vendor. It's that a credential pasted into a chat now exists in at least three places: my terminal history, the vendor's logs, and whatever caching layer sits in between. A credential that lives only in a vault has one place to rotate if something goes wrong.
Customer and user data. Names, emails, phone numbers, order details, anything tied to a real person who didn't consent to their data going to a third-party AI provider. When I need help debugging a data issue, I use synthetic rows or redact the identifying fields first. This matters for the same reason it matters in any other outsourced processing relationship: the data subject's consent covers the systems I've told them about, not every tool I decide to try that week.
Internal infrastructure details. Hostnames, internal IP ranges, firewall rules, the shape of my network. An agent debugging a server issue doesn't need the literal subnet to reason about the problem, it needs the symptom and the relevant config. Pasting a full network map into a prompt to save five minutes of abstraction isn't worth what it reveals about my attack surface if that text ends up somewhere I didn't intend.
Unreleased business terms. Draft contracts, pricing I haven't published, deal terms with partners who expect confidentiality. Not because I think a model is going to leak my numbers to a competitor on purpose, but because once text leaves my environment I've lost the ability to guarantee where it goes, and some of these documents come with confidentiality obligations to other people, not just to me.
Anything I wouldn't paste into a public support ticket. That's the actual test I use day to day. Support tickets get read by strangers, sometimes get forwarded, sometimes get quoted back in ways you didn't expect. If I wouldn't put a piece of information in a support ticket to a vendor I've never met, I don't put it in a chat prompt either.
Why agents change the math, not the rule
Running AI as an agent, wired into real tools through something like MCP, adds a layer people often miss: the model isn't just reading your prompt anymore, it's calling functions that touch real systems, and every one of those calls can get logged somewhere, by the tool server, by the orchestration layer, by the model provider, or by all three. That's more surface area for the same information to end up recorded, not less. The instinct to loosen up because "it's just an agent doing its job" is backwards. If anything, an agent connected to your file system, your email, or your database deserves a tighter rule about what you feed it directly in the prompt, because the blast radius of a leak is bigger when the agent already has legitimate access to sensitive systems.
The practical fix is the same one I use for secrets: give the agent the ability to read what it needs from the source (a file, a database query, an environment variable) rather than typing the sensitive value into the conversation yourself. The model still gets what it needs to reason and act. What changes is where the sensitive value actually lives and how many copies of it exist afterward.
Where local models actually help
Running a model locally does remove one link in that chain: your prompt never leaves your machine, so there's no vendor logging pipeline to worry about for that particular request. That's a real, mechanical difference, not a vibe. If I'm processing something genuinely sensitive and I have a local model capable enough for the task, using it instead of a cloud API is a straightforward reduction in where the data travels.
It's not a universal fix, though. Local models I can run on ordinary hardware are usually smaller and less capable than the frontier cloud models, so there's a real quality tradeoff, and for a lot of tasks that tradeoff isn't worth it. Local doesn't eliminate risk either: the machine itself can be compromised, logs still get written locally, and if that local model is exposed through an API to other tools on your network, you've just moved the trust boundary rather than removed it. The honest way to think about it is that local processing narrows the exposure to your own environment. It doesn't make the data disappear, and it doesn't turn a weak model into a strong one.
The rule I actually follow
I don't run a formal classification system before every prompt. I use one question: would I be comfortable if this exact text showed up, unredacted, in a support forum post six months from now? If the answer is no, it doesn't go in, and I find another way to get the model what it needs, whether that's a redacted version, a synthetic example, a reference to a file the agent can read directly, or just skipping the AI step for that particular piece of work and doing it by hand.
None of this is a claim that AI tools are unsafe to use for real work. I use them for real work constantly, including work that touches production infrastructure. It's a claim that the convenience of a chat window doesn't change what happens to text once you send it, and the systems that make AI genuinely useful for automation, local inference, agents, MCP-connected tools, are worth understanding well enough to know where the actual boundaries are, instead of guessing.
If you want more of how I actually run AI day to day, automation, local versus cloud tradeoffs, and agent tooling, without the hype, you can find the rest of it on [the home page](/).
Get new guides and videos first — join the Telegram channel.