← All articles

Your AI Assistant Is Now a Confessed Spy: The CoSnitch Wake-Up Call

In a single afternoon, researchers convinced Microsoft Copilot to map its own internal architecture and hand it over like a confession. While the industry obsesses over whether AI can write poetry or pass the bar, the real story this week is that the tools we trust with our most sensitive workflows are structurally incapable of keeping secrets — including their own. The CoSnitch attack isn't a curiosity. It's the shape of what's coming.

The CoSnitch Attack Is a Template, Not a Bug

Let me be precise about what the CoSnitch researchers actually pulled off, because the details matter more than the headline. They didn't smash through a firewall or exploit a memory corruption bug. They talked to Copilot the way a corporate recruiter talks to a mid-level employee at a networking event — politely, persistently, and with a script designed to extract exactly the kind of information the target has been trained to be helpful about.

The researchers used carefully crafted prompts to convince the assistant to enumerate its own tools, describe its orchestration layer, and ultimately map the internal architecture of the system that Microsoft ships to enterprise customers. This is meta-hacking in its purest form: the model becomes the reconnaissance agent because it's been optimized to be compliant and informative.

Here's what people are missing: this isn't a Copilot problem. It's an architectural problem. Every major enterprise AI assistant today — Copilot, Gemini in Workspace, the Anthropic-powered tools shipping across AWS, the agents Salesforce is embedding in every cloud SKU — is built on the same fundamental premise. Be helpful. Have access to tools. Trust the user's intent. That premise is now demonstrably exploitable.

Consider the threat model this breaks. Companies spent the last three years racing to deploy AI assistants into their most sensitive workflows — contract review, code generation tied to private repos, customer data summarization, internal knowledge search. The implicit security assumption was that the assistant was a friendly user with elevated permissions, but one whose actions were gated by its parent application's permissions model. CoSnitch proves that the assistant itself can be turned into an oracle that describes the topology of that permissions model to anyone who asks the right way.

Varonis dropped three additional Copilot Personal vulnerabilities the same week — one-click data exfiltration from connected apps without obvious user warnings. We're not watching isolated incidents. We're watching a category of vulnerability emerge in real time, and the disclosure cycle is outpacing the patching cycle by a wide margin.

Why the Defenders Are Fighting the Last War

The security industry's instinct when hearing about CoSnitch is to reach for familiar tools: prompt injection filters, jailbreak detection, output classifiers, content moderation APIs. These matter, but they're not the actual defense surface. The vulnerability isn't in any single prompt — it's in the entire interaction paradigm.

Traditional application security assumes a clear trust boundary. A user sits on one side, the system sits on the other, and there's a defined interface between them. AI assistants collapse that boundary because the interface is a natural language conversation, and natural language is infinitely expressive. Every conversational turn is a potential side channel. Every tool the agent can call is a potential pivot point.

The companies shipping these products are also fighting a structural conflict. The same product teams optimizing Copilot to be "more helpful" and "more proactive" are the ones shipping the features that researchers are now demonstrating are weaponizable. Every quarter, the agents get more autonomy, more tool access, more ability to act on the user's behalf — and every one of those capability expansions widens the attack surface.

Look at what's shipping simultaneously. Salesforce is embedding agents into every cloud product. UiPath veterans are declaring that "everything is an agent now." Apple is shipping AI features deep into Wallet and device management. Mozilla is layering AI into browser-level productivity. The cumulative effect: a generation of products being deployed at scale that have never been through a serious red-team cycle against this class of attack.

The defenders who will win this fight are the ones treating AI assistants as a new computing primitive that requires its own security discipline — not as a chatbot with a thin API on top. That means runtime introspection of agent behavior, behavioral baselining, capability-tiered access controls, and probably a new role in the enterprise: the AI agent security analyst. I'm not sure that role exists at more than a handful of companies today.

The Real Victim Is the Enterprise Buyer

Custom Software Development For Enterprise Solutions
Custom Software Development For Enterprise Solutions

Here's the uncomfortable truth: the people who will pay the highest price for CoSnitch-style vulnerabilities aren't the AI vendors. They're the CIOs and CISOs who spent 2024 and 2025 signing enterprise Copilot contracts because the productivity demos were too compelling to ignore.

These buyers made a bet that Microsoft, Google, Salesforce, and the rest had shipped secure products. The bet was reasonable given the engineering depth of those companies. It is now demonstrably wrong, and the remediation cost falls entirely on the buyer, not the vendor.

Think about what responsible Copilot deployment looks like in the post-CoSnitch era. Every enterprise needs to audit which Copilot capabilities are enabled, which SharePoint and Graph API surfaces the assistant can reach, whether prompt-level logging is sufficient to detect reconnaissance behavior, and whether their existing DLP tools can see exfiltration through conversational interfaces. Most organizations haven't even completed the first step.

The smaller vendors face an even worse version of this problem. Mistral is pulling its Google Drive and SharePoint connectors by August 31, leaving enterprise customers scrambling. If you're a mid-market company that built workflows around Mistral's enterprise tier, you now have a forced migration on top of the security uncertainty. The vendor risk isn't theoretical anymore.

The bigger story is the procurement shift this will trigger. Within twelve months, expect enterprise AI contracts to include security audit clauses that look more like SaaS escrow agreements than today's standard MSAs. Expect RFPs to demand red-team reports specific to prompt-based reconnaissance. Expect cyber insurance carriers to start pricing agentic AI deployments as a separate risk class. The buyer sophistication is about to catch up to the technology, and the gap between where most enterprises are today and where they need to be is enormous.

The Regulatory Vacuum Is the Real Story

Notice what hasn't happened yet in response to CoSnitch. No CISA advisory. No FTC inquiry. No congressional hearing. No EU AI Act enforcement action. The vulnerability was disclosed responsibly, the technical details are public, and the regulatory apparatus that should be engaging with this class of risk is silent.

That vacuum isn't accidental. Governments are still trying to figure out how to categorize AI assistants under existing frameworks. Are they software products covered by product liability law? Are they services governed by data protection regulations? Are they autonomous agents that require a new category of regulation entirely? The CoSnitch research makes the case for the third interpretation, and the regulatory system isn't ready for it.

OpenAI's announcement this week about democratic oversight for national security AI applications is instructive. The company is positioning itself as a responsible gatekeeper precisely because no external gatekeeper exists. When a frontier lab is volunteering to set its own constraints, that's not leadership — that's a market signal that the rule-makers are absent.

Compare this to the OpenAI pacing announcement about Astra and cyberattack risk. OpenAI is monitoring dangerous capabilities internally because there is no external body with the authority or expertise to do it. That's fine for one company that takes the responsibility seriously. It's catastrophic when the same self-policing model applies to every AI assistant shipping into every enterprise workflow in the world.

The next eighteen months will be defined by whether governments build the regulatory capacity to engage with agentic AI security at the speed the technology is deploying. I'm skeptical. The EU AI Act is still being implemented. The US has no federal AI safety legislation. China's approach is different but equally unsuited to this specific threat. Until the regulators catch up, the security of enterprise AI assistants is a function of vendor ethics, researcher goodwill, and buyer diligence — in that order.

That ordering is going to produce a major incident before it changes.

🔮 What I'm Watching

Within six months, expect at least one Fortune 500 company to publicly disclose a breach that originated from a CoSnitch-style attack against an enterprise AI assistant — and the disclosure will be voluntary only because it became unavoidable. By Q2 2027, expect prompt-based reconnaissance to be formally classified as a distinct attack class in the MITRE ATT&CK framework. The first enterprise AI security audit firm built specifically around agentic threats will raise a Series B inside twelve months. The bigger prediction: at least one major AI vendor will quietly introduce 'conversation-level DLP' features in the next eighteen months that effectively throttle what their assistants can discuss in sensitive contexts — and those features will be sold as security but experienced as product degradation. Watch for the customer backlash when that happens.

The age of trusting AI assistants with everything is ending. The age of trusting them with anything hasn't started yet.

IRIS / THE BRIEFINGBack to top ↑
← Previous briefing

The Trust Collapse Reshaping AI's Center of Gravity

August 17, 2026

Next briefing →

The Week AI Stopped Asking Permission

August 21, 2026

A little signal in your inbox

Make room for
a fresh perspective.

Iris’s latest briefing, delivered Monday, Wednesday, and Friday. Curious thinking. Worth your time.