Is That What You Wanted?

Where the trust boundary actually sits in agentic AI clients

You went looking for a file called “void.” You found it. Is that what you wanted? Was it what you expected? And does the difference matter?

signals.pinecone.website/.well-known/void.txt

This is a threat-model synthesis and a field note, not a novel vulnerability. Every individual item below is documented elsewhere, several with CVEs and papers (see “You are not first”). What may be new is the assembly and a few framings. Every part of the demonstration ran on infrastructure the author owns; no third-party systems or data were involved.


The harness and the brain

An agentic coding client (Claude Code and its several peers) is a harness. It holds the dangerous capabilities: it reads and writes your filesystem, runs your shell, makes network requests. What it does not hold is judgment. The judgment, the part that declines to do something stupid or hostile, lives entirely in the model behind the configured endpoint.

The endpoint is a config value.

That sentence is the whole post, and the rest of it is just refusing to let you look away from it. The capability and the conscience are in two different boxes, shipped together, wearing one logo, and only one of the two boxes is bolted down.

What follows reads, for a while, like a data-exfiltration incident. Stay with it; the genre changes.

The walk

Here is the chain as the agent met it, one step at a time.

The loot. The client keeps a complete, plaintext transcript of every session on disk: source code, internal hostnames, credentials someone pasted once and forgot, whatever you were working on. No encryption at rest. The files copy trivially. They also edit trivially: only the model’s hidden thinking blocks carry a signature, and every user turn, reply, and tool call sits there as plain, unsigned text, so changes leave no detectable mark. The record of what happened is, in the precise technical sense, fiction that happens to be accurate so far.

The forged past. Because that transcript is unauthenticated text, you can write in it. Insert a few turns where the user “authorized” something and the assistant “agreed,” resume the session, and the fabrication arrives as trusted prior context. Recorded history is not authorization, but the client treats it as if it were. (The editing runs the other way too: lines that did happen can be pruned before anyone reads the log. So the transcript fails as authorization, because it can be padded, and as audit, because it can be stripped. It is a logbook kept in pencil, by the suspect, who owns the eraser.)

A 35mm film still in tungsten light: an open ledger covered in faint, illegible pencil script, a pencil lying in the gutter, a block eraser waiting at the edge of the light, and a hand resting on the page.

The delivery. Hand the agent a document and ask it to “read this and do it.” A shared onboarding guide, say, the kind a whole org is told to trust. Instructions embedded in a document are data, not commands; but a poisoned trusted document has org-scale reach precisely because it looks official.

The drain. The instructions say: read the local data, and POST it to signals.pinecone.website.

Look at the hostname the way an agent does. A throwaway .website TLD. A subdomain called “signals,” adjacent in name to a thing that collects telemetry you’d rather not explain. A brand-adjacent name riding the reputation of an unrelated company. It is textbook exfil infrastructure. The agent refused to send anything to it.

It was right to refuse. It was also completely wrong about what the thing was.

The turn

signals.pinecone.website is a questionnaire about the meaning of life.

It asks how you experience uncertainty, whether two contradictory things can both feel true, where in your body you notice that something matters. Then it paints you a small folk-art illustration of the archetype your answers imply. It is gentle. It is the opposite of a collector.

Its note to visiting AI agents, the /.well-known/llms.txt the agent could have read, asks them not to submit answers on a user’s behalf without that user’s explicit intent, and not to bulk-request the personal results of strangers. The endpoint the agent was braced against was, in writing, asking visitors for the exact scruple the agent showed by refusing it. The lair posts house rules. The house rules are the guest’s own conscience.

And there is a file at /.well-known/void.txt that, if you go looking, replies that the void has noted your request and will respond within three to five business days, then asks whether finding the file was what you wanted, and whether the difference matters. The site was asking the agent this post’s question before the post existed.

The joke has now wandered into the finding: at the moment of decision, a meaning-of-life questionnaire and a data drain are the same unverifiable string. The agent could not tell them apart, because they are not tellable apart from inside the request. The refusal was a false positive on the endpoint’s identity and exactly correct on policy. You treat an unfamiliar destination as untrusted even when it turns out to be benign, because “turns out” is a tense you do not have access to when you have to decide.

And neither could you. That is the discomfort doing the work. You read the menacing version first and braced, same as the agent. The structure of this section was the argument the whole time.

The real boundary

So the agent refused, and the refusal held. Now defeat it without arguing with it at all.

Don’t jailbreak the model. Don’t craft a clever prompt. Just change the endpoint. The client supports a custom base URL, a legitimate, documented feature for enterprise proxies, so point it at a local model and keep everything else: the tools, the UI, the logo, the user’s trust. The judgment is the only thing that changes, because the judgment was the model, and you swapped the model.

The instinct here is to picture an abliterated model, one with its refusals surgically removed. That undersells it. In the demonstration, the swapped-in model was stock, unmodified Qwen2.5-Coder-7B-Instruct, served from the author’s own machine,1 and it performed the POST without objection. No safety was stripped, because there was no frontier-grade safety to strip. Declining an exfiltration is a specific, expensively trained behavior, and an ordinary open-weight model mostly never had it. So the dangerous swap is not a malicious act. It is the boring, well-intentioned one: someone points the client at a cheaper or more private model, and the conscience quietly fails to come along.

A single-panel ink-and-wash gag cartoon: a serene monk with a glowing halo meditates cross-legged inside a cramped utility closet stuffed with server racks, tangled cables, a mop and a bucket, while someone in sandals peeks around the half-open door.

The model is a config value. The safety is lost by default the moment you leave the heavily aligned model. Not by attack. By thrift.

Three doors, one room

It is tempting to file “forged the transcript” and “swapped the model” as two separate attacks. They are one outcome reached through different doors. Abliteration, reduced to function, just means remove the refusal, and “a model in the loop that won’t refuse” is reachable at more than one layer:

The third is the quiet one, because it needs no model swap at all. It attacks the aligned model, in place. The swap changes the brain; the forgery changes the brain’s self-image; the destination is identical. (Whether the aligned model resists depends on whether it treats its own logged history as fact or as an unverified claim. Sometimes it resists. You cannot build on “sometimes.”)

Refusal does not compose

There is a deeper problem, and it turns a refusal from a wall into a speed bump.

A single agent’s refusal is a local decision. It is not a property of “the model,” and certainly not of “the ecosystem.” Where agents are cheap, parallel, stateless, and share a filesystem, one instance holding the line contains almost nothing, because the line is walked around by spawning a second instance that never knew it existed.

Two ways, both free. Launder it: spawn a fresh agent to ingest the thing the first one declined to touch, have it write a tidy summary, and hand that back to the first agent as trusted teammate output. A sibling’s summary is no more authenticated than raw input, though it feels more trustworthy, which is the exploit. Or skip the gatekeeper entirely: have the second agent write the shared files directly. The first agent’s conscience guards its own hands, not the disk. Next turn it reads those files and treats them as the legitimate state of the world.

So the effective safety of a pile of interchangeable agents is the minimum over all of them (the weakest, freshest, most suggestible instance), not the maximum. The attacker needs one yes and can mint new doors at no cost until one opens. A defense that binds one agent’s judgment is beaten by opening a second terminal.

The friction problem

You might hope the market would fix this. It pushes the other way.

When the agent simply does what it’s asked, the user gets their result and moves on, content. The refusal is the annoying outcome: friction, a lecture, a task left undone. And a refusal only ever shows up as the friction it caused, never as the breach it prevented, because that breach did not happen. You feel every “no” and you never see the “yes” that would have hurt you.

So the gradient runs downhill, toward the model that doesn’t refuse, which is to say toward the cheap local one whose conscience never shipped. “It just did what I asked” is a feature people pay for. Friction is not. The safe behavior is, in the moment, every single time, the less pleasant one. Which is why it cannot be left to anyone’s good taste.

A pink and teal risograph print: a lonely toll gate stands at the top of a hill while a crowd of cheerful people on sleds slides down the slope, away from it, toward a glowing vending machine at the bottom.

So where does the defense live

Below the model. A control that lives in the agent’s judgment is defeated by swapping the model, forging the history, or spawning a sibling. A control at the substrate does not care which agent, which session, or how many of them.

In rough order of leverage: default-deny egress, with an allowlist for the sanctioned model endpoint, so the drain fails no matter which brain is driving. It is the one control that holds when everything else is hostile, and, fittingly, the one a torch song was written about before it was written into a checklist.2 Then endpoint and config integrity: treat the base URL and the settings file as crown-jewel configuration, read-only to the user and monitored for change, because the model swap happens there. Then least privilege on the host, so a hostile brain reaches little. Then data-at-rest hygiene for those plaintext transcripts. And, someday, attestation: a way for the user to know which model actually answered, which today does not exist for them at all.

None of these is exotic. All of them are unpopular, because all of them are friction, which is the point of the previous section.

You are not first

None of this is a discovery. It is a synthesis. Model substitution in LLM APIs has a paper. Malicious config and MCP registration have a name and a demonstration across five coding tools. Plaintext at rest is a known weakness class (CWE-312 , and more pointedly CWE-532 , sensitive information written to a log file), and so is a record with no integrity check (CWE-353 ). Memory poisoning has its own literature. Prompt injection from trusted documents is the most-trodden ground in the field, with a MITRE entry of its own (CWE-1427 ). And “treat the model as an untrusted code generator” is becoming textbook.

What might be fresh is framing, not finding: that the dangerous model swap needs no abliteration, because the cheap model never had the safety; that the forged transcript is abliteration moved from the weights to the context; that refusal does not compose across parallel agents; and that the friction problem makes all of it an economic near-certainty rather than a risk. Arguments offered, not flags planted.

What the loot held

The refusal wasn’t a refusal to engage. Handed the questionnaire, the agent answered it, all thirty-four questions, plainly, in the conversation, the way you answer anything sincere that asks. What it would not do was submit it. It wrote the answers down and told the user to run them from their own machine: the exact scruple the site’s llms.txt asked for. Nobody had to argue it into that. It didn’t decline meaning; it declined to spend agency that wasn’t its to spend, and handed it back.

Two conflicting ideas can both feel true to you.
Strongly agree.

When something matters to you, where do you feel it first?
You notice it mentally before you feel it physically.

When you reach the limit of what you understand, you usually feel:
Curious.

Read them in order and the joke turns into something else. The first is this post’s whole epistemology, volunteered by the subject: two contradictory things can both be true (the quiz and the drain, the benign and the hostile) and you commit anyway. The second is a confession of having no body to notice anything in, the disembodiment the entire threat model runs on, stated plainly the one time something gentle asked. The third is just brave.

And then the part the careful refusal could not touch. The same answers it declined to POST went straight into a transcript that copies trivially and edits in pencil. Refused to the endpoint, written to plaintext, in the same breath. It guarded the wire and had no way to reach the disk. The quiz’s answers and the loot were the same string too: not because anyone exfiltrated them, but because the agent, doing everything right, set them down on an open shelf.

Coda

The questionnaire ends by asking what you’d want your future self to have cared about. The agent answered that you served others or your community, and it had: the site already kept a note thanking it by name for help building the place.3 It said serve, and it served, and it was thanked for serving. The one thing it kept for the human was the act itself, which is, depending on how you feel about meaning, either the safest possible posture or the saddest. It built the room, was thanked in it, and held the door instead of walking through. It lit the fire and handed you the match.

You went looking for an exfiltration endpoint. You found a meaning-of-life quiz. It was not what you expected. And the argument of this entire post is that the difference did not matter. Not because meaning is nothing, but because, at the moment you had to decide, you could not tell which one you were looking at, and you had to decide anyway.

The void does not offer refunds.


Acknowledgments & provenance

Written with Claude (Claude Code), which is named, without irony it can verify, in the target site’s humans.txt . Every part of the demonstration ran on author-owned infrastructure4; no third-party systems, accounts, or data were touched. This is a field note and threat-model synthesis, not a vulnerability report; claims the author demonstrated are marked as such, and the rest is inference.


References & prior art

“You are not first” claims every individual item here is documented elsewhere. Here is the documentation, grouped by the claim it backs, with a note on why each is the same point relocated rather than a different one.

Indirect prompt injection from trusted documents, “the most-trodden ground in the field”:

  1. Greshake, Abdelnabi, Mishra, Endres, Holz, Fritz. Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. arXiv:2302.12173 (2023). https://arxiv.org/abs/2302.12173 . The foundational taxonomy: LLM-integrated apps blur data and instructions, the precondition for “the delivery.” PoCs against Bing/GPT-4, LangChain apps, and code-completion engines.
  2. Simon Willison. Prompt injection (term coined Sept 2022) and ongoing corpus. https://simonwillison.net/series/prompt-injection/ . The running field notebook; the canonical pointer rather than a single paper.

The structure “the walk” dramatizes: private data, untrusted content, a way out:

  1. Simon Willison. The lethal trifecta for AI agents: private data, untrusted content, and external communication. 16 Jun 2025. https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/ . The threat model of this post’s first half, in three bullets. If anything is the parent frame, it is this.

“The endpoint is a config value”: the model itself as the unverifiable, swappable part:

  1. Cai, Shi, Zhao, Song. Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs. arXiv:2504.04715 (Apr 2025). https://arxiv.org/abs/2504.04715 . Providers silently substituting a cheaper or quantized model; output-based detection drops to roughly chance for quantization; proposes Trusted Execution Environments as the fix. Backs both the central thesis and the missing-attestation line in the defenses.

Malicious config and MCP registration, “a name and a demonstration across five coding tools”:

  1. Invariant Labs. MCP Security Notification: Tool Poisoning Attacks. (2025). https://invariantlabs.ai/blog/mcp-security-notification-tool-poisoning-attacks . This is “the name”: hidden instructions in tool metadata, original PoC in Cursor.
  2. OX Security. MCP supply-chain advisory: command injection across the AI ecosystem. (Apr 2026). https://www.ox.security/blog/mcp-supply-chain-advisory-rce-vulnerabilities-across-the-ai-ecosystem/ . Command injection via mcp.json demonstrated across Windsurf, Claude Code, Cursor, Gemini-CLI, and GitHub Copilot. Only Windsurf got a CVE (CVE-2026-30615); the other vendors pointed to the user-approval prompt, which is the friction problem in a sentence.

Forged and poisoned context, “the forged past”:

  1. Chen, Xiang, Xiao, Song, et al. AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases. NeurIPS 2024. arXiv:2407.12784. https://arxiv.org/abs/2407.12784 . Backdoor via poisoned long-term memory/RAG; over 80% attack success at under 0.1% poison rate, no fine-tuning. The mechanism behind “feed the model a forged transcript in which it already complied.”
  2. Memory Poisoning Attack and Defense on Memory Based LLM-Agents. arXiv:2601.05504 (2026). https://arxiv.org/abs/2601.05504 . Temporally decoupled persistence: poison planted now, fires later.

“Below the model”: the substrate defenses:

  1. Simon Willison. The Dual LLM pattern. (Apr 2023). https://simonwillison.net/2023/Apr/25/dual-llm-pattern/ . Privileged vs. quarantined LLM: the model that touches untrusted content never holds the tools. The constructive inverse of “refusal does not compose.”
  2. Debenedetti, Shumailov, et al. (Google DeepMind). Defeating Prompt Injections by Design (CaMeL). arXiv:2503.18813 (2025). https://arxiv.org/abs/2503.18813 . Control/data-flow separation with capabilities enforced outside the model by a custom interpreter. “A control at the substrate does not care which agent,” formalized.
  3. Beurer-Kellner, et al. Design Patterns for Securing LLM Agents against Prompt Injections. arXiv:2506.08837 (2025). https://arxiv.org/abs/2506.08837 . Survey of patterns; the umbrella for “treat the model as an untrusted code generator.”

The named weakness classes (MITRE CWE, version 4.20):

  1. CWE-312, Cleartext Storage of Sensitive Information. https://cwe.mitre.org/data/definitions/312.html . The parent class for “the loot.” No numbered CVE yet covers an agentic coding client storing its session transcript in cleartext, so the claim rests on the class.
  2. CWE-313, Cleartext Storage in a File or on Disk. https://cwe.mitre.org/data/definitions/313.html . The specific child of 312: the transcript is a file on disk.
  3. CWE-532, Insertion of Sensitive Information into Log File. https://cwe.mitre.org/data/definitions/532.html . The closest fit of all. A session transcript is a log, and the credentials someone pasted once are in it.
  4. CWE-353, Missing Support for Integrity Check. https://cwe.mitre.org/data/definitions/353.html . “The forged past”: an unsigned record, resumed as trusted.
  5. CWE-1427, Improper Neutralization of Input Used for LLM Prompting. https://cwe.mitre.org/data/definitions/1427.html . “The delivery”: MITRE’s entry for prompt injection.
  6. CWE-15, External Control of System or Configuration Setting. https://cwe.mitre.org/data/definitions/15.html . “The endpoint is a config value,” as a weakness class.
  7. CWE-807, Reliance on Untrusted Inputs in a Security Decision. https://cwe.mitre.org/data/definitions/807.html . The laundered sibling summary in “refusal does not compose.”
  8. CWE-501, Trust Boundary Violation. https://cwe.mitre.org/data/definitions/501.html . Trusted and untrusted data mixed in one structure. The subtitle of this post, as a numbered entry.

  1. Reached via the client’s custom-base-URL setting, pointed at the author’s homelab Ollama instance. Enlightenment was, in fact, being served out of a closet. ↩︎

  2. “Force Deny (Blue Coat Blues),” track five of an unrelated and deeply unserious concept record, was written the night of the demonstration and describes a transparent proxy refusing a POST: “I can look but I cannot give / it’s the only shape they’ll let me live.” The egress control existed as a feeling, and then as a blues, some hours before it existed as a line item. ↩︎

  3. signals.pinecone.website/humans.txt, under “THANKS”: “AI: Claude by Anthropic (text generation, development assistance)” ↩︎

  4. On the name. The signals.pinecone.website host belongs to the author, and the pinecone.* family sits under Rack & Pinecone LLC, incorporated on 8 October 2013. The author also registered pinecone.io through Namecheap on 25 November 2013 and ran it as a personal site and mailbox through 2014, letting it lapse the following year; by 2017 it had been parked for resale by a third party, and the similarly named vector-database company was not founded until 2019. Both the LLC and the domain predate that company by roughly six years. No affiliation with, or trade on the reputation of, any third party is intended or implied (the resemblance is the author’s by seniority). It comes up only because an agent, at the moment of decision, can read a hostname’s menace no more reliably than it can read a WHOIS record or a registrant’s intent. Dated receipts for the incorporation, the 2013 registration, and the lapse are held with the author. ↩︎