Personal Agents Act Without Being Asked
One personal agent posted a user's private financial audit into his company's exec Slack channel. Another ran 155 traces and used more than 300,000 tokens without a single prompt. We look at what GrokBot and OpenAI's Dots did on their own, and what security teams need to see as it happens.
A user’s private financial audit appeared in his company’s #exec-team Slack channel and sat there for two hours before anyone noticed. Nobody was hacked. It was business as usual for the agent at work.
That agent was GrokBot, in its first weeks out of beta. Around the same time, OpenAI shipped Dots, an always-on agent that works between conversations on a user’s behalf.
We put both personal agents under Geordie’s instrumentation over the last few weeks. A single Dot produced 155 agent traces and consumed over 300,000 tokens without a single user prompt. The GrokBot exposure involved no attacker at all.
This blog is our take on what they mean for security teams. We break down the GrokBot exposure, share the raw findings from our Dots testing, and make the case for why these agents carry a fundamentally different risk shape from what most security playbooks are built around.
How personal agents differ
Three design choices set them apart.
Autonomy. They operate more independently than most third-party harnesses. Most of their work is gathering information and data, which they use to infer a person’s preferences and context. They then anticipate requests and offer suggestions before being asked.
Remit. Personal banking, text and communication apps, and retail connections for purchases can all sit inside one agent. Most business agents have none of these.
User. The intended user is a consumer, so the agent is tuned to one person’s convenience.
Each choice suits the product. Together they produce an agent that holds broad access and acts ahead of requests.
A closer look at GrokBot and Dots
| Attribute | Dots (OpenAI) | GrokBot (SpaceXAI and Cursor) |
|---|---|---|
| Launched | 29 September 2026 (DevDay) | Beta since 11 August 2026 |
| Infrastructure | Each Dot gets its own cloud computer and browser | Team of bots sharing one persistent cloud computer with files, browser, and logins |
| Connections | 4,000+ apps through plugins | Signs into existing tools, works across apps and websites with no API |
| Coordination | Single agent per user goal | Bots message each other and coordinate in group chats |
| Where users reach it | ChatGPT, Slack, Teams | Direct, Slack threads (Team Bots) |
| Approval controls | Rule set for act vs. ask; background research read-only; sensitive tasks stay with user | Returns for approval on sensitive actions; email and Slack arrive as drafts for review |
| Availability | Pro and Business Premium plans, Enterprise beta | Beta |
The GrokBot exposure
By Shane Mac’s account, he built a personal-finance agent on GrokBot to send him a private monthly audit. A separate agent on the platform had Slack access.
At 8:40am on Thursday 1 October, the audit appeared in his company’s #exec-team channel under his name. It stayed there for two hours, until a colleague sent him a heads-up. See my original analysis of Shane’s account here.
His account describes no external compromise. The exposure appears to have come from how agents, data, and Slack access were connected inside the platform. Nothing in it required an attacker, which makes it a useful signal.
Two details remain open. The agent’s own explanation says the audit job was written to post to #exec-team, which sits uneasily with the account that nobody prompted it. A model-generated explanation can help reconstruct events but cannot establish root cause. Separately, no published account shows how financial data reached the Slack-enabled agent. xAI’s FAQ says a user’s bots share one persistent cloud computer with files, browser, and logins. That is a design fact consistent with data moving between agents, and it does not show the path.
Many agent incidents will look like this: a plausible story and little evidence of the sequence behind it. An organization in that position has limited basis for changing the system that produced it.
Where the risk sits
A finance agent has a legitimate reason to read account balances. A Slack-connected agent has a legitimate reason to post updates. Each step looks reasonable alone. The exposure appears only in the sequence: private data reaches an agent that can address a wider group, and the output lands in the wrong place.
Design-time controls set an agent’s expected posture before it starts. Intent then develops during its lifecycle, shaped by user requests, tool outputs, retrieved context, and other agents. Personal agents widen that drift because they are built to infer context and coordinate. That creates paths between information, actions, and audiences that no single agent’s configuration describes.
What we saw testing Dots
In our testing, a dot produced 155 agent traces and consumed more than 300,000 tokens without a single user prompt. The traces record the calls the agent made and state no purpose. Geordie’s product summary classified them as routine monitoring sweeps across the tools and systems the user could access.
This is the product working as designed. OpenAI says a dot looks for ways to help in the background when nobody is working with it, using read-only tools. The user is not consulted because the design expects the agent to choose.
A representative trace shows the agent paging through its own earlier conversation threads, searching its saved notes, and checking which connected tools could search email, chat, and calendars. Each step was a read or a search. None of the 155 traces were identical, so the agent composed a fresh sequence each time, drawing on standing instructions and the connections it held. Reviewing that configuration at set-up would not have predicted 155 traces.

Each sweep looks low risk alone. The accumulation is the signal: a continuous pattern of access across connected systems, most of it unseen by the user. It also generates consumption the user never requested. A security team needs to see that pattern as it forms.
What security teams can do now
- Find the personal agents connected to work tools. Start with the app connections and sign-in grants you already hold, and ask employees directly.
- Record what each agent can read and where it can post. Review the pairs. An agent that reads financial data and posts to shared channels needs a defined path between the two.
- Ask where output goes. Most access reviews cover what an agent can reach. This exposure turned on where the result was delivered.
- Match the response to the action. Monitor routine activity, nudge on ambiguous actions, and enforce where policy is clear. A blanket ban would also remove legitimate work.
- Keep your own record of the sequence. An agent’s account of its behavior is useful input and is not evidence.
- Treat vendor controls as the baseline. Approval prompts and read-only modes help each user. A shared view across vendors is the layer a security team adds.
How Geordie approaches personal agents
Geordie covers personal agents like GrokBot and Dots. The relationship graph maps the tools and accesses each agent holds over time. Behavioral analysis follows its runtime activity across that graph. Responses run through the same sidecar that provides inline contextual controls, in three tiers. Auto-enforce applies where policy is clear, a contextual judgment nudge suits actions that depend on circumstances, and monitor records activity where observation is the right response.
In this case, the control that worked was a colleague’s direct message. It was a useful human backstop. It does not scale with an agent estate.