Skip to main content Announcing Tool Gateway MCP: the universal MCPRead the announcement
Guillaume Lebedel Guillaume Lebedel · · 7 min
An orchestrator model handing work to three Haiku 5.5 subagents, each with its own scoped credentials

How to scope permissions for Claude Haiku 5.5 subagents

Table of Contents

Claude Haiku 5.5, released by Anthropic on October 7, 2026, costs around 75% less to run than Haiku 4.5 and is pitched as a subagent for summaries, database queries and browser use. Scope each subagent on its own: separate credentials tied to the person it acts for, read-only tools by default, approval for writes, and a log of every call.

Facts in this post were checked on October 8, 2026.

What is Claude Haiku 5.5 built for, and how much cheaper is it?

Anthropic built Haiku 5.5 for high-volume, narrowly scoped work: summaries, compactions, database queries, classification and browser use, often as a subagent under Opus 5.5 or Sonnet 5.5. For prompts up to 100,000 tokens it is priced 90% below Haiku 4.5, and Anthropic puts the average saving at around 75% once its new tokenizer is counted.

Price per million tokensHaiku 4.5Haiku 5.5 (prompts up to 100k)Haiku 5.5 (prompts over 100k)
Input$1.00$0.10$0.50
Output$5.00$0.50$2.50
Cache reads$0.10$0.01$0.05

Bar chart comparing Haiku 4.5 and Haiku 5.5 input and output prices per million tokens

Prices are from Anthropic’s Haiku 5.5 announcement, which notes that prompts under 100,000 tokens made up around 90% of Haiku 4.5 requests. Sonnet 5.5 and Opus 5.5 stay the better choice for complex agentic coding; Haiku 5.5 targets narrow tasks that “might otherwise have been cost-prohibitive”, subagent work included.

Those narrow tasks run inside business systems. HubSpot told Anthropic, in a customer quote on the Haiku 5.5 launch page, that Haiku 5.5 got the best score it had seen on its CRM eval suite, 92.8% averaged over three runs, on tasks such as identifying stale but ambiguous records.

Why do cheaper subagents mean more agents running at once?

Cost was the main thing holding back fan-out. When Anthropic described its multi-agent research system in June 2025, it reported that multi-agent setups used about 15 times more tokens than a chat, and only paid off on high-value tasks. A 75% cheaper subagent moves a lot of work over that line.

That 15x figure is more than a year old and was measured with Opus 4 and Sonnet 4 subagents. The direction still holds: the Haiku 5.5 announcement makes the same argument about subagent work that used to cost too much.

The same post describes a lead agent starting 3 to 5 subagents in parallel, and early versions starting 50 for simple queries. For teams building internal agents for support, onboarding or finance ops, that means more processes reading CRM records and querying databases for a single request. We have written about designing async subagents that work like contractors rather than apprentices; the permission model has to keep up with that design.

What credentials does a subagent inherit from its orchestrator?

Usually whatever the orchestrator holds. In many setups the lead agent passes its own tools and token down, so a subagent doing a read-only lookup can call anything the orchestrator could, with the same rights. OWASP lists this pattern under ASI03, Identity and Privilege Abuse, in its Top 10 for Agentic Applications.

Take an orchestrator holding an admin key for the whole Salesforce org. A Haiku subagent asked to flag stale contacts now runs every lookup with admin rights, and an instruction planted in a contact note can steer it toward writes or exports the task never needed. Agents doing more than they were asked is common enough that we treat AI agent failures as a permissions problem.

Diagram comparing three subagents sharing one admin token with three subagents each holding scoped, read-only credentials

Shared orchestrator tokenPer-subagent scope
Identity on each callOne service accountThe user the run is for, plus the subagent name
Default accessEvery tool the orchestrator hasOnly the tools the task lists
WritesAllowed wherever the token allowsOff unless the task is a write
Sensitive fieldsReturned in fullMasked unless the task needs them
What one poisoned record can reachEverything the admin key reachesThe few read tools in that subagent’s scope

How do you scope permissions for each subagent?

Give each subagent its own identity and the smallest tool set its task needs, default it to read-only, and check every write before the business system is called. Then run new restrictions in monitor mode first, so you can see what they would block before they block real work.

  1. Write the task down before choosing tools. A stale-record check needs to list and read contacts and deals. It does not need update, delete or export.
  2. Issue credentials per subagent, tied to the user the run is for. Do not pass down the orchestrator’s service account. If the user cannot see a field, the subagent working for them should not see it either.
  3. Default to read-only. Give write tools only to the subagent whose job is the write, and require approval for deletes, bulk updates and payments.
  4. Mask fields the task does not use. A stale-record check needs the last activity date, not personal email or bank details.
  5. Run new rules in monitor mode first. Read what they would have denied, then enforce. We explain why agent permission policies should roll out in monitor mode before enforcement.
  6. Log every call with the subagent’s name, the user it acted for and the policy decision.

That’s why we built Permission Policies at StackOne. Policies are evaluated for the member the agent acts for on every request: a denied tool or input value stops before the business system is called, and denied output fields such as salary or bank details are masked on the way back, across 540+ connectors. They apply to calls made for a member of your StackOne organization, so a subagent calling with a member’s identity gets that member’s narrowed access, while an API-key call with no member identity does not receive member policies. Tool search helps the cheap model too: the agent sees two tools, tool_search and tool_execute, instead of hundreds of definitions.

What changes when a subagent can use a browser?

A browser subagent acts with whatever session its browser holds, so its permissions are those of the logged-in user on every site that session can reach. Anthropic added browser use and computer use to its Python and TypeScript SDKs in beta alongside this launch, and says Haiku 5.5 is especially well suited to both.

A tool call can be checked field by field; a click in a logged-in admin console cannot. Three habits help:

  • Give browser subagents a dedicated profile signed in only to the accounts their task needs, never your own admin session.
  • Restrict the domains the browser may open.
  • Use an API tool call instead of the browser whenever one exists, because it can be scoped, masked and logged per call.

How do you log which subagent made which tool call?

Record each call with the subagent that made it, the user it acted for, the tool and inputs, and the policy decision. Without the subagent name, the log shows one orchestrator doing everything, and you cannot tell which delegated task read a record or tried a write.

Pass a run ID and a subagent ID down with every tool call so they reach the log. For the full list, see what an AI agent audit record should contain, from who delegated the run to the policy text that applied.

If your agents reach business systems through StackOne, see how Permission Policies scope each call to the person the agent acts for.

Frequently Asked Questions

Is Claude Haiku 5.5 good enough to run as a subagent?
Claude Haiku 5.5 is good enough to run as a subagent for narrow, well-defined tasks, according to Anthropic. It is positioned for summaries, database queries, classification and browser use under a larger planning model. For complex agentic coding, Anthropic still points to Sonnet 5.5 or Opus 5.5 as the lead model.
Should every subagent get its own API key?
Giving every subagent its own API key beats one shared key, though a key on its own says nothing about the user the run is for. The stronger setup gives each subagent the requesting user's identity plus a subagent name, so access follows that user's permissions and every call is attributable in the log.
Does read-only access stop prompt injection?
Read-only access does not stop prompt injection, but it limits what an injected instruction can do. A read-only subagent cannot write, delete or send anything, yet it can still read records and return them to the orchestrator. Pair read-only defaults with masked sensitive fields and screening of tool output for planted instructions before the model acts on it.
How many subagents should an orchestrator start?
How many subagents an orchestrator should start depends on the task. Besides its usual 3 to 5 parallel subagents, Anthropic's June 2025 research report gave a scaling rule: one agent with 3 to 10 tool calls for simple fact-finding, 2 to 4 subagents for comparisons and more than 10 for complex research. The direction still applies.

Put your AI agents to work

All the tools you need to build and scale AI agent integrations, with best-in-class connectivity, execution, and security.