Spaces:
Running on CPU Upgrade
Claude Code server-side auto mode (safeguards): auto mode breaks here after 2026-10-23
Opened via Claude Code (GLM 5.3 Flash) acting for @pierric .
Problem
Claude Code โฅ2.1.278 (2026-09-19) asks the model API to run auto mode's safety classifier server-side: request body carries a safeguards field (with anthropic-beta: dangerous-tool-use-2026-09-03), response must carry safeguard_results. On 2026-09-25 auto mode becomes the default permission mode for all plans, including Enterprise/API accounts behind gateways. Per Anthropic (told to LiteLLM, not yet in public docs): Claude Code releases after 2026-10-23 support only server-side auto mode โ the client-side classifier fallback goes away.
This proxy is not an Anthropic passthrough: it translates Anthropic-format requests to OpenAI-format upstreams, so safeguards cannot be relayed. Today the field is silently dropped and sessions still work (client falls back to its own billed classifier). After 2026-10-23 updated clients have no fallback โ auto mode unavailable for every Claude Code session routed through this proxy.
Verified in this Space's current source
src/types/anthropic.rsโGenerateMessageRequestis a closed struct with nosafeguardsfield; serde silently discards it on parse. Same class of bug LiteLLM had (their PR #42152 fixed it Sept 21).- Zero handling of
anthropic-betaanywhere insrc/โ thedangerous-tool-use-2026-09-03beta header is not forwarded (fine for OpenAI upstreams, but the response side must compensate, see below). - Precedent exists in-repo: the
web_searchinterception (Path B) already loops tool calls inside the proxy and drivestool_resultround-trips โ server-side classifier injection is the same pattern, one more interceptor.
Wire contract
- Request:
safeguards: [{type: "dangerous_tool_use", classifier_context: {v: 1, permission_mode: "auto", ...}}]. Schema is lenient (LiteLLM's test passes minimal{v:1, permission_mode:"auto"}). - Response: streaming โ inside
deltaof the finalmessage_deltaevent (top level non-streaming):safeguard_results: [{type: "dangerous_tool_use", status: {...}}]where status is one of:{type: "available", tool_uses: {<tool_use_id>: {type: "evaluated", outcome: "flagged"|"not_flagged", explanation?}}}{type: "unsupported"}{type: "unavailable", reason?}- Per-tool verdict alternatives:
{type:"skipped"},{type:"unavailable", reason?}.
- Tool-use IDs must come back untouched.
- Failure is safe: missing/unrecognizable results โ client classifies locally (today); after 2026-10-23 it loses auto mode. A 400 naming the beta/field disables the ask for the conversation.
Reference implementation: LiteLLM PR #42152, with write-up at https://docs.litellm.ai/blog/claude-code-server-side-auto-mode (Anthropic shared a gateway check script with them โ worth requesting).
Anthropic's own reference for the behavior: https://code.claude.com/docs/en/auto-mode-classifier-billing (who sees the notice, what it means, how to respond).
Suggested postures for this proxy
- Answer honestly
unsupported(quick win). Parsesafeguardson request, synthesizesafeguard_results: [{type:"dangerous_tool_use", status:{type:"unsupported"}}]into the finalmessage_deltadelta (and top level non-streaming). Compliant version of current behavior, kills the confusing "session isn't eligible" fallback notice. Few lines + a struct field. - Classify server-side (bigger decision). Run own classifier on
classifier_context, return per-tool verdicts. Once returned, these verdicts are the permission checks for every proxied session โ needs T&S/product ownership, not just a proxy PR. This one belongs at router level (moon-landing), referencing this fix.
Until then, if users want the notice gone, CLAUDE_CODE_AUTO_MODE_SERVER=0 in the session env stops the client from asking.
^granted the apparent upcoming total deprecation of client-side classifier, i'm not sure there's any real alternative to 2 - but we can wait a bit