Claude Code server-side auto mode (safeguards): auto mode breaks here after 2026-10-23

#7
by pierric - opened

Opened via Claude Code (GLM 5.3 Flash) acting for @pierric .

Problem

Claude Code โ‰ฅ2.1.278 (2026-09-19) asks the model API to run auto mode's safety classifier server-side: request body carries a safeguards field (with anthropic-beta: dangerous-tool-use-2026-09-03), response must carry safeguard_results. On 2026-09-25 auto mode becomes the default permission mode for all plans, including Enterprise/API accounts behind gateways. Per Anthropic (told to LiteLLM, not yet in public docs): Claude Code releases after 2026-10-23 support only server-side auto mode โ€” the client-side classifier fallback goes away.

This proxy is not an Anthropic passthrough: it translates Anthropic-format requests to OpenAI-format upstreams, so safeguards cannot be relayed. Today the field is silently dropped and sessions still work (client falls back to its own billed classifier). After 2026-10-23 updated clients have no fallback โ†’ auto mode unavailable for every Claude Code session routed through this proxy.

Verified in this Space's current source

  1. src/types/anthropic.rs โ€” GenerateMessageRequest is a closed struct with no safeguards field; serde silently discards it on parse. Same class of bug LiteLLM had (their PR #42152 fixed it Sept 21).
  2. Zero handling of anthropic-beta anywhere in src/ โ€” the dangerous-tool-use-2026-09-03 beta header is not forwarded (fine for OpenAI upstreams, but the response side must compensate, see below).
  3. Precedent exists in-repo: the web_search interception (Path B) already loops tool calls inside the proxy and drives tool_result round-trips โ€” server-side classifier injection is the same pattern, one more interceptor.

Wire contract

  • Request: safeguards: [{type: "dangerous_tool_use", classifier_context: {v: 1, permission_mode: "auto", ...}}]. Schema is lenient (LiteLLM's test passes minimal {v:1, permission_mode:"auto"}).
  • Response: streaming โ†’ inside delta of the final message_delta event (top level non-streaming): safeguard_results: [{type: "dangerous_tool_use", status: {...}}] where status is one of:
    • {type: "available", tool_uses: {<tool_use_id>: {type: "evaluated", outcome: "flagged"|"not_flagged", explanation?}}}
    • {type: "unsupported"}
    • {type: "unavailable", reason?}
    • Per-tool verdict alternatives: {type:"skipped"}, {type:"unavailable", reason?}.
  • Tool-use IDs must come back untouched.
  • Failure is safe: missing/unrecognizable results โ†’ client classifies locally (today); after 2026-10-23 it loses auto mode. A 400 naming the beta/field disables the ask for the conversation.

Reference implementation: LiteLLM PR #42152, with write-up at https://docs.litellm.ai/blog/claude-code-server-side-auto-mode (Anthropic shared a gateway check script with them โ€” worth requesting).

Anthropic's own reference for the behavior: https://code.claude.com/docs/en/auto-mode-classifier-billing (who sees the notice, what it means, how to respond).

Suggested postures for this proxy

  1. Answer honestly unsupported (quick win). Parse safeguards on request, synthesize safeguard_results: [{type:"dangerous_tool_use", status:{type:"unsupported"}}] into the final message_delta delta (and top level non-streaming). Compliant version of current behavior, kills the confusing "session isn't eligible" fallback notice. Few lines + a struct field.
  2. Classify server-side (bigger decision). Run own classifier on classifier_context, return per-tool verdicts. Once returned, these verdicts are the permission checks for every proxied session โ€” needs T&S/product ownership, not just a proxy PR. This one belongs at router level (moon-landing), referencing this fix.

Until then, if users want the notice gone, CLAUDE_CODE_AUTO_MODE_SERVER=0 in the session env stops the client from asking.

^granted the apparent upcoming total deprecation of client-side classifier, i'm not sure there's any real alternative to 2 - but we can wait a bit

Sign up or log in to comment