Haiku 5.5 adds tunable effort levels (`low` to `max`) that control token spend and thinking depth—replace 4.5 in cost-sensitive tasks like summarization and classification without changing your API calls.
Summary
Token-per-task costs drop when you dial reasoning down for simple work, while `xhigh` and `max` effort modes handle complex reasoning without model swaps. Faster latency at standard speed suits live support and browser use.
Why it matters
Token-per-task costs drop when you dial reasoning down for simple work, while `xhigh` and `max` effort modes handle complex reasoning without model swaps. Faster latency at standard speed suits live support and browser use.
Implementation verdict
Direct drop-in replacement for Claude Haiku 4.5 via `anthropic/claude-haiku-5.5` model string on AI Gateway. Set effort with `reasoning` param in AI SDK or `reasoning_effort` in Chat Completions API. Worth testing now in summarization and classification pipelines to measure token savings; requires zero code refactor.
Sources
Dev Signal
Get briefs like this in your inbox — free, every weekday.
100+ sources compressed into one 4-minute read. Ranked, cited, implementation-ready.