Claude Haiku 5.5 launches on Amazon Bedrock today
Haiku 5.5 costs 75% less than Haiku 4.5, supports effort controls for cost-vs-intelligence tuning, and runs as subagents paired with Opus 5.5 for hierarchical workflows.
Enables cost-effective scaling of repetitive tasks (classification, routing, code review, document summarization) without sacrificing multi-step tool use or computer-use capabilities. Effort controls let you tune inference budgets per task instead of picking one setting for entire workloads.
Replaces Haiku 4.5 for volume work; requires AWS account with Bedrock access, Boto3/Anthropic SDK, and IAM permissions (bedrock:InvokeModel). Ready now—three invocation paths available (Boto3 InvokeModel, Converse API, Anthropic SDK via bedrock-runtime). Start with Bedrock Playground, then move to code.
- “costs around 75 percent less than Claude Haiku 4.5 for most tasks”
- “the fastest and most efficient model in the Claude 5.5 family, built for subagents and high-volume, cost-sensitive work”
- “the first Haiku model with effort controls, so you can tune cost against intelligence for each task”
- “Claude Haiku 5.5 takes on the fast layer of subagents”
- “available today on Amazon Bedrock”
claude-haiku-5.5bedrockcost-optimizationmulti-agentaws
Claude Haiku 5.5 ships with adaptive reasoning levels
Haiku 5.5 adds tunable effort levels (`low` to `max`) that control token spend and thinking depth—replace 4.5 in cost-sensitive tasks like summarization and classification without changing your API calls.
Token-per-task costs drop when you dial reasoning down for simple work, while `xhigh` and `max` effort modes handle complex reasoning without model swaps. Faster latency at standard speed suits live support and browser use.
Direct drop-in replacement for Claude Haiku 4.5 via `anthropic/claude-haiku-5.5` model string on AI Gateway. Set effort with `reasoning` param in AI SDK or `reasoning_effort` in Chat Completions API. Worth testing now in summarization and classification pipelines to measure token savings; requires zero code refactor.
- “built for high-volume, cost-sensitive tasks such as summaries, context compaction, database queries, and classification”
- “the fastest Claude model at standard speed, which suits live customer support and browser use”
- “supports `low`, `medium`, `high`, `xhigh`, and `max` effort, which set how much the model thinks and how many tokens it uses”
- “At `xhigh` and `max`, thinking must stay on”
- “Set effort with `reasoning` in the AI SDK or `reasoning_effort` in Chat Completions”
claudeinference-optimizationtoken-efficiencyreasoning-levelsanthropic
Decisions API routes structured inference through AI Gateway
OpenAI's Decisions API now accessible via AI Gateway's `/v1/decisions` endpoint—returns probabilities and choices instead of text, replacing separate classification pipelines with typed question-answering against shared input.
Consolidates routing, triage, and scoring into a single request with predicate, choice, and score question types. Eliminates prompt engineering overhead for classification tasks by treating decisions as a first-class API primitive.
Ready to use now. Requires OpenAI SDK 7.30.0+ (JS) or 3.26.0+ (Python); point baseURL to AI Gateway and call `decisions.create()`. Replaces chained prompt-based classification patterns. Trial immediately if you're currently parsing LLM output for routing decisions.
- “Decision models answer typed questions about a shared input and return probabilities, choices, and scores instead of generated text, which suits routing, triage, guardrails, and rubric scoring.”
- “You can ask predicate (yes/no), choice, and score questions against the same input in one request.”
- “decisions.create requires the OpenAI SDK 7.30.0 or later for JavaScript, or 3.26.0 or later for Python.”
decisions-apistructured-outputrouting-inferenceai-gatewayclassification
OpenAI Ultrafast tier launches on Vercel AI Gateway
GPT-6 Astra requests now route through Ultrafast on US/global infrastructure via `serviceTier: 'ultrafast'` parameter, trading 6× token cost for reduced latency on interactive workloads.
Developers building latency-sensitive coding assistants or chat interfaces can now optimize for speed over cost at the parameter level. Regional fallback to standard tier removes hard geographic constraints, but EU requests automatically downgrade.
Replaces manual regional routing logic. Requires specifying `serviceTier` in OpenAI SDK calls or Chat Completions API. Worth testing now for real-time use cases, but verify 6× pricing impact against your latency SLA—worthless if you're already meeting targets at standard tier.
- “Ultrafast supports US and global processing”
- “Requests pinned to unsupported regions, such as the EU, run at the standard (`default`) tier”
- “Requests served at Ultrafast are billed at 6× the standard per-token rate”
- “For workflows with frequent tool calls, OpenAI recommends the Responses API over a persistent WebSocket connection to reduce overhead between turns”
openailatency-optimizationgpt-6-astraservice-tierscost-tradeoff
Claude Haiku 5.5 matches Luna pricing under 100k tokens
Haiku 5.5 hits price parity with GPT-6 Luna ($0.10/$0.50) for short contexts, but tokenizer inflation and steep 5x pricing cliff above 100k tokens create a narrow competitive window.
Cost-sensitive inference workloads under 100k tokens now have true parity benchmarking; beyond that threshold, token efficiency and pricing trade-offs force explicit model selection per use case rather than blanket adoption.
Replaces Haiku 4.5 for cost-optimized inference. Requires re-baselining token counts (1.25x inflation vs prior version) and profiling context lengths before commit. Test now if you're already on Anthropic; switch only if workloads consistently exceed 100k tokens where Luna becomes 5x cheaper.
- “Claude Haiku 5.5”
- “$0.10/$0.50—up to 100,000 tokens. Beyond 100,000 tokens the price increases 5x to $0.50/$2.50”
- “same long prompt uses around 1.25x as many tokens with Haiku 5.5 compared to Haiku 4.5”
- “Above 100,000 tokens, Luna looks like a much better deal”
- “Max 5x users will get $100 in credits per month, Max 20x users will get $200, and Team subscribers will receive up to $500”
anthropicpricing-analysisinference-optimizationtoken-efficiencymodel-selection