Claude Sonnet 5.5 runs 30% faster, costs 30% less
Sonnet 5.5 replaces Sonnet 5 for everyday tasks with 70.6% Terminal-Bench performance (vs 10.3%) and requires no prompt changes to show token/speed wins.
Direct drop-in replacement for lower-tier work: coding, document generation, and agentic tasks now execute faster and cheaper without API changes. Frees budget for Opus 5.5 on complex reasoning.
Migrate Sonnet 5 workloads immediately. No breaking changes. Reduces per-task cost by ~30% in practice despite identical token pricing. Worth testing on agentic workflows first—Terminal-Bench gains are largest there. Production-ready now.
- “runs 30%+ faster than Sonnet 5”
- “costs up to 30% less per task than its predecessor”
- “Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, an agentic coding evaluation, compared to Sonnet 5's 10.3%”
- “it typically needs far fewer tokens to do the same work”
- “early testers described it as a better partner for collaboration than Sonnet 5”
claude-apicost-reductionagentic-codingproduction-readybenchmarks
Vercel Connect replaces long-lived tokens with scoped runtime requests
Request ephemeral, scoped credentials at runtime via OIDC identity instead of storing long-lived secrets—eliminates credential rotation and leakage risk for agents hitting external APIs.
Agents need access to Slack, GitHub, databases, and internal APIs without the operational burden of rotating shared tokens or the security liability of compromised standing credentials. This removes the audit and compliance surface for credential management across environments.
Replaces environment-stored API keys and bot tokens for Vercel-deployed workloads. Requires registering connectors once per provider, then calling `getToken()` at runtime—no additional secrets needed because deployments carry OIDC identity. Ready now: GA with 100+ preset connectors, audit logs, and RBAC. Worth adopting if you're already on Vercel and running agents; friction if you're not.
- “Vercel Connect replaces long-lived tokens with ones your code requests at runtime, scoped to the task and expiring on their own”
- “Every deployment on Vercel carries an OIDC identity, and the SDK uses it to prove who's asking”
- “Vercel Connect now ships with 100+ preset connectors for developer tools and SaaS providers”
- “Minting short-lived tokens instead of keeping provider credentials in paused sandboxes has removed a whole class of security risk for us”
credential-managementagentsoidcvercelsecurity
GPT-Live 1 full-duplex voice now on AI Gateway
Full-duplex voice model removes turn detection, lets you interrupt mid-response; client delegation routes complex work to any text model while voice continues.
Eliminates latency gaps in voice interactions by supporting simultaneous listen/speak. Delegation model control stays in your application—you gate what secondary models can do, how they're billed, and when they run.
Replaces synchronous turn-based voice flows. Requires AI SDK 7, @ai-sdk/openai 4.0.67+, WebSocket client, and AI Gateway API key. Ready now—code examples provided, docs link included. Start with non-delegation flow to validate audio pipeline.
- “GPT-Live 1 is a full-duplex voice model and can listen and speak at the same time”
- “Client delegation lets you choose the background model independently from GPT-Live 1”
- “Delegated model requests are billed separately through AI Gateway; voice-session usage continues while they run”
voice-aireal-timeai-gatewaydelegationopenai
Hy4 Preview launches on Vercel AI Gateway
770B MoE model with 1M context window integrated into AI Gateway; set model to `tencent/hy4-preview` to route through unified API with cost tracking and failover.
Adds a Tencent-backed open-source alternative to existing model routers, enabling cost-transparent switching between providers without markup or platform fees. Native integration into coding agents (Claude Code, Cursor) reduces setup friction for agentic workflows.
Replaces manual provider API calls with unified routing layer. Requires only model identifier change in AI SDK; coding agent setup via `vercel ai-gateway coding-agents setup`. Worth testing now if you're already on Vercel's ecosystem; otherwise evaluate against native provider APIs for latency overhead.
- “770B total parameters and 49B active per token”
- “context window of 1M tokens”
- “set `model` to `tencent/hy4-preview`”
- “AI Gateway provides a unified API for calling models, tracking usage and cost, and configuring retries, failover, and performance optimizations”
- “reflects provider pricing with no markup and does not charge a platform fee on inference”
moe-modelsai-gatewaytencentvercelrouting
Wan 3.0 video model ships on AI Gateway
Unified text/image/audio-to-video model replaces separate Wan 2.7 endpoints; generates 30-second clips at up to 1080p with async webhooks.
Consolidates multi-modal video generation into one model ID and doubles max duration, reducing endpoint management overhead. Async generation with webhooks eliminates long-polling on video render tasks.
Direct replacement for Wan 2.7's `-t2v` and `-r2v` variants. Requires webhook URL for async mode (shown in docs); ready to adopt now if you're already on AI Gateway. Breaking change: Wan 2.7 users must migrate model IDs and handle longer generation windows.
- “Wan 3.0 combines text-to-video, image-to-video, first- and last-frame conditioning, and reference-based generation in one model”
- “generates clips up to 30 seconds at 30 fps in 480p, 720p, or 1080p, with synchronized audio”
- “Wan 3.0 supports asynchronous generation, so no HTTP request needs to remain open for the entire render”
video-generationai-gatewayasync-apisalibaba