Two new GPT-6 variants ship on Vercel's AI Gateway: Sol for sustained coding work, Luna for high-volume agentic tasks at lower cost than Astra.
Developers routing production inference through AI Gateway can now choose models optimized for either quality (Sol) or throughput (Luna) without leaving the platform. Both variants claim improved factual reliability and less jargon, reducing hallucinations in coding contexts where accuracy compounds.
Drop-in model string replacement for existing OpenAI API calls (`openai/gpt-6-sol` or `openai/gpt-6-luna`) across AI SDK, Chat Completions, and Codex agents. Requires Vercel CLI setup and API Gateway credential provisioning. Ready to try now if already on Vercel infrastructure; pricing comparison against Astra needed before production routing.
“GPT-6 Sol and GPT-6 Luna from OpenAI are now available on AI Gateway”
“GPT-6 Sol is suited to complex professional workflows and sustained coding tasks where quality and room to iterate both matter”
“GPT-6 Luna is the lower-cost option for high-volume agentic workflows, coding, and everyday tasks”
“Both Sol and Luna communicate more directly than their GPT-5.6 counterparts, with less jargon and fewer low-value details”
“less likely to make misleading claims about work completed during coding tasks”
Get issues like this in your inbox — free, every weekday.
Quick Signals
Gemini 3.8 Live models available on AI Gateway
WebSocket-based realtime audio API with parallel reasoning and 97-language auto-switching replaces polling-based voice integration patterns.
Eliminates latency gaps in voice assistant flows by streaming audio directly and running Extended Thinking reasoning in parallel with speech output, reducing response lag in production voice interfaces.
Drop-in replacement for polling-based voice APIs if you're already on Vercel's AI SDK. Requires WebSocket setup and short-lived token auth. Worth trying now if building voice features—the parallel reasoning variant is novel but adds provider lock-in via thinkingConfig.
“google/gemini-3.8-live-extended-thinking adds multi-step reasoning that runs in parallel with speech, allowing it to acknowledge requests and narrate progress without interrupting the conversation”
“supports real-time audio, visual grounding, automatic switching across 97 languages, and background tool calls while the conversation continues”
“Use either model through the AI SDK's realtime API”
realtime-apivoice-aiwebsocketgemini-3.8ai-gateway
Claude Desktop now integrates Ollama as gateway
Claude Desktop can now route inference through Ollama, letting you swap providers without code changes; KV cache bugs fixed to prevent token reprocessing.
Reduces friction for developers running local models or multi-provider setups. Cache restore point fixes eliminate wasteful reprocessing on cancelled requests, cutting latency and compute cost on long prefills.
Data Point
Naive Bayes beats LLMs on labeled text classification
Complement Naive Bayes matches or outperforms LLMs at 40–486x higher throughput on commodity CPUs when labeled data exists; decision boundary is task-dependent (~10^4 labels for parity on topic classification).
For teams doing text classification with labeled datasets, LLM inference cost and latency are often unjustified. This quantifies the exact threshold where classical methods become objectively superior, enabling data-driven model selection instead of default-to-LLM patterns.
Replaces zero-shot LLM pipelines for labeled text tasks. Requires: labeled training data (10K+ samples), CPU-based inference setup, and task-level benchmarking to confirm parity. The authors provide a Kubernetes Helm operator for automated model selection. Worth evaluating now if you're running batch text classification on resource budgets.
“NB reaches 89.1% accuracy, statistically indistinguishable from the zero-shot 27B LLM (89.0%) and better than the 397B frontier model (84.8%), at thousands of samples/sec on a commodity CPU”
“small-LLM batched inference is 40-486x slower than NB CPU inference”
“NB reaches LLM parity around $N \sim 10^4$ labels for topic classification, while zero-data sentiment favors the LLM at all N tested”
Replaces manual gateway configuration for Ollama users. Requires Claude Desktop update (v0.33.0+) and Ollama configured as third-party provider. Worth adopting now if you're already running Ollama; otherwise standard Claude Desktop path unchanged.
“Developers can now easily configure Claude Desktop to seamlessly work with Ollama as a third-party gateway provider”
“a cancelled prefill keeps every restore point it crossed, so retries resume where they stopped instead of restarting from scratch”
“on models with recurrent layers this previously forced a request matching 46k of 47k tokens to reprocess from zero”
Python SDK lets you publish/consume queue messages alongside JavaScript, with automatic retries and delivery guarantees via topic-based routing.
Eliminates runtime friction in polyglot stacks—Next.js producers can now fan out to Python consumers without separate infrastructure. Reduces boilerplate for background job handling in Python backends.
Replaces manual job queue setup (Celery, RQ) within Vercel projects. Requires `pip install vercel`, pyproject.toml configuration, and decorator-based subscriber definition. Ready to try now in beta; cross-runtime message routing is functional.
“Python SDK for Vercel Queues is now available in beta”
“Messages can be enqueued or consumed from either JavaScript or Python”
“automatic retries, sharding, and delivery guarantees”
“independent consumer groups process them in parallel”
Vercel Agent integrates into Slack code channels for real-time collaborative debugging, PR review, and approval workflows with full audit trails.
Eliminates context-switching between deployment tools and chat; teams can follow Agent work, review code, and enforce approval gates without leaving Slack. Reduces coordination overhead for incident response and migrations.
Replaces manual Slack-to-Vercel context switching for debugging and review workflows. Requires Slack Pro/Enterprise tier and Vercel app installed. Worth testing now if your team uses Slack heavily for incident response; approval gate requirement means it won't fully automate deployment pipelines.
“Vercel Agent now works in Slack code channels, a new kind of channel launched today for working with a coding agent”
“Agent is read-only by default and never exceeds the requester's permissions”
“Before making a change, it drafts a plan and waits for approval”
“For every action, Vercel records who requested and approved it and what Agent ran”
“available in public beta for Pro and Enterprise teams”
`connectMCPTransport` lets TanStack AI agents call OAuth-protected MCP servers with token refresh handled by Connect, eliminating credential storage.
Removes credential management overhead from agent workflows and centralizes OAuth token lifecycle. Consent errors surface as redirects before model execution, not tool failures.
Replaces manual OAuth token handling in MCP client setup. Requires TanStack AI, Vercel Connect account, and MCP server endpoint. Ready now—code sample provided; integrate via `@vercel/connect/tanstack-ai` subpath.
“Agents built with TanStack AI can now call OAuth-protected MCP servers through Vercel Connect, with no credentials for you to store or rotate”
“The provider is called before every MCP request, so the token is always fresh”
“If the user has not granted access, `createMCPClient` fails with a consent challenge before the model runs”