Gemini 3.5 Transcribe delivers sub-second streaming transcription
Two separate APIs—Live for real-time voice agents (gemini-3.5-transcribe-live), Interactions for post-recorded audio with speaker attribution—both handle custom vocabulary and 85+ languages natively.
Replaces Chirp 3 with 70% latency reduction and measurable accuracy gains (4.0% WER streaming, 2.6% non-streaming). Enables voice-first workflows in existing Gemini API integrations without architectural changes.
Public preview now in Google AI Studio and Gemini API. Requires switching from Chirp 3 model identifiers and choosing between Live (streaming) or Interactions (batch) paths. Ready to try for voice agents and transcription pipelines; production readiness depends on your WER tolerance and language mix.
- “Delivers continuous, bidirectional streaming with sub-second latency for interactive voice apps via the Live API using gemini-3.5-transcribe-live”
- “achieves an average Word Error Rate (WER) of 4.0% for streaming and 2.6% for non-streaming use-cases”
- “time to final transcription, for example, improves by 70%”
- “Automatically detects and transcribes over 85 languages”
- “In public preview in the Gemini API via Google AI Studio and Google Antigravity”
speech-to-textreal-time-apigeminivoice-agentstranscription
Cloudflare launches K2 durable event streaming
K2 is a serverless event log built on R2 object storage that decouples producers and consumers, trading ~1 second produce latency for unlimited retention and horizontal scaling across 335 edge cities.
Eliminates event loss from producer-consumer misalignment without running Kafka at the edge. Enables independent consumer scaling and long-term retention without dedicated infrastructure management.
Replaces Kafka for edge event buffering; Queues for high-scale fan-out; requires HTTP API or Worker binding integration. Public beta now—worth prototyping if you need durable multi-consumer streams, but 1s p99 latency rules out sub-second requirements.
- “Today we are launching Cloudflare K2 in public beta to solve this problem.”
- “this adds up to about 1 second of produce latency at the 99th percentile of response times”
- “K2 implements a partitioned, durable log on top of R2 object storage, which allows it to scale to huge volumes of storage”
- “Object storage systems like R2 combine extremely durable storage (11 9s!) with strongly consistent APIs”
- “Cloudflare's global infrastructure presents some challenges: we get relatively small slices of machines, those machines are relatively ephemeral, and networking is often over the public Internet”
- “it's close to users wherever they are in the world and has an incredible capacity to scale horizontally”
- “Pipelines runs on the Cloudflare edge, which spans a huge number of servers across over 335 cities”
event-streamingedge-computingcloudflareobject-storagekafka-alternative
Delta replaces pull requests with agent-aware threads
DeltaDB extends Git with delta-based versioning to preserve agent reasoning and collaborative context within threads instead of forcing code review via diffs.
Eliminates context loss when reviewers must re-examine agent-generated code without access to decision history. Developers and agents can continue work mid-thread without committing, shrinking the review surface from full diffs to focused decisions.
Replaces GitHub pull request workflow. Requires downloading Delta client (macOS/Linux/Windows) or using web version; Git repositories remain compatible. Public beta is free. Early, but 33-person team ships 570 changes without PRs—worth testing if you pair with AI coding agents regularly.
- “we disabled pull requests on Delta's own repository. We now build and collaborate on Delta entirely within Delta”
- “Delta is built on DeltaDB, which extends Git's content-based versioning with incremental versions based on deltas”
- “33 of us have landed 570 changes to main since we turned off pull requests”
git-workflowai-agentscollaborative-codingversion-control
Fish Audio models free on Vercel Gateway thirty days
Text-to-speech and transcription via AI SDK 7 with low-latency streaming and word-level timestamps; billing begins September 19 unless you use the `-free` suffix.
Removes vendor lock-in friction for audio features—test production models at zero cost before committing to per-character/per-hour rates. Speech-to-text returns word-level timing, unblocking real-time captioning workflows.
Replaces Fish Audio SDK calls with unified AI SDK functions (`generateSpeech`, `transcribe`). Requires Node.js 18+, one npm install, and model-name awareness for post-trial billing. Worth trying now if you're evaluating audio infrastructure; use `-free` suffix to auto-cutoff on September 19.
- “every Fish Audio model is free on AI Gateway for the next 30 days, through September 18”
- “fish-audio/s2.1-pro (text-to-speech): Built for low-latency streaming; clones a voice from a reference recording”
- “fish-audio/transcribe-1 (transcription): Returns the text along with the duration of the audio and timestamped segments, down to individual words”
- “Speech and transcription ship in the current AI SDK 7 release”
audio-generationspeech-to-textvercel-ai-gatewayai-sdkfree-tier
Ling 3.0 Flash Fin launches free on AI Gateway
Finance-tuned 256K-context model with 32K output tokens and function calling now available gratis through September 25—pick your model ID to control billing behavior post-trial.
Developers building financial analysis tools gain a cost-free month to benchmark domain-specific reasoning and multi-step tool workflows. The dual model ID pattern lets you test before committing to production costs.
Replaces manual Ling 3.0 Flash fine-tuning for finance use cases. Requires only model ID swap in existing AI Gateway code; no new auth or SDK changes. Worth trying immediately if you're building earnings report analysis or financial research agents—the free window is short.
- “Ling 3.0 Flash Fin from Inclusion AI is now available on AI Gateway and free to use through September 25”
- “It has a 256K-token context window, produces up to 32K output tokens, and supports reasoning and function calling”
- “Use `inclusionai/ling-3.0-flash-fin-free`. This model ID returns an error after the offer ends, preventing future charges”
llm-releasefinancial-aiai-gatewayfree-tierfunction-calling