Quick Tunnels gain email access control via flag
cloudflared 2026.9.3 adds --allowed-mail to gate Quick Tunnel access by email address with one-time PIN verification, no account required.
Agents can now share local dev servers safely without exposing them publicly. Developers can enforce access policy at the connector level while keeping guest lists off Cloudflare's servers.
Replaces manual DNS/dashboard access control for temporary shares. Requires cloudflared 2026.9.3+ and updating agent instructions to pass --allowed-mail flags. Ready now: single-flag syntax, works in wrangler, scales from one email to wildcard domains.
- “Starting with cloudflared 2026.9.3, you can add --allowed-mail to the command, and your Quick Tunnel only lets in the email addresses and domains you choose.”
- “your guest list never leaves your machine. Cloudflare learns that a tunnel requires email authentication. It doesn't learn who you invited.”
- “Add --output json and every log line becomes a JSON object, so the agent can pick out the URL without scraping text.”
- “Visitors prove they own one of those addresses with a one-time PIN from Cloudflare Access. Nobody, on either side, needs a Cloudflare account.”
cloudflare-tunnelaccess-controlagent-workflowsdev-tools
Kolibri delivers 1M-token context at 3B active parameters
MoE model optimized for German/English with 78B total, 3B active parameters—on-premise deployable with abstention training for regulated sectors, available on Hugging Face under Apache 2.0.
Enables sovereign deployment in compliance-heavy domains (public admin, aerospace, manufacturing) without exfiltrating data to third-party inference. On-device inference reduces latency and regulatory friction for teams building mission-critical EU/German workloads.
Replaces closed-model inference for regulated German-language tasks. Requires on-premise GPU capacity and familiarity with MoE routing. Worth evaluating now if you target public sector or automotive verticals—benchmarks show parity with 4x larger active-parameter models (Nemotron 3 Super 120B-A12B). Production-ready; full model weights and tech report available.
- “78B total parameters, 3B active”
- “context lengths of up to 1M tokens”
- “Kolibri matches models with up to four times its active parameter count, such as Nemotron 3 Super”
- “21.3% of the pre-training tokens are German”
- “trained to say "I don't know" when the answer isn't in the context”
- “The model can be downloaded with the full weights on Hugging Face and used under the Apache 2.0 license terms”
mixture-of-expertsgerman-languagesovereign-ailong-contexton-premise
ThinkingBox grades agents on backend state, not tool calls
Measure agent reliability by executing tasks 20 times against isolated backends and checking terminal state changes, not parsing trajectories—reveals 67% of clean-looking failures have wrong field values or missing side effects.
Single-attempt benchmarks hide consistency gaps: Claude Opus 5.5 retains only 71% of its pass@1 score across 20 repeats. You need to know whether your agent will execute stateful workflows correctly every time, not just once.
Replaces pass@1-only leaderboard thinking with pass@20 and observed 20/20 metrics. Requires isolated MCP tool sessions and executable state checks on real backend objects. Available now via Hugging Face—worth running on your own workflows before production deployment to surface silent failures current evals miss.
- “grades agents on terminal backend state and side effects”
- “79,853 attempts failed the executable checks. Of those failures, 67.24% still terminated cleanly, invoked a state-changing tool, and reported no final tool error”
- “wrong field values in 77.61% of them, unintended extra effects in 43.30%, and missing required effects in 25.36%”
- “Only three hold on to most of their pass@1 scores: GPT-6 Astra retains 78% of its single-attempt rate, and Claude Opus 5.5 and Claude Opus 5 each retain 71%”
- “Kimi-K3 has the broadest coverage of any model we tested. It solves 93.89% of the benchmark at least once”
agent-benchmarkingstateful-workflowsreliability-testingmcp-toolsevaluation
Radicle patches cleartext protocol vulnerabilities
All Radicle versions leak private repository contents over the network—network traffic is unencrypted and peer authentication is broken—requiring immediate repo blocking and a major version bump to fix.
If you're syncing private repos via Radicle, assume they're compromised; anyone observing the network path between nodes reads the data. You must stop seeding private repos now and rotate any credentials stored in them.
Stop using private repositories over Radicle network immediately. Use `rad ls --private --all` to list affected repos, then `rad block <RID>` on each. Treat all private repos synced to other nodes as leaked. Fix requires replacing the transport layer with iroh (backwards-incompatible major release). Not production-safe for private code until patched.
- “Network traffic between nodes is not encrypted and not authenticated”
- “All versions of Radicle that were released to date are vulnerable”
- “Anyone who can observe the network path between two nodes can read the data they exchange as the data is sent in plain text”
- “The resolution involves replacing Radicle's networking protocol (currently a custom protocol using Noise) with iroh”
- “We recommend to stop using private repositories until a fix is released”
radiclesecurity-vulnp2ptransport-layergit-collaboration
Multi-provider routing masks silent document drops
Failover logic that retries across providers without checking capability support will silently drop attachments, charge for hallucinated responses, and hide failures behind success codes.
Document processing APIs that silently discard inputs and return confident false results are worse than errors—they corrupt production data while appearing successful. Developers building multi-model routers must validate provider capability, not just key presence.
Replaces naive failover (check key exists) with capability-aware routing (ask provider if it handles this file type before sending). Requires per-provider feature detection methods and explicit format support matrices. Worth implementing immediately if you route across models with different input support.
- “Looking at a key cannot tell you it was revoked. Only a call finds out.”
- “Both build their request from the prompt and the message history alone. The document was dropped without a word.”
- “A confident invention on a paid endpoint is worse than the error it replaced, because the error is visible and the invention is not.”
- “a candidate must be able to SERVE the call, not merely hold a key”
multi-model-routingerror-handlingdocument-processingapi-designprovider-detection