Quarter-size model matching Inkling performance, controllable thinking effort, native multimodal reasoning—set model to `thinkingmachines/inkling-small` in AI SDK.
Summary
Reduces inference cost and latency for agentic coding and tool-use workflows while maintaining reasoning quality. Zero Data Retention option available per-request, enabling compliance-first deployments without vendor lock-in.
Why it matters
Reduces inference cost and latency for agentic coding and tool-use workflows while maintaining reasoning quality. Zero Data Retention option available per-request, enabling compliance-first deployments without vendor lock-in.
Implementation verdict
Drops into existing AI SDK calls as a model swap. Requires no setup beyond selecting the model identifier and optionally enabling ZDR flag. Worth trying now for cost-sensitive coding agents and multimodal document workflows; benchmark against your current model first.
Sources
Dev Signal
Get briefs like this in your inbox — free, every weekday.
100+ sources compressed into one 4-minute read. Ranked, cited, implementation-ready.