Profession-specific system prompts in Scientific Agents corpus increased token output 1.5–2.3× and per-call costs 2.2–4.5× without measurable accuracy improvement across 9 benchmarks, while actively degrading performance on tool-using tasks by 10 percentage points.
Summary
Teams optimizing LLM inference costs should benchmark prompt length against accuracy gains before deploying verbose domain profiles. This work quantifies the real trade-off: longer prompts don't reliably improve reasoning on scientific tasks and burn budget on repetitive API calls.
Why it matters
Teams optimizing LLM inference costs should benchmark prompt length against accuracy gains before deploying verbose domain profiles. This work quantifies the real trade-off: longer prompts don't reliably improve reasoning on scientific tasks and burn budget on repetitive API calls.
Implementation verdict
Replace verbose domain prompts with minimal baselines for Gemini 3.5 Flash on text-based science tasks. The only exception: if your provider has API reliability issues, longer prompts may improve first-pass success rates due to retry mechanics, not intelligence. Not production-ready for cost-conscious deployments—benchmark your own tasks before committing to profile injection.
Sources
Dev Signal
Get briefs like this in your inbox — free, every weekday.
100+ sources compressed into one 4-minute read. Ranked, cited, implementation-ready.