Agentic LLMs like Sonar Deep Research outperform base models by 60% on lit reviews, but still lose to humans 77% of the time; expert-calibrated evaluator (LitJudge) now available to measure synthesis quality.
Summary
If you're building or evaluating AI-assisted research tools, existing LLM-as-judge methods misalign with expert preferences (ρ=0.467). LitJudge fixes this gap (ρ=0.78), letting you optimize for criteria that actually matter: structure, synthesis, research suggestions.
Why it matters
If you're building or evaluating AI-assisted research tools, existing LLM-as-judge methods misalign with expert preferences (ρ=0.467). LitJudge fixes this gap (ρ=0.78), letting you optimize for criteria that actually matter: structure, synthesis, research suggestions.
Implementation verdict
This replaces generic ROUGE/overlap metrics for lit review evaluation. Requires domain expertise to validate outputs; code and data are public. Worth adopting now if you're building research assistants, but don't expect to replace humans—current best systems win only 23% of head-to-head matches on overall utility.
Sources
Dev Signal
Get briefs like this in your inbox — free, every weekday.
100+ sources compressed into one 4-minute read. Ranked, cited, implementation-ready.