Two rerank models now available via unified API—accuracy-focused Rerank 3 and latency-optimized Lite variant, both handling 32K tokens per query-document pair.
Summary
Replaces separate model integrations with a single API call for relevance reordering before LLM inference, reducing boilerplate in RAG pipelines. Lite variant trades accuracy for sub-100ms latency, enabling cost-conscious retrieval without swapping implementations.
Why it matters
Replaces separate model integrations with a single API call for relevance reordering before LLM inference, reducing boilerplate in RAG pipelines. Lite variant trades accuracy for sub-100ms latency, enabling cost-conscious retrieval without swapping implementations.
Implementation verdict
Ready to use today. Set `model: 'voyage/rerank-3'` or `'voyage/rerank-3-lite'` in the `rerank()` function call. Requires AI Gateway API key or BYOK setup. Worth trying now if you're already on Vercel's platform; evaluate latency-accuracy tradeoff for your query patterns before committing to Lite.
Sources
Dev Signal
Get briefs like this in your inbox — free, every weekday.
100+ sources compressed into one 4-minute read. Ranked, cited, implementation-ready.