What it does
llm-router runs locally next to Claude Code, Codex and Gemini CLI. It reads each prompt and sends it to the cheapest model that can handle it, compresses tokens, and falls back to another provider when one fails. Routed answers are advisory by default, so Claude still takes the turn unless you change the setting.
Use cases
- 01Send simple questions to a free or local model
- 02Fall back to another provider when one is unavailable
- 03Reduce token use in long coding sessions