[
    {
        "role": "user",
        "content": "what%20is%20the%20best%20method%20to%20reduce%20claude%20token%20usage%3F"
    },
    {
        "role": "assistant",
        "content": "Here are proven strategies to reduce Claude\u2019s (Anthropic) token usage\u2014meaning both input and output tokens\u2014while preserving quality and effectiveness. Many of these apply to Claude Code, the API, or usage via the web UI.\n\n---\n\n## What to measure (so you know what you're optimizing)\n\n- **Usage dashboard / `/usage` command**: See token counts, model used, per-session costs. ([code.claude.com](https://code.claude.com/docs/en/costs?utm_source=openai))  \n- **Context window usage**: Monitor how much of the allowed history is being used. Long conversation history = more repeated tokens. ([support.claude.com](https://support.claude.com/en/articles/11647753-how-do-usage-and-length-limits-work?utm_source=openai))  \n- **Effort level** and **model choice**: These two directly affect how many tokens are consumed. ([platform.claude.com](https://platform.claude.com/docs/en/build-with-claude/effort?utm_source=openai))  \n\n---\n\n## Key strategies to cut token usage\n\nHere are the top levers people use; many come from both internal Claude docs and community experiences.\n\n1. **Pick the smallest model that suffices**  \n   - Claude Sonnet (and Haiku) are much cheaper than Claude Opus for many tasks like editing, reviewing, or routine analysis. Use Opus only when you need its extra reasoning power. ([support.claude.com](https://support.claude.com/en/articles/12429409-manage-usage-credits-for-paid-claude-plans?utm_source=openai))  \n   - Smaller models often converge faster on simpler tasks, saving both tokens and time. ([claude.com](https://claude.com/blog/claude-model-and-effort-level-in-claude-code?query=Liskit&utm_source=openai))  \n\n2. **Lower the effort level when possible**  \n   - Claude supports effort levels such as `low`, `medium`, `high`, `xhigh`, and `max`. Lower effort responses consume fewer tokens. Use `max` only when required. ([platform.claude.com](https://platform.claude.com/docs/en/build-with-claude/effort?utm_source=openai))  \n\n3. **Clean context / remove unused history**  \n   - Use commands like `/clear` at natural task boundaries. Old context that\u2019s no longer needed still costs on every new message. ([code.claude.com](https://code.claude.com/docs/en/costs?utm_source=openai))  \n   - Use `/compact` to summarize long conversation histories, keeping only what\u2019s necessary. ([code.claude.com](https://code.claude.com/docs/en/costs?utm_source=openai))  \n\n4. **Be careful with tool usage and function calls**  \n   - Tools, extra actions, and non-critical plugins often add output and input tokens. Disable them when not needed. ([support.claude.com](https://support.claude.com/en/articles/11647753-how-do-usage-and-length-limits-work?utm_source=openai))  \n   - Limit file reads and large attachments\u2014only include what\u2019s relevant. Instead of uploading an entire repo, reference specific files. ([warrenmar.com](https://www.warrenmar.com/2026/06/12/reducing-token-usage.html?utm_source=openai))  \n\n5. **Prompt and content engineering**  \n   - Keep prompts concise. Avoid pasting large snippets if only part is relevant. Reference files or documentation, rather than copying. ([warrenmar.com](https://www.warrenmar.com/2026/06/12/reducing-token-usage.html?utm_source=openai))  \n   - Use clear instructions: define scope, constraints, stop\u2010conditions, etc. That helps Claude avoid unnecessary or verbose output. ([yu.phd](https://yu.phd/notes/agent/efficiency-techniques/?utm_source=openai))  \n\n6. **Caching and compression**  \n   - Prompt caching: reuse static prompts or examples that don\u2019t change across sessions. Don't resend them every time. ([builder.aws.com](https://builder.aws.com/content/3FRlppwY0rQsApCRxEksJP0s6hX/the-token-efficiency-playbook-10-methods-to-spend-less-on-llm-inference?utm_source=openai))  \n   - Tool\u2010output compression: shorten or filter tool or API call outputs before sending them into Claude. That reduces both input and output tokens. ([lineman.io](https://lineman.io/guides/reduce-claude-code-token-usage?utm_source=openai))  \n\n7. **Break large tasks into phases**  \n   - Instead of feeding Claude everything at once, do steps: define \u2192 plan \u2192 implement \u2192 debug. Each step only needs relevant context. ([yu.phd](https://yu.phd/notes/agent/efficiency-techniques/?utm_source=openai))  \n\n8. **Set quotas, usage caps, and govern usage in teams / orgs**  \n   - For enterprise usage: set spend caps, define which models are allowed for which tasks. Educate users. ([support.claude.com](https://support.claude.com/en/articles/14782391-claude-enterprise-consumption-guide?utm_source=openai))  \n   - Role\u2010based access: only give access to high\u2010token models to people who really need them. ([support.claude.com](https://support.claude.com/en/articles/14782391-claude-enterprise-consumption-guide?utm_source=openai))  \n\n---\n\n## Trade-offs to be aware of\n\n- Lower effort or smaller models sometimes reduce output quality or depth; complex reasoning might need more tokens to do well. ([claude.com](https://claude.com/blog/claude-model-and-effort-level-in-claude-code?query=Liskit&utm_source=openai))  \n- Over\u2010compressing context or removing too much history can omit needed information. Be careful to preserve important parts. ([arxiv.org](https://arxiv.org/abs/2607.12161?utm_source=openai))  \n- Frequent model switching or overuse of caching/compression layers can add complexity to your workflow. Sometimes saving tokens could cost time or introduce errors.\n\n---\n\nIf you want, I can put together a personalized checklist based on how *you* use Claude\u2014e.g. for code editing, research, etc.\u2014so you can hit big savings right away. Do you want that?"
    }
]