Macha

How Big a Context Window Does an AI Support Agent Need in 2026?

Abbas, Customer Support & AI, Macha

Written by

Ankeet Guha, Co-founder & CTO, Macha

Reviewed by

Published May 11, 2026

Updated September 24, 2026

A typical AI support run uses 15,000 to 25,000 tokens, far below the 131,072 to 1,050,000 token windows of the models Macha runs. Window size only starts to matter on very long threads and agents that chain many tool calls.

Key takeaways

  • A typical Macha support run that reads a ticket, checks an order and drafts a reply uses 15,000 to 25,000 tokens, under every model's context window.
  • GPT-5.4, Macha's default model, has a 1,050,000 token context window, while GPT-OSS 120B on Groq has the smallest at 131,072 tokens.
  • GPT-5 Mini and GPT-5.4 Mini both list 400,000 token context windows in OpenAI's docs, so the old 128K versus 400K split no longer applies.
  • Macha compacts a conversation once it crosses 55% of the model's context window, summarizing the middle while keeping the first and most recent messages.
  • An agent chaining 5 to 6 tool calls in one run can consume 20,000 to 40,000 tokens from tool results alone.
How Big a Context Window Does an AI Support Agent Need in 2026?

Almost every model a support team would pick in 2026 has a big enough context window: a typical Macha support run uses 15,000 to 25,000 tokens, and the smallest window among Macha's models is 131,072 tokens on GPT-OSS 120B.

What is a context window?

A context window is the total amount of text, measured in tokens (roughly 4 characters of English per token), that a model can process in one request. Everything the model reads and writes has to fit inside it:

  • The system prompt (agent instructions, tool definitions, knowledge context)
  • The conversation history (every previous message in this interaction)
  • Tool results (data returned from your help desk, APIs, knowledge base searches)
  • The model's own response

If the total goes over the window, the request errors out or older information has to be dropped.

How big are the context windows of the models Macha runs?

ModelContext windowSource
GPT-5.4 (Macha's default)1,050,000 tokensOpenAI model docs
Claude Sonnet 51,000,000 tokensAnthropic models overview
GPT-5400,000 tokensOpenAI model docs
GPT-5.4 Mini400,000 tokensOpenAI model docs
Claude Sonnet 4.5200,000 tokensAnthropic models overview
GPT-OSS 120B (on Groq)131,072 tokensGroq model docs

Figures are from OpenAI, Anthropic and Groq, checked 24 September 2026. GPT-5 Mini and GPT-5.4 Mini both list 400,000 tokens, so the old "128K vs 400K" gap between them no longer exists.

When does context window size matter in support?

How much do long ticket histories add?

A customer who has gone back and forth 15 times generates a lot of context. Each message, plus the agent's replies, tool calls and tool results, adds to the count. On a 131K model a very long conversation can get close to the limit.

How much do tool results add?

When your agent fetches a ticket with a long description, searches a knowledge base and reads a multi-page document, each result adds thousands of tokens. An agent that chains 5 to 6 tool calls in one run can consume 20,000 to 40,000 tokens from tool results alone.

How much do agent instructions take up?

Detailed instructions (WISMO classification rules, response templates, business logic) run 2,000 to 5,000 tokens. Add tool definitions for 10 or more tools and you're at 8,000 to 10,000 tokens before the first customer message.

What happens when the window fills up?

Macha manages the limit with conversation compaction. When the estimated token count crosses 55% of the model's context window, the middle of the conversation is summarized, keeping the first message (the original intent) and the most recent messages (the current context) intact.

The arithmetic shows how rarely that fires. At 55%, compaction starts around 72,000 tokens on GPT-OSS 120B, 220,000 on GPT-5.4 Mini and 577,500 on GPT-5.4. The summarized middle loses detail: if the agent needs an order number from 10 messages back, it may not have it after compaction.

Which model should you choose for long conversations?

For most support agents, any model in the table is enough. A typical autonomous run (read ticket, check order, draft response) uses 15,000 to 25,000 tokens, well under even the 131,072-token window. A 1M-class model like GPT-5.4 or Claude Sonnet 5 earns its place when an agent handles very long threads, chains many tool calls, or reads tickets with extensive history.

Pick the model on task complexity and conversation length, not raw window size. The choice doesn't change what Macha charges either: billing is per ticket, one charge per conversation, whichever model the agent runs on.

Macha

About Macha

Macha is an AI agent platform that works on top of the help desk you already use — Zendesk, Freshdesk, Gorgias, or Front — and connects to the rest of your stack, even your own internal systems. Its AI agents resolve tickets and automate entire workflows end to end, all set up in plain English, no code. Learn more about Macha →

Zendesk
5.0 on Zendesk Marketplace

Loved by support teams worldwide

See what support teams are saying about Macha AI.

The application seems excellent to me! We are still testing, and we need support for some details and they were extremely efficient too!

Daniela Costa

Daniela Costa

Head of Support, Seabra

Macha has been a great addition to our support toolkit. It generates clear, well-organized responses that fit naturally into our workflow. One feature we particularly appreciate is its ability to automatically reply in the same language as the ticket.

Marius F

Marius F

Support Head, Zentana

We've been using Macha for a little while now and it's been really great addition so far! It's powerful, convenient, and makes getting work done a lot easier for our agents.

Alexander Wedén

Alexander Wedén

Head of Support

Support team is very helpful and responsive. Really enjoy how lightweight this is within Zendesk itself vs other more intrusive tools.

Cathleen Wright

Cathleen Wright

Zendesk Admin, Cortex IO

So far it's pretty good! Our queries are a little nuanced, so we can't always use it, but it's got enough utility for us. It can even incorporate our bilingual country with greetings in a second language.

Jae Oliver

Jae Oliver

Head of Support, Wise

Really enjoying using Macha, it has made a noticeable difference to our support team in a short amount of time. I really like the ticket summary feature, saves us a lot of time.

Harry Jackson

Harry Jackson

Head of Support, Crumb

Macha AI is a great addition to my workspace! It's powerful, convenient, and it really makes productivity so much easier for our agents!

Dave G

Dave G

Head of Support, Cyber Power Systems

Very impressed! AI integration for Zendesk has certainly come a long way and Macha seems to set the standard for now. This will for sure save lot of time in our support team.

Pauli Juel

Pauli Juel

Head of CS, Dokument24

Macha has been working great for us so far! The auto-responses are accurate and our resolution time has dropped significantly.

Lana T

Lana T

Zendesk Admin, Swotzy

Macha AI is a great addition. The knowledge base feature means our agents always have the right answers at their fingertips.

Mischa Wolf

Mischa Wolf

Head of Support, Topi

We're enjoying this integration so far. It's made our support team more efficient and our customers get faster responses.

Paula G

Paula G

Head of Customer Support, Xly Studio

The team enjoys using it. It saves considerable time on common questions and the integration options are excellent.

Kilian Leister

Kilian Leister

Support Head, Didriksons

Ready to supercharge your team with AI?

Get started in minutes. Connect your tools, configure your agents, and let AI handle the rest.

$50 in free credits · no time limit, no credit card