Which AI Model Should Your Support Agent Use? GPT-5.6, Claude Sonnet 5 and GPT OSS Compared (2026)
Most Macha support agents should start on GPT-5.4 and move up to GPT-5.6 Terra or Claude Sonnet 5 only for complex tickets. Macha charges per ticket whichever model runs, so the choice is about quality, speed and image reading.
Key takeaways
- GPT-5.4 is the default model for new Macha agents and suits ticket classification, reply drafting and tool calling, with switching recommended only when an agent hits a limitation.
- Macha charges per ticket regardless of which AI model handles it, so choosing between GPT-5.6 Terra, Claude Sonnet 5 or GPT OSS 20B is a decision about quality and speed, not budget.
- GPT-5.6 Terra became Macha's flagship model on 27 July 2026, with a context window of roughly 1M tokens and the strongest multi-step reasoning in the line-up.
- Groq's GPT OSS 120B and GPT OSS 20B replaced Llama 3.3 70B, Llama 3.1 8B and Mixtral 8x7B on 28 August 2026, and neither GPT OSS model reads image attachments.
- Each Macha agent sets its model independently, so a triage agent can run on GPT OSS 20B while an escalation agent runs on GPT-5.6 Terra.
For most support agents on Macha, start on GPT-5.4, the default for new agents. Move to GPT-5.6 Terra or Claude Sonnet 5 for complex, high-stakes tickets, and to a fast model such as GPT-5.4 Mini or GPT OSS 20B for tagging and routing. Macha charges per ticket, and the model doesn't change that charge: one conversation is billed once however many replies, tool calls or lookups it takes. So the choice comes down to quality, speed and whether the agent needs to read images.
| Model | Provider | Best for | Reads images |
|---|---|---|---|
| GPT-5.6 Terra | OpenAI | Best answer quality, multi-step reasoning, tool selection | Yes |
| GPT-5.6 Luna | OpenAI | High-volume queues that want near-Terra quality, faster | Yes |
| GPT-5.4 (default) | OpenAI | Everyday agents with strict instruction-following | Yes |
| GPT-5 | OpenAI | Balanced replies at volume | Yes |
| GPT-5.4 Mini | OpenAI | Simple, fast tasks; 400K context window | Yes |
| Claude Sonnet 5 | Anthropic | Nuanced writing, tricky escalations | Yes |
| Claude Sonnet 4.5 | Anthropic | Careful reasoning, structured output | Yes |
| GPT OSS 120B | Groq | Cheap, fast, tool-capable work | No |
| GPT OSS 20B | Groq | The fastest option for tagging and intent detection | No |
Which AI models can a Macha agent run on?
Macha offers nine models from three providers, set per agent. The line-up changed in summer 2026, so older guides name models you can no longer pick.
OpenAI
- GPT-5.6 Terra: Macha's flagship since 27 July 2026, with a context window of roughly 1M tokens. Sharpest on multi-step reasoning and tool choice.
- GPT-5.6 Luna: the mid-tier. Faster than Terra, with a small step down in quality.
- GPT-5.4: the default for new agents. Better than GPT-5 at strict rules like "do this exactly once" and complex branching.
- GPT-5: balanced, and quick enough to run at volume.
- GPT-5.4 Mini: small and fast, with a 400K context window.
Anthropic
- Claude Sonnet 5: added on 26 August 2026, replacing Claude Sonnet 4. Strong at careful reasoning and nuanced writing.
- Claude Sonnet 4.5: still available, with consistent, structured output.
Groq (open-weight models, fast inference)
- GPT OSS 120B: OpenAI's larger open-weight model, served by Groq. A quick all-rounder for tool calls. No image vision.
- GPT OSS 20B: the fastest option, for tagging and routing. No image vision.
Llama 3.3 70B, Llama 3.1 8B and Mixtral 8x7B left the picker on 28 August 2026. GPT-5 Mini and GPT-4o Mini were retired earlier, and agents on them moved to GPT-5.4 Mini.
Which model should a support agent start on?
For most support agents: GPT-5.4
New agents start here. It's fast enough for real-time replies and follows instructions closely enough for classification, reply drafting and tool calling. Switch only when you hit a limitation.
Which model is best for complex or high-stakes tickets?
If your agent handles escalations, refund decisions or multi-step reasoning (check the order, compare it with the return policy, decide eligibility, draft the reply), a stronger model makes fewer mistakes. GPT-5.6 Terra is the pick when you want the best answer on every ticket. Claude Sonnet 5 is the alternative when the reply's tone and wording matter as much as the logic. Neither changes what the ticket costs, so choose on answer quality.
Which model is fastest for high-volume, simple tasks?
Tagging, intent classification and routing don't need deep reasoning. Use the fastest model that gets it right: GPT OSS 20B or GPT-5.4 Mini. The price per ticket is the same, so you're choosing speed.
What about multilingual support?
The Llama model this guide used to recommend for European languages is retired. If you handle tickets in German, French, Spanish or Italian, run past tickets in each language through two or three models with Agent Simulations and compare the replies.
Which models can read image attachments?
If your tickets include screenshots, product photos or receipts, you need a vision-capable model. The OpenAI and Anthropic models support vision. The Groq-hosted GPT OSS models do not, and return a "not supported" message when they meet an image.
Can different agents use different models?
Yes. Each agent's model is set independently: a triage agent on GPT OSS 20B for speed, an escalation agent on GPT-5.6 Terra or Claude Sonnet 5, a WISMO agent on GPT-5.4. Existing agents keep their model until you change it.
Resolve tickets automatically with AI agents
Macha's AI agents work on top of the help desk you already use — no code.
Intercom
Shopify
Stripe
Slack
Notion
Google Workspace
Confluence

