Can AI Agents Read Image Attachments in Zendesk Tickets? How Macha's Vision Works (2026)
Macha's AI agents can open image attachments on Zendesk tickets, such as screenshots, receipts and product photos, and answer based on what the image shows. The agent needs the Read Attachment tool and a vision-capable model.
Key takeaways
- Macha's AI agents read JPEG, PNG, GIF and WebP attachments on Zendesk tickets through the Read Attachment tool, which sends the image to a vision-capable model.
- GPT-5.4 Mini and Claude Sonnet 4.5 are the vision-capable models confirmed in Macha's changelog for reading screenshots, product photos and receipts.
- The Groq models GPT OSS 120B and 20B, which replaced Llama 3.3 70B on 28 August 2026, cannot read images and return a not-supported message.
- Image vision for Zendesk attachments shipped on 17 April 2026 and needs only a Zendesk connector plus the Read Attachment tool on the agent.
Macha's AI agents read image attachments on Zendesk tickets through the Read Attachment tool: the agent downloads a JPEG, PNG, GIF or WebP file, sends the pixels to a vision-capable model such as GPT-5.4 Mini or Claude Sonnet 4.5, and replies based on what the image shows. The feature shipped on 17 April 2026 and needs only a Zendesk connector and the Read Attachment tool on the agent.
| Question | Answer |
|---|---|
| Help desk | Zendesk |
| Formats | JPEG, PNG, GIF, WebP |
| Setup | Zendesk connector plus Read Attachment |
Why couldn't AI agents handle image attachments before?
Support tickets don't just contain text. Customers attach screenshots of error messages, photos of damaged products, receipts for refund requests and tracking page screenshots. Before April 2026, Macha's agents could only process text attachments: PDFs, spreadsheets, plain text files. An image came back as "This file type cannot be read."
So every ticket with an image needed a human to open it before anything else could happen.
How does inline image vision work?
When an agent encounters an image attachment on a Zendesk ticket, it:
- Downloads the image from Zendesk using your connector's authentication
- Sends it to the AI model as part of the conversation, so the model sees the image pixels, not a filename
- Responds based on what it sees: describing the content, extracting text, identifying products or flagging issues
The image is processed once, and the model's description becomes part of the conversation context for the rest of the run.
What can the AI see and do in an image?
With a vision-capable model, your agent can:
- Read text in screenshots: error messages, tracking numbers, order confirmations
- Identify products: match a photo to your catalogue, spot damage or defects
- Parse receipts: pull amounts, dates and transaction IDs from photos of receipts
- Describe what it sees: "This is a screenshot of a DHL tracking page showing the parcel was delivered on April 14th"
How does image reading fit into an autonomous workflow?
Image vision runs inside Macha's trigger system. When a new Zendesk ticket arrives with an image attachment:
- The trigger fires and invokes your AI agent
- The agent reads the ticket: subject, description, custom fields, conversation history
- The agent sees the image attachment and calls the Read Attachment tool
- The model analyses the image and folds what it saw into the response
- The agent posts a reply or internal note based on the full context, text and visual
To have a person check the output first, point the agent at internal notes.
Which models support vision?
Not every model in the picker can process images. Macha's changelog confirms these:
| Model | Vision | Notes |
|---|---|---|
| GPT-5.4 Mini | Yes | Confirmed in the April 2026 release; low cost, 400K context window |
| Claude Sonnet 4.5 | Yes | Confirmed in the April 2026 release; detailed image descriptions |
| GPT OSS 120B / GPT OSS 20B (Groq) | No | Text only; the tool returns a "not supported" message |
GPT OSS 120B and 20B replaced Llama 3.3 70B on 28 August 2026 and cannot read images either. For newer models in the picker (GPT-5.6 Terra, GPT-5.6 Luna, GPT-5.4, Claude Sonnet 5), test an image ticket before relying on vision. On a non-vision model, the tool suggests which models to switch to instead of failing the run.
How do you turn image reading on?
Macha has one plan with every feature, priced by ticket volume from $299 a month for 750 tickets, so image vision is included wherever there is a Zendesk connector. Connect Zendesk, add the Read Attachment tool to your agent, and pick a vision-capable model. The agent detects image attachments and processes them with that model.
Add AI agents to your Zendesk
Macha reads the ticket, drafts the reply and takes the action, inside the Zendesk you already run.
Intercom
Shopify
Stripe
Slack
Notion
Google Workspace
Confluence

