Multimodal AI is artificial intelligence that works with several media types at once – text, images, sound and video – in the same model. It can describe an image, talk about a document and turn a conversation into a diagram.

The first language models could do only one thing: text in, text out. Multimodal models break down that boundary. You can photograph a whiteboard and get minutes, speak to the model instead of typing, or ask it to turn a spreadsheet into an illustration.

The consequence is that AI moves from the desk out into the rest of the work – meetings, presentations, production and documentation. That widens the possibilities, but also the questions about quality, rights and judgement that text-based AI has already raised.

In practice

A concrete example: you record a client meeting, and the model writes the minutes, pulls out the decisions, suggests a follow-up email and makes a slide for the steering group – four media types in one workflow.

How to explain it to senior management

“Multimodal means that AI now works in all our formats. That is why we can no longer limit the ground rules to the written word.”

I explore what this means for creative work on the page about AI in creative work.