Multimodal
A model that handles more than one data type, such as text plus images.
A multimodal model processes and generates content across more than one data type, such as accepting both text and images as input, then producing a text response that reasons over both. On the AIF-C01 exam, this capability is associated with Amazon Bedrock, which hosts foundation models like Anthropic Claude and Amazon Titan that accept image-plus-text input, enabling visual question answering, chart interpretation, and document understanding. Note the distinction between multimodal input (a model that reads images) and multimodal output (a model that also generates images); not every model does both, so read whether the scenario needs understanding or generation before answering.
PlayPrepHQ study notes are written and reviewed against primary exam sources. How we create & review content →