SIGNAL//DESK
AI/MLsrc: curated AI glossary

Multimodal

A multimodal AI is like a person who can see, hear, and read all at once, allowing it to understand and create things using a mix of pictures, sounds, and words instead of just typing.

A multimodal model is an AI architecture designed to ingest or output multiple data types—such as text, images, and audio—by mapping them into a unified vector space, allowing the system to relate different media formats to one another.

A model that processes and/or generates more than one modality (text, image, audio, video) within a shared representation, enabling cross-modal tasks.

evolution

  1. 2015 · history
    Show and Tell

    Google researchers introduced a neural network that could automatically generate captions for images, marking a foundational step in vision-language integration.

  2. 2017 · history
    Transformer Architecture

    The introduction of the Transformer model provided the unified mathematical framework necessary to process diverse data types as sequences.

  3. 2021 · history
    CLIP

    OpenAI released CLIP, which demonstrated that contrastive learning could effectively bridge the gap between visual concepts and natural language.

  4. 2022 · history
    DALL-E 2

    The release of DALL-E 2 showcased high-fidelity text-to-image generation, bringing multimodal AI into mainstream public awareness.

  5. 2023 · history
    GPT-4V

    The integration of vision capabilities into large language models enabled native multimodal reasoning and interaction.


← all terms