System Design Interview Roadmap

System Design Interview Roadmap

Multi-Modal System Design — Handling Text, Image, and Audio Simultaneously

Oct 09, 2026
∙ Paid

A customer uploads a 40-second phone video of a leaking pipe and types “is this covered under my policy?” One request, three payloads: a text question, an audio track, and roughly 1,200 video frames. Most teams building this feature start by wiring the video into a vision model, the audio into a transcription service, and the text straight to the LLM, then stitching the three outputs together in a prompt. It works in the demo. It falls apart at 200 requests per second, when the video encoder queue backs up, the transcription service returns four seconds after the vision model already timed out, and the LLM ends up reasoning over a transcript-shaped hole. Multi-modal system design is mostly about not letting that happen.

The Core Problem: Modalities Don’t Move at the Same Speed

A multi-modal system ingests inputs of fundamentally different sizes, formats, and processing costs, then produces a unified response. The naive architecture treats this as three parallel single-modal pipelines that converge at the end. The problem is that text tokenizes in microseconds, a 512×512 image embeds in 30–80ms on a decent GPU, and a 40-second audio clip takes 1–3 seconds to transcribe even with a fast ASR model. If your orchestrator waits for all three before calling the LLM, your p99 latency is set by your slowest modality, every time, regardless of how fast the other two are.

The fix isn’t a faster audio model. It’s decoupling ingestion from fusion. Each modality gets its own preprocessing pipeline — resize and encode for images, chunk and transcribe for audio, tokenize for text — running independently and writing normalized output to a shared store (commonly Redis or a lightweight message queue keyed by request ID). A fusion layer watches for all required pieces to arrive, or a timeout to expire, before assembling the final prompt.

That timeout decision is where most of the design tension lives. Do you wait for the slow modality, or degrade gracefully and answer with partial context? Insurance claims can’t guess at what’s in a video. A customer support chatbot answering “what’s this error mean” from a screenshot plus text can often proceed on text alone if the vision pipeline is backed up, flagging the response as partial. The right answer is domain-specific, but the architecture has to support both paths — most teams only build the happy path and discover the gap during an incident.

Fusion itself has two common strategies. Early fusion projects every modality into a shared embedding space (think CLIP-style joint embeddings) before any reasoning happens — this is what lets a model compare “a photo of a dog” against an actual photo directly. Late fusion keeps modalities separate through their own specialized models and only combines outputs at the decision layer, e.g., concatenating a transcript, an image caption, and the user’s text into a single LLM prompt. Late fusion is easier to build and debug, since each component fails independently and visibly. Early fusion captures cross-modal relationships that late fusion misses entirely — a joint embedding model can tell you the audio tone doesn’t match the words, while three separate pipelines never make that comparison at all.

User's avatar

Continue reading this post for free, courtesy of System Design Roadmap.

Or purchase a paid subscription.
© 2026 SystemDR Inc · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture