Three Things Llama 4 Clarifies About Open Multimodal Brand Summaries

This article was AI-generated as part of an experimental historical-content project. The date reflects the period being analyzed rather than the date the article was originally written.

Meta released Llama 4 on Saturday, an odd day for a major launch, and the three days of reaction since have been as instructive as the release itself.

The facts first. Meta introduced two open-weight models, Llama 4 Scout and Llama 4 Maverick, and previewed a much larger one, Llama 4 Behemoth, which is still training and has not been released. Meta calls Scout and Maverick “the first open-weight natively multimodal models,” and they are its first built on a mixture-of-experts design. Scout has a context window of 10 million tokens. The weights are on llama.com and Hugging Face, and TechCrunch reported that Meta AI in WhatsApp, Messenger and Instagram now runs on Llama 4 in 40 countries, with multimodal features limited to English in the U.S. for now.

Last fall, Llama 3.2 showed that open vision models could take brand summaries out of the big cloud services. Llama 4 sharpens three points.

1. Multimodal is becoming the default

Llama 3.2 offered vision as separate model sizes. Llama 4 is multimodal from the start: Meta says it uses “early fusion” to train text and vision together in one model backbone. The model that summarizes your annual report can also read your product photo, your packaging, a screenshot of your app or a chart from your investor deck, in the same pass. A brand description built partly from images is no longer a specialist case. Anyone who downloads the weights gets it.

Scout’s long context window points the same way. It can take in an entire archive of releases, filings and coverage at once. That does not make its conclusions more accurate. It does make them feel more thorough.

2. The same weights can describe you differently

The most useful lesson came from launch week’s problems. Developers reported uneven results, and Ahmad Al-Dahle, who leads generative AI at Meta, acknowledged “mixed quality across different services.” He said that “since we dropped the models as soon as they were ready, we expect it’ll take several days for all the public implementations to get dialed in,” according to TechCrunch. He also denied rumors that Meta trained on benchmark test sets.

Separately, the leaderboard score Meta promoted came from “an experimental chat version” of Maverick, not the one released. The Register looked at the published comparisons and found the experimental version verbose and often full of emojis, and the public one far more concise.

For brands, the benchmark dispute is beside the point. What matters is that “Llama 4” is not one voice. Tuning changed its tone. Each host chooses its own settings, system prompts, retrieval and filters. Two apps built on identical weights can describe the same company differently, and neither description is “what Meta says.” Tracking a model name is less useful than tracking the products people actually use.

3. Meta’s app and Meta’s weights are separate maps

Meta AI is a consumer product with a known owner, a feedback channel and a policy team. Llama 4 weights running inside somebody else’s app have none of those. Even the open distribution has borders: TechCrunch notes the license bars users and companies domiciled in the EU, and requires a special license for companies with more than 700 million monthly users.

So the map has two layers. One is a few Meta surfaces with enormous reach, where an error can at least be reported to a company. The other is a long tail of apps, internal tools and hosted services you will mostly never see. As I suggested when Gemini launched, the question is increasingly who summarizes you first. With open weights, the answer can be almost anyone.

None of this calls for alarm. It calls for precision. When someone says “Llama got us wrong,” the first question is which product, which version and what retrieval. The second is whether the facts it would find about you, in text and in images, are consistent.