AI Mindset · Model Cheatsheets
Meta AI & Llama

Two Meta AI Stories You Should Not Confuse

Meta now runs one proprietary API and one download-only open-weight family. The Meta Model API serves Muse Spark 1.1 and no Llama models; Meta retired its hosted Llama API on July 6, 2026 and points developers to third-party hosts instead. Muse Spark 1.1 also powers the agentic Meta AI capabilities announced on July 24. The business decision starts by choosing the track.

Verified August 3, 2026Muse Spark 1.1Meta Model API: Muse onlyLlama 4: download onlyAgentic Meta AI, July 24
1 / Mental Model

Meta’s own API is now Muse only; Llama is download only

The clean split of “proprietary Muse for consumers, hosted Llama for developers” no longer holds. Meta retired the hosted Llama API on July 6, 2026 and its migration guidance sends developers to third-party hosts. The endpoint Meta still operates serves Muse Spark 1.1 and no Llama models.

Choose the Meta track

Start with the outcome and control requirement.

Use Meta’s assistant, apps or API
Muse Spark 1.1, hosted by Meta

Use the proprietary Muse family when the value comes from Meta AI, social content, messaging, voice, camera, shopping and glasses, or when a Meta-operated endpoint is acceptable for a build.

  • No downloadable Muse weights
  • Meta Model API serves Muse Spark 1.1 only
  • The only model track Meta still hosts
Hosted by Meta

Muse Spark 1.1

Built by Meta Superintelligence Labs, it powers the Meta AI assistant and meta.ai, and it is the model behind the agentic capabilities announced on July 24.

  • Reasoning and multimodality
  • Voice, camera and personal context
  • The only model on the Meta Model API
Retired

The hosted Llama API

Meta shut the Llama API down on July 6, 2026. Meta’s deprecation notice states that on that date the service shuts down and API requests return a sunset response with redirect guidance.

  • No Meta-run Llama inference
  • Migration guidance points to third-party hosts
  • Not a deprecation of the models themselves
Download only

Llama 4

Scout and Maverick are natively multimodal open-weight models. They remain available for download through the Meta Llama downloads page and through partners.

  • Deployment control
  • Text and image understanding
  • Llama Guard and Stack ecosystem
2 / What’s Current

Muse Spark 1.1 is the live model; Llama 4 is still the open baseline

Meta’s newest public releases are all on the Muse side: version 1.1 of the model, the Meta Model API preview and the agentic assistant capabilities announced on July 24. There has still been no documented open-weight successor to Scout and Maverick, and no material first-party change appeared between July 28 and August 3, 2026.

July 6
Meta’s hosted Llama API shut down
July 9
Muse Spark 1.1 and the Meta Model API public preview
July 24
Agentic Meta AI announced, powered by Muse Spark 1.1
10M
Supported context window for open-weight Llama 4 Scout

Muse Spark 1.1 Current

A multimodal reasoning model built for agentic tasks. Meta reports major gains in tool and computer use, coding and multimodal understanding, and describes the model as actively managing its own 1M-token context: remembering actions, retrieving earlier work and compacting without losing steps it will need later.

  • Meta reports strong multi-application computer use
  • Meta’s description cites large code migrations and enterprise bug fixing
  • Vendor claims — evaluate on your own workload
  • Powers Meta AI and meta.ai, reconfirmed July 24

Meta Model API Public preview

Meta’s own developer endpoint serves Muse Spark 1.1 through an OpenAI-compatible interface. It carries no Llama models, so it is a proprietary Muse endpoint rather than a general Meta model service. The model is live in Thinking mode in the Meta AI app and on meta.ai.

  • Still not downloadable weights
  • OpenAI-compatible package eases migration
  • Public preview as of Meta’s July 9 post; no GA announcement found

Agentic Meta AI July 24

Meta announced assistant capabilities that act rather than only answer: daily briefings, recurring tasks, web research and slide generation. Meta states these are powered by Muse Spark 1.1.

  • Rolling out that day in select markets
  • WhatsApp support “in the coming weeks”
  • Do not plan around universal availability

Muse Image Shipped

Meta’s image generation model, combining prompt reasoning, web context and multiple visual references inside the Meta AI experience.

  • Generate and edit
  • Blend multiple photo references
  • Text, QR codes and presets

Muse Video Announced, not shipped

Announced alongside Muse Image on July 7, 2026 as “coming soon to creators and Meta AI,” and described by Meta as already in development. Nothing since has shipped it.

  • No availability date published
  • Do not scope work against it
  • Treat as roadmap, not capability

Llama 4 Scout Open weight

A 17B-active, 109B-total model with 16 experts, native multimodality and a supported 10M context window.

  • Efficient deployment target
  • Fits on one H100 with Int4 quantization according to Meta
  • Long-context and image grounding

Llama 4 Maverick Open weight

A 17B-active, 400B-total model with 128 experts for stronger general multimodal, reasoning and coding work.

  • Higher serving requirement than Scout
  • General-purpose open model
  • Available from Meta and partners

Llama Guard 4

Meta’s safety model accompanies the Llama 4 ecosystem for classification and protection workflows.

  • Part of a broader safety system
  • Requires task-specific evaluation
  • Does not replace application controls

Behemoth Not released

Meta described the Llama 4 teacher model as still training in the original announcement. Do not include it as an available model.

  • No current downloadable model
  • No procurement assumption
  • Update only from a new official release
ModelAccessActive / total parametersContextPrimary use
Muse Spark 1.1Meta products; Meta Model API public previewNot publicly disclosed1M, actively self-managedMeta AI assistant, agentic tasks and product experiences
Muse ImageMeta AI and selected Meta appsNot publicly disclosedVisual conversation contextImage generation and editing
Muse VideoAnnounced July 7, 2026; not yet availableNot publicly disclosedNot publicly disclosedVideo generation, on the roadmap
Llama 4 ScoutDownload and partners; no Meta-hosted API17B / 109B10M supportedEfficient long-context multimodal deployment
Llama 4 MaverickDownload and partners; no Meta-hosted API17B / 400B1M listed in Meta’s launch materialsStronger open-weight generalist
3 / Meta AI / Muse

Meta wants intelligence to appear where social and personal context already lives

Muse Spark is designed for Meta’s distribution advantage: billions of users, messaging, public content, cameras, shopping, social graphs and wearable devices. Since July 24 the assistant also runs multi-step tasks rather than only answering, and Meta has added teen-safety controls around those conversations.

Agentic

Briefings and recurring tasks

Announced July 24 and powered by Muse Spark 1.1: daily briefings, recurring tasks, web research and slide generation. Rolling out that day in select markets, with WhatsApp in the coming weeks.

Safety

Parental alerts for teen distress

From July 16, Meta AI can alert parents when a teen’s conversation shows self-harm signals. Live in the US, UK, Australia and Canada, with global availability stated for year end.

Conversation

Meta AI app and web

Rich conversation, personal context, web information and direct creation in one consumer surface.

Messaging

WhatsApp and Messenger

Assistant access inside conversations where plans, questions and shared decisions already appear.

Social

Instagram, Facebook and Threads

Context from public posts, creators, groups and social content can shape search and recommendations.

Wearables

AI glasses

Voice, camera and live visual assistance bring Meta AI into the physical moment.

Search

AI Mode in Facebook

Answers can be grounded in public perspectives from groups, Reels and other Meta content.

Shopping

Marketplace plus web

Meta AI can combine local Marketplace listings with wider web products and map context.

Teams

Side chats and mentions

Meta is testing ways to bring private AI assistance into group discussions without exposing the side conversation.

Creation

Muse Image

Generate, edit and share visuals directly into stories, chats and feeds.

A phone-first work moment

Meta’s work relevance appears in small, contextual actions rather than formal enterprise documents.

Capture
Use voice or camera in context

Ask about a product, place, message, photo or situation at the moment it appears.

  • Less context reconstruction
  • Immediate multimodal input
  • Personal-data sensitivity
4 / Muse Image

Image generation grounded in the social and personal world

Muse Image’s distinctive proposition is not only output quality. It can reason over the brief, combine several references and use web context inside the Meta AI experience. One launch capability has already been withdrawn, so check the current feature set before designing a workflow around it.

Prompt reasoning

Muse Spark helps plan layout, retrieve relevant context and coordinate visual references before image creation.

  • Multi-step generation
  • Useful for informational visuals
  • Review inferred facts

Multi-photo blending

Combine selfies, vacation images, pets, products and other references into a single creation.

  • Check identity consistency
  • Respect consent
  • Avoid confidential uploads

Text and QR generation

Meta says Muse Image can render clean text and functional QR codes for practical visual outputs.

  • Test every QR code
  • Proofread all text
  • Do not trust generated regulatory copy

Presets

Suggested prompts make common transformations and styles accessible without writing a detailed brief.

  • Useful for adoption
  • Can create generic sameness
  • Apply brand judgment

Sketch-based edits

Circle, sketch or annotate directly on the image to describe changes while keeping conversational context.

  • Precise local direction
  • Iterative editing
  • Inspect unintended changes elsewhere

Withdrawn: profile mentions

The launch feature that let users pull public Instagram account content into a generated image was removed on July 15. Meta’s update reads: “We’ve heard the feedback that this feature missed the mark, so it’s no longer available.”

  • Eight days from ship to withdrawal
  • Do not document it as available
  • Re-check any workflow built on it

Room redesign shopping

Use a room photo and ask for redesigns grounded in real products from the web or Marketplace.

  • Visualize before buying
  • Verify dimensions and availability
  • Separate concept from specification

Direct distribution

Share creations into chats, stories and feeds, reducing friction between generation and publication.

  • Faster workflow
  • Higher accidental-publication risk
  • Use an approval pause

Advertiser path

Meta announced Muse Image support for Advantage+ creative, extending the model into paid-media production.

  • Brand and claim review
  • Test across placements
  • Track model-generated asset provenance
Campaign concept
Create three visual territories for this campaign using only the product images I upload. Explain the strategic idea before generating each route. Do not invent product features or customer claims.
Room visualization
Redesign this room for a hybrid-working professional. Keep the exact room geometry, use products that are currently available, and list the dimensions I must verify before purchase.
Social infographic
Create a 4:5 infographic from these verified facts. Keep text concise and legible. After generating, transcribe every word in the image so I can proofread it.
Local edit
Keep composition, people and lighting unchanged. Remove only the marked object and rebuild the background naturally. Do not alter faces, clothing or product details.
5 / Llama 4

Open weights trade convenience for control

Llama 4 remains relevant because organizations can download, host and customize it. Since Meta shut its hosted Llama API on July 6, 2026, that is the only path: the weights stay available for download, and the serving decision belongs to you or a third-party host.

Scout or Maverick?

Choose from the workload and infrastructure.

Extreme context or efficient serving
Llama 4 Scout

Use Scout when the supported 10M context, smaller total parameter count or single-H100 efficiency is the dominant requirement.

  • 17B active / 109B total
  • 16 experts
  • Long documents and large code/data contexts

Long-context analysis

Scout’s supported 10M context can enable codebase, archive and multi-document workloads that are otherwise split across requests.

  • Test retrieval at real lengths
  • Build coverage checks
  • Do not assume equal attention

Multimodal understanding

Scout and Maverick accept text and image inputs for grounding, visual question answering and application-specific workflows.

  • Use exact visual tasks
  • Evaluate localization and detail
  • Add application guardrails

Fine-tuning and customization

Open weights allow organizations to adapt behavior, domain language and task performance beyond prompting.

  • Curate training data
  • Prevent regression
  • Maintain evaluation baselines

Llama Stack

Meta provides standardized building blocks for inference, evaluation, tools and application development.

  • Reduce integration fragmentation
  • Still choose serving architecture
  • Review current project status

Llama Guard 4

A companion safeguard model supports content and risk classification around Llama applications.

  • One safety layer, not the whole system
  • Tune policy to the use case
  • Monitor false positives and misses

Cloud and edge partners

Llama is available through model hosts, clouds and edge ecosystems, creating several operational paths. With Meta’s own API retired, one of these is now required rather than optional.

  • Avoid provider-specific lock-in
  • Check quantization and context differences
  • Document the deployed artifact

Open models beyond Llama

Meta’s open-model story is wider than Llama. On July 21 Meta described SAM 3 and DINOv3, both open source, deployed in SYNAPS-I, a Lawrence Berkeley–led Department of Energy multi-lab project under the Genesis Mission.

  • Vision models, not chat models
  • Fine-tuned across 300 A100 GPUs
  • Evidence of open models in serious scientific work
SAM 3
Open segmentation model deployed in SYNAPS-I
DINOv3
Open vision model deployed alongside it
300
A100 GPUs used to fine-tune the models, as reported by Meta
15 min
Reported turnaround for labeled 3D volumes, against roughly a month of manual annotation
DimensionScoutMaverick
Active parameters17B17B
Total parameters109B400B
Experts16128
Supported context10M1M in launch materials
Deployment positioningSingle H100 GPU with Int4 quantizationSingle H100 host, meaning one multi-GPU server rather than one GPU
Best fitEfficient extreme-context workStronger general open-weight work
6 / Deployment

The open-weight advantage appears only if the organization can operate it

A downloadable model can improve control, customization and portability. It can also create a permanent model-serving, security and evaluation responsibility.

A responsible Llama deployment

Treat the model as a software supply chain, not a one-time download.

Define
State why open weights matter

Name the control, cost, latency, data-location or customization requirement that a managed API cannot meet.

  • Avoid ideology-only decisions
  • Quantify success
  • Define acceptable maintenance burden

Hosting decision

Choose the control point the organization genuinely needs.

No infrastructure team
Use a managed model endpoint

A cloud or inference partner can provide Llama without transferring the full serving burden to the organization.

  • Check model and quantization
  • Review data and region
  • Retain migration options
7 / Behavior Playbook

Eight habits for keeping Meta’s two tracks clear

The most important behavior is naming the product, model and access path precisely. “Use Meta AI” and “deploy Llama” are different decisions.

01

Name Muse or Llama

Do not say “Meta’s model” when the tracks have different access and purposes.

02

Use Muse for distribution

Choose Meta AI when social, messaging, camera, voice or creator context is the advantage.

03

Use Llama for control

Choose open weights when hosting, customization or data location is the requirement, and budget for the hosting Meta no longer provides.

04

Say public preview

Meta described the Meta Model API as a public preview on July 9. No GA announcement has been published since.

05

Test long-context coverage

A 10M window needs retrieval and reasoning evaluation at the real workload size.

06

Treat social grounding as biased

Public Meta content reveals perspectives, not a representative population.

07

Review before direct sharing

Generation-to-feed convenience needs a deliberate publication pause.

08

Operate the open model

Pin, secure, evaluate and maintain every deployed Llama artifact.

8 / Watch Outs

Reach, personal context and open weights create different risk profiles

Meta’s consumer products raise privacy, representation and publication questions. Llama shifts safety, security and lifecycle responsibility toward the deployer.

Muse availability

The assistant is widely distributed, but Meta’s last dated statement on the Meta Model API, from July 9, calls it a public preview. Confirm the current label before committing.

Staged rollout

The July 24 agentic capabilities started in select markets, with WhatsApp promised in the coming weeks. Meta product and model features reach countries and surfaces at different times.

Features can be withdrawn

Muse Image shipped with public-profile mentions on July 7 and lost them on July 15. A shipped consumer AI feature can disappear in eight days, so do not build a dependency on one.

Public-content bias

Groups, posts, Reels and profiles provide useful context without representing all users or reality.

Minors and duty of care

Parental alerts for teen distress reached the US, UK, Australia and Canada in July, with global coverage stated for year end. Coverage is uneven until then.

No Meta-hosted Llama

The hosted Llama API shut down on July 6, 2026. Any architecture diagram that shows Meta serving Llama traffic is out of date.

Direct sharing

The shortest path from generation to publication increases accidental or insufficiently reviewed output.

Llama license, unverified this week

Open weight is not the same as unrestricted open source. The Llama 4 license terms, including the 700 million monthly-active-user threshold, were last confirmed on July 27, 2026 and could not be re-checked since: Meta serves these developer pages as client-rendered bodies that an automated check cannot read. Anyone relying on the threshold for a commercial decision should open the license and read it directly rather than trust this guide.

Model operations

Self-hosting transfers serving, logging, patching, abuse controls and evaluation to the organization.

Roadmap speculation

Do not publish codenames, projected dates or unreleased models as product facts. Muse Video is announced but not shipped, and no Llama 4.1, 4.5 or 5 exists on any first-party source.

9 / Sources

First-party evidence behind this guide

These official Meta pages anchor the current Muse and Llama tracks, the July 2026 product changes and the retirement of the hosted Llama API. Re-checked on August 3, 2026: no material first-party Meta change was found between July 28 and August 3, so every dated claim in this guide still rests on the July sources below and nothing was added for the sake of movement.

AI Mindset

Explore the model cheatsheets