Geo News
Community curated by people like you
LatestAICryptoHealthWorld AffairsUS Politics
Black Forest Labs releases FLUX 3 multimodal foundation model
00

Black Forest Labs releases FLUX 3 multimodal foundation model

Jul 26, 2026

Black Forest Labs has released FLUX 3, a multimodal foundation model that learns from images, video, and audio in a single architecture. The model can generate 20-second video clips with native audio and, in human preference tests, was preferred over Luma Ray 3.2 in 93% of comparisons. The same backbone also powers a robot policy, FLUX-mimic. Access is being rolled out in stages via early access.

FLUX 3 multimodal architecture

  • ▪The model's design is based on the principle that different modalities are projections of the same reality and should constrain each other during training
  • ▪Black Forest Labs (BFL) has released FLUX 3, a multimodal foundation model that learns from images, videos, and audio
  • ▪FLUX 3 uses a single architecture and one set of weights for video, audio, and action prediction

Self-Flow training method

  • ▪FLUX 3 is built on Self-Flow, a method from Black Forest Labs that combines a flow matching objective with a self-supervised feature reconstruction objective
  • ▪The public Self-Flow reference implementation on GitHub is an ImageNet research model and is not the full FLUX 3 model
  • ▪The Self-Flow method was first introduced in March 2026, with FLUX 3 representing a significant scale-up of compute and data
  • ▪Video prediction accounts for over 95% of the training compute for FLUX 3, while audio uses less than 0.5% of the tokens

FLUX 3 Video capabilities

  • ▪FLUX 3 Video generates clips up to 20 seconds long with native audio in a single pass
  • ▪Additional capabilities include generative video-audio continuation, multilingual dialogue, and chaining clips into multi-shot sequences
  • ▪Black Forest Labs reports the model has particular strength in generating human facial expressions and matching sounds to physical events
  • ▪The model supports text-to-video, image-to-video, video-to-video, and keyframe-to-video generation

Performance benchmark results

  • ▪FLUX 3 was also preferred over Runway Gen-4.5 (77%), Grok Imagine Video (69%), and Kling v3 Pro (60%)
  • ▪In human preference tests, FLUX 3's 10-second, 720p text-to-video clips were preferred over Luma Ray 3.2 in 93% of comparisons
  • ▪Against Seedance 2.0 and Gemini Omni Flash, FLUX 3 was preferred in 52% of comparisons, indicating similar performance

Early access rollout

  • ▪Access to FLUX 3 is being released in stages, with video and action prediction capabilities available first in early access
  • ▪Image generation will be released after the early access period, and open weights for the model will be released last

Robot policy integration

  • ▪The FLUX-mimic robot policy runs in under 80 milliseconds on a single NVIDIA RTX 5090 GPU
  • ▪The same model architecture used for FLUX 3 also powers FLUX-mimic, a robot policy

1 source

Marktechpost
Black Forest Labs Releases FLUX 3: A Multimodal Flow Model for Image, Video, Audio and Robot Action Prediction
View source article

Featured stories

View more in Generative audio

Google announces Gemini 4 Argon in limited release to cybersecurity partners

Sep 30, 2026 · 12 sources

OpenAI and Synopsys partner to develop AI model for chip design

Sep 30, 2026 · 3 sources

Anthropic targets November IPO at $25-26 billion valuation after delays

Sep 30, 2026 · 2 sources

DeepSeek releases software tools for Huawei AI chips to challenge Nvidia

Sep 30, 2026 · 4 sources

Story comments

Loading comments…

Related Projects

Black Forest LabsFLUX

Topics

Generative audioAI foundation modelsAI images & videosMultimodal modelsAI tools & products

Featured stories

View more in Generative audio

Google announces Gemini 4 Argon in limited release to cybersecurity partners

Sep 30, 2026 · 12 sources

OpenAI and Synopsys partner to develop AI model for chip design

Sep 30, 2026 · 3 sources

Anthropic targets November IPO at $25-26 billion valuation after delays

Sep 30, 2026 · 2 sources

DeepSeek releases software tools for Huawei AI chips to challenge Nvidia

Sep 30, 2026 · 4 sources