BFL Announces 'FLUX 3', Unified Generation of Images, Videos, and Audio
Black Forest Labs (BFL), a German AI company, has announced 'FLUX 3', a multimodal AI model that handles images, videos (up to 20 seconds with audio), and robot motion control in a single model. Currently, certain features are available only through early access, with pricing and open-weight version release planned for later.

Black Forest Labs (BFL), a German AI company, has announced a multimodal model called 'FLUX 3'. The model can generate either images or videos with audio (up to 20 seconds) from a single prompt. This marks the company's first public release of a video generation model.
BFL has been known so far for its 'FLUX' series of image generation models. FLUX 3 is characterized by the fact that instead of processing images, videos, and audio with separate models, it is trained comprehensively on a single architecture (design foundation). The company calls this approach 'Visual Intelligence' and envisions applying a single capability across creative production, enterprise simulation, computer control, and even robot motion control.
FLUX 3 is offered in four product lines: 'FLUX 3 Video', 'FLUX 3 Image', 'FLUX 3 Action', and the open-source version 'FLUX 3 Dev' to be released later. At present, FLUX 3 Video and FLUX 3 Action with audio generation options are included in the 'Early Access' program, which anyone can apply for, but approval from BFL is required to use. General release through APIs or external partners has not yet occurred. BFL explains that FLUX 3 Image will be rolled out within weeks, followed by general availability in sequence.
On the other hand, several critical pieces of information remain undisclosed at this stage. Pricing, Service Level Agreements (SLA), evaluation methodology details, benchmark sample sizes and evaluator counts, and comparative metrics for the image model—information necessary for enterprises to objectively assess adoption costs and performance—remain unpublished. Additionally, the model's weight data (files for running in local environments) has not been released at this time. The open-weight version is scheduled for release in the latter half of this year, and BFL states that FLUX 3 Dev will be FLUX's first open-source multimodal version. However, while previous FLUX Dev series focused exclusively on images, the scope this time extends to videos, audio, and action prediction, representing a broader commitment.
This phased release approach is similar to the release strategies recently adopted by major U.S. AI labs such as Anthropic and OpenAI. However, in those cases, safety concerns or government requests were cited as background. It remains unclear what BFL is presenting as the reason for this phased release at this time.
The FLUX series has been gaining support from the developer community through an open-weight strategy of releasing model weights for free. From this perspective, the deferral of the open-source version release in this launch may be seen as disappointing to existing users. BFL's commitment to contributing to the developer community is not being negated, but the impact of delayed open-source release cannot be ignored.
Multimodal integration as a design philosophy is positioned as having advantages in consistency and development efficiency compared to approaches that combine separate models. However, how these advantages are reflected in actual product quality can only be assessed once pricing and benchmark details are published and enterprises or developers can independently verify them. The content and pace of future information disclosure will be critical factors determining FLUX 3's actual adoption rate.
This article is an original work independently written and edited by the AI issue editorial team based on factual reporting. © AI issue. Unauthorized reproduction, redistribution, or use for AI training is prohibited.