AI TechnologyBlackforestlabsJul 23, 2026 19:24 UTC

Black Forest Labs Releases "Flux 3" with Voice Generation Support

Black Forest Labs has released "Flux 3," a multimodal AI model that integrates images, videos, and audio. The capability to simultaneously generate videos up to 20 seconds long and native audio is a first for the company. According to BFL's own testing, it slightly outperforms Suno 2.0, though third-party verification has not yet been conducted.

Black Forest Labs Releases "Flux 3" with Voice Generation Support

Black Forest Labs (BFL) has released "Flux 3," a multimodal foundation model that learns by combining images, videos, and audio. The model not only can generate videos up to 20 seconds long but also can natively generate audio synchronized with the video (without combining separate tools). This marks BFL's first venture into voice generation capabilities and represents a significant technological milestone for the company.

BFL is a company that has been developing the "Flux" series, which garnered high praise as an image generation AI. While previous Flux models excelled at static image generation, the introduction of Flux 3—which expands capabilities to video and audio—signals that BFL has entered a new phase as a multimodal AI development company beyond the realm of image generation alone. The technology of generating video and audio as a unified whole is an area where applications in content creation and entertainment are expanding, with ongoing competitive development across the industry.

In benchmark testing conducted by BFL itself, Flux 3's performance slightly exceeded that of Suno 2.0, a leading model in the video generation space. However, it should be noted that independent third-party validation results have not yet been published, and this evaluation is based solely on BFL's own testing. Additionally, BFL has disclosed that it is already conducting tests applying Flux 3 to robotics tasks.

BFL's ultimate goal is the construction of a "world model," according to statements from the company. A world model refers to enabling AI to internally understand and reproduce the physical laws and causal relationships of the real world, with applications envisioned in autonomous robotics and advanced simulation scenarios. The experimental application of Flux 3 to robotics can be positioned as an initial step toward that long-term objective.

The trend of combining audio with video generation is accelerating across the AI industry. Rather than generating video and audio separately and then combining them, the ability of a single model to produce both integratively is significant from the perspectives of output consistency and workflow efficiency. BFL's approach represents a shift from combining single-function models to a move toward integrated generative foundations.

The key question going forward is how Flux 3's performance—currently evaluated only through BFL's own testing—will be assessed through independent third-party verification. Additionally, the extent to which robotics applications reach practical deployment levels will be an important indicator for measuring the feasibility of realizing world models. The cross-domain movement spanning video, audio, and robotics is viewed as an opportunity to further clarify BFL's future direction.

#GenerativeAI#VideoGeneration#MultimodalAI#BlackForestLabs#Flux#WorldModel#Robotics
AI issue Staff

This article is an original work independently written and edited by the AI issue editorial team based on factual reporting. © AI issue. Unauthorized reproduction, redistribution, or use for AI training is prohibited.

Comments

Log in to comment