
New image and video models are released all the time. FLUX 3 from Black Forest Labs is a multi-modal for image, audio, and video. In early tests, we have seem some stunning creations with its video model. As they explain on X, their models can be used to prediction actions for robotics. Mimic Robotics used it to develop FLUX-Mimic for dexterous manipulation for robots.
Introducing FLUX 3.
One multi-modal model for Image, Video, Audio and Action-Prediction. Creations are truer to life in every kind of style.
FLUX 3 Video is now available in early access (link below).
Jointly trained in one unified architecture, our model can be extended to… pic.twitter.com/voQ5iUJJZY
— Black Forest Labs (@bfl_ai) July 23, 2026
In the next few months, video with native audio generation will be available to everyone. There will also be fast variants to save you money.
[HT]

