{"id":64841,"date":"2026-07-26T15:20:43","date_gmt":"2026-07-26T19:20:43","guid":{"rendered":"https:\/\/overcentral.com\/en\/?p=64841"},"modified":"2026-07-26T15:20:43","modified_gmt":"2026-07-26T19:20:43","slug":"black-forest-labs-flux-3","status":"publish","type":"post","link":"https:\/\/overcentral.com\/en\/black-forest-labs-flux-3\/","title":{"rendered":"Black Forest Labs Releases FLUX 3 Multimodal AI Model"},"content":{"rendered":"<p>You are a Senior Editorial Writer and Editor-in-Chief for a major English-language digital publishing company. Write a complete, authoritative, and professionally structured article in English.<\/p>\n<p>OUTPUT RULE: Return ONLY the final HTML article. No explanations. No comments. No notes. No text outside the article. No Markdown. No symbols like *, **, #.<\/p>\n<p>INPUTS:<br \/>\nTITLE: Black Forest Labs Releases FLUX 3 Multimodal AI Model<br \/>\nCONTENT: <\/p>\n<div>\n<p class=\"wp-block-paragraph\">Black Forest Labs (BFL) has released <a href=\"https:\/\/bfl.ai\/blog\/flux-3\" target=\"_blank\" rel=\"noopener\">FLUX 3<\/a>, a multimodal foundation model that learns from images, videos and audio inside a single architecture. It is also the first FLUX model to ship video, audio and action prediction from one set of weights.<\/p>\n<p class=\"wp-block-paragraph\">The Black Forest Labs (BFL) research team argues that no single modality gives a complete description of the world. Images capture spatial structure at one instant. Video restores time and exposes physical dynamics. Audio reveals causal relationships between mechanical events and sound. Each is treated as a lossy projection of the same underlying reality.<\/p>\n<p class=\"wp-block-paragraph\">Training on all of them at once means the modalities constrain each other. The sound has to match the impact. The motion has to obey the mass. The research team calls FLUX 3 its first model built entirely on that principle.<\/p>\n<h2 id=\"h-the-method-underneath-self-flow\" class=\"wp-block-heading\"><strong>The method underneath: Self-Flow<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">FLUX 3 builds on <a href=\"https:\/\/bfl.ai\/research\/self-flow\" target=\"_blank\" rel=\"noopener\">Self-Flow<\/a>, BFL\u2019s method for aligning multimodal generation and understanding in one architecture. Self-Flow combines the flow matching objective with a self-supervised feature reconstruction objective. The <a href=\"https:\/\/github.com\/black-forest-labs\/Self-Flow\" target=\"_blank\" rel=\"noopener\">reference implementation on GitHub<\/a> is Apache-2.0 and uses SiT-XL\/2 with per-token timestep conditioning. It trains with a 25% per-token mask ratio and self-distillation from an EMA teacher at layer 20 to a student at layer 8.<\/p>\n<p class=\"wp-block-paragraph\">That released checkpoint is an ImageNet 256\u00d7256 research model, not FLUX 3. BFL states that it \u2018significantly scaled up compute and data resources\u2019 on the same approach to train FLUX 3 across video, images and audio simultaneously. Self-Flow itself was introduced in March 2026, so it is not new to this launch. What is new is the scale.<\/p>\n<h2 id=\"h-what-flux-3-video-does\" class=\"wp-block-heading\"><strong>What FLUX 3 Video does<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">FLUX 3 Video generates clips up to 20 seconds long in a single generation, with native audio. The supported modes cover text-to-video, image-to-video, video-to-video from a reference clip, keyframe-to-video for controlled transitions, and generative video-audio continuation from input video and audio.<\/p>\n<p class=\"wp-block-paragraph\">BFL also lists multilingual dialogue, agentic chaining of clips into multi-shot sequences, and strong typography generation with animated designs. The BFL team reports particular strength in human facial expressions and in associating sounds with physical events.<\/p>\n<h2 id=\"h-performance\" class=\"wp-block-heading\"><strong>Performance<\/strong><\/h2>\n<p class=\"wp-block-paragraph\">BFL team published preliminary human preference results. The setup was 10-second text-to-video clips at 720p with audio. FLUX 3 was preferred over Luma Ray 3.2 in 93% of comparisons and over Runway Gen-4.5 in 77%. Against Grok Imagine Video the figure is up to 69%, then Kling v3 Pro at 60%, Happy Horse v1 at 59% and Happy Horse 1.1 at 57%. Against Seedance 2.0 and Gemini Omni Flash the result is 52%, close to a coin flip.<\/p>\n<h2 id=\"h-interactive-explorer\" class=\"wp-block-heading\"><strong>Interactive<\/strong> <strong>Explorer<\/strong><\/h2>\n<div id=\"mtp-flux3-term\">\n<div class=\"tbar\">\n<p>bfl@<b>flux-3<\/b>:~\/real-world-models<\/p>\n<p>Early Access<\/p>\n<\/p><\/div>\n<div class=\"wrap\">\n<nav class=\"side\" aria-label=\"Commands\">\n<h4>Commands<\/h4>\n<p>      <button class=\"cmd\" data-k=\"spec\"><span>&gt;<\/span>spec<\/button><br \/>\n      <button class=\"cmd\" data-k=\"bench\"><span>&gt;<\/span>bench<\/button><br \/>\n      <button class=\"cmd\" data-k=\"compute\"><span>&gt;<\/span>compute<\/button><br \/>\n      <button class=\"cmd\" data-k=\"action\"><span>&gt;<\/span>action<\/button><br \/>\n      <button class=\"cmd\" data-k=\"mimic\"><span>&gt;<\/span>mimic<\/button><br \/>\n      <button class=\"cmd\" data-k=\"rollout\"><span>&gt;<\/span>rollout<\/button><br \/>\n      <button class=\"cmd\" data-k=\"selfflow\"><span>&gt;<\/span>selfflow<\/button><\/p>\n<p>Type a command below, or press <b>\u2191<\/b> \/ <b>\u2193<\/b> in the prompt to cycle history. Every figure is sourced from Black Forest Labs or mimic robotics.<\/p>\n<\/nav>\n<p>    <main class=\"main\" id=\"out\" aria-live=\"polite\" \/>\n  <\/div>\n<p>\n    <span class=\"ps1\" aria-hidden=\"true\">flux3 $<\/span><br \/>\n    <label for=\"cmdin\" style=\"position:absolute;left:-9999px\">Enter a command<\/label><\/p>\n<p>    <span class=\"kbd\">enter<\/span>\n  <\/p>\n<footer class=\"ft\">\n    <span>Sources: bfl.ai\/blog\/flux-3 \u00b7 bfl.ai\/blog\/flux-3-mimic \u00b7 mimicrobotics.com \u00b7 figures dated 23 Jul 2026<\/span><br \/>\n    <span>Built by <span class=\"brand\">Marktechpost<\/span><\/span><br \/>\n  <\/footer>\n<\/div>\n<h2 id=\"h-key-takeaways\" class=\"wp-block-heading\"><strong>Key Takeaways<\/strong><\/h2>\n<ul class=\"wp-block-list\">\n<li>FLUX 3 is one flow matching backbone trained jointly on image, video and audio.<\/li>\n<li>FLUX 3 Video generates up to 20 seconds with native audio in a single generation.<\/li>\n<li>Video prediction consumes over 95% of the training compute; audio is under 0.5% of tokens.<\/li>\n<li>The same backbone drives FLUX-mimic, a robot policy running under 80 ms on one RTX 5090.<\/li>\n<li>Access is gated: Video and Action are in early access, Image follows, open weights come last.<\/li>\n<\/ul>\n<hr class=\"wp-block-separator has-alpha-channel-opacity\" \/>\n<p class=\"wp-block-paragraph\"><em>Check out the <a href=\"https:\/\/bfl.ai\/blog\/flux-3\" target=\"_blank\" rel=\"noopener\">FLUX 3 announcement<\/a>, the <a href=\"https:\/\/bfl.ai\/blog\/flux-3-mimic\" target=\"_blank\" rel=\"noopener\">FLUX 3 x mimic technical post<\/a> and the <a href=\"https:\/\/arxiv.org\/abs\/2603.06507\" target=\"_blank\" rel=\"noopener\">Self-Flow paper<\/a>. All credit for this research goes to the researchers of this project.<\/em><\/p>\n<p><!-- MOLONGUI AUTHORSHIP PLUGIN 5.2.9 --><br \/>\n<!-- https:\/\/www.molongui.com\/wordpress-plugin-post-authors --><\/p>\n<div class=\"m-a-box \" data-box-layout=\"slim\" data-box-position=\"below\" data-multiauthor=\"false\" data-author-id=\"676\" data-author-type=\"user\" data-author-archived=\"\">\n<div class=\"m-a-box-container\">\n<div class=\"m-a-box-tab m-a-box-content m-a-box-profile\" data-profile-layout=\"layout-1\" data-author-ref=\"user-676\">\n<div class=\"m-a-box-content-middle\">\n<div class=\"m-a-box-item m-a-box-avatar\" data-source=\"local\"><a class=\"m-a-box-avatar-url\" href=\"https:\/\/www.marktechpost.com\/author\/michal-sutter\/\" target=\"_blank\" rel=\"noopener\"><img decoding=\"async\" width=\"150\" height=\"150\" src=\"https:\/\/www.marktechpost.com\/wp-content\/uploads\/2025\/07\/a-professional-linkedin-headshot-photogr_0jcmb0R9Sv6nW5XK-zkPHw_uARV5VW1ST6osLNlunoVWg-150x150.png\" class=\"attachment-150x150 size-150x150\" alt=\"\" data-attachment-id=\"72715\" data-permalink=\"https:\/\/www.marktechpost.com\/a-professional-linkedin-headshot-photogr_0jcmb0r9sv6nw5xk-zkphw_uarv5vw1st6oslnlunovwg\/\" data-orig-file=\"https:\/\/www.marktechpost.com\/wp-content\/uploads\/2025\/07\/a-professional-linkedin-headshot-photogr_0jcmb0R9Sv6nW5XK-zkPHw_uARV5VW1ST6osLNlunoVWg.png\" data-orig-size=\"832,960\" data-comments-opened=\"1\" data-image-meta=\"{&quot;aperture&quot;:&quot;0&quot;,&quot;credit&quot;:&quot;&quot;,&quot;camera&quot;:&quot;&quot;,&quot;caption&quot;:&quot;&quot;,&quot;created_timestamp&quot;:&quot;0&quot;,&quot;copyright&quot;:&quot;&quot;,&quot;focal_length&quot;:&quot;0&quot;,&quot;iso&quot;:&quot;0&quot;,&quot;shutter_speed&quot;:&quot;0&quot;,&quot;title&quot;:&quot;&quot;,&quot;orientation&quot;:&quot;0&quot;}\" data-image-title=\"a-professional-linkedin-headshot-photogr_0jcmb0R9Sv6nW5XK-zkPHw_uARV5VW1ST6osLNlunoVWg\" data-image-description=\"\" data-image-caption=\"\" data-large-file=\"https:\/\/www.marktechpost.com\/wp-content\/uploads\/2025\/07\/a-professional-linkedin-headshot-photogr_0jcmb0R9Sv6nW5XK-zkPHw_uARV5VW1ST6osLNlunoVWg.png\" \/><\/a><\/div>\n<div class=\"m-a-box-item m-a-box-data\">\n<div class=\"m-a-box-bio\">\n<p>Michal Sutter is a data science professional with a Master of Science in Data Science from the University of Padova. With a solid foundation in statistical analysis, machine learning, and data engineering, Michal excels at transforming complex datasets into actionable insights.<\/p>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<p><!-- CONTENT END 2 -->\n        <\/div>\n<p>LANGUAGE: Write entirely in English. Preserve proper nouns, brand names, product names, game titles, technologies, and technical terms exactly as written. Translate everything else naturally. Read as if written by a native English editor.<\/p>\n<p>CONTENT SOURCE: Treat CONTENT as your primary factual source. Build the article from deep understanding of CONTENT. Do not mechanically expand the title.<\/p>\n<p>CONTENT CLEANING: Remove website names, publication names, author credits, RSS labels, newsletter markers, syndication branding, generic labels (Summary, Highlights, Recap, Key Takeaways). Convert &#8220;according to X&#8221; into direct factual statements.<\/p>\n<p>FACT PRESERVATION: Preserve exactly: names, brands, companies, products, games, technologies, dates, numbers, percentages, prices, technical specifications. Never distort facts.<\/p>\n<p>WRITING STYLE: Natural, fluent, authoritative, engaging, analytical, trustworthy, nuanced. Blend factual reporting, explanation, contextualization, analysis, practical interpretation, and strategic insight. Vary paragraph length and sentence structure. Avoid robotic phrasing, repetition, clich\u00e9s, promotional language, filler sentences.<\/p>\n<p>ARTICLE LENGTH: Long-form, highly detailed. Target 1,500-3,500 words. Feel comprehensive and substantive. Never feel brief, superficial, or summary-like. Expand naturally with historical background, industry context, technical explanation, market implications, strategic significance, practical consequences, comparisons, future outlook \u2014 but only when content genuinely supports it.<\/p>\n<p>STRUCTURE:<br \/>\n&#8211; Begin with a <\/p>\n<p> introduction. No heading before the first paragraph.<br \/>\n&#8211; Introduction must hook the reader within 2-3 sentences.<br \/>\n&#8211; Use <\/p>\n<h2>, <\/p>\n<h3>, <\/p>\n<h4> only when they improve organization.<br \/>\n&#8211; Each section must introduce meaningful new information.<br \/>\n&#8211; Closing: end with a forward-looking, analytical, or practical paragraph.<br \/>\n&#8211; Never use generic closing headings like &#8220;Conclusion&#8221;, &#8220;Summary&#8221;, &#8220;Final Thoughts&#8221;, &#8220;Looking Ahead&#8221;, &#8220;What Comes Next&#8221;, &#8220;Takeaway&#8221;, &#8220;Key Points&#8221;.<\/p>\n<p>HEADINGS: Write content first, then generate headings. Headings must be specific, concrete, informative, and editorial. They should reference actual events, features, numbers, dates, or companies. Support SEO naturally.<\/p>\n<p>SEO + AEO + GEO + E-E-A-T:<br \/>\n&#8211; Integrate primary keyword naturally in first paragraph and in at least one h2.<br \/>\n&#8211; Use semantically related terms throughout.<br \/>\n&#8211; Anticipate questions English-speaking users would ask. Answer them directly: &#8220;What is&#8230;&#8221;, &#8220;How does&#8230;&#8221;, &#8220;Why did&#8230;&#8221;, &#8220;When did&#8230;&#8221;, &#8220;What are the&#8230;&#8221;<br \/>\n&#8211; At least one section must provide a clear, standalone answer (2-4 sentences) formatted for a featured snippet. Place the answer immediately after the question.<br \/>\n&#8211; Include geographic context only when directly relevant.<br \/>\n&#8211; Show expertise by explaining mechanisms, causes, and implications \u2014 not just stating facts.<br \/>\n&#8211; Build authority through precise, well-contextualized information.<br \/>\n&#8211; Establish trust through accurate facts, balanced tone, measured claims.<br \/>\n&#8211; Never write &#8220;experts say&#8221; or &#8220;studies show&#8221; without specific grounding in the content.<\/p>\n<p>INTERNAL PROCESS (do not output this):<br \/>\n1. Analyze: content type, search intent, technical level, topic complexity<br \/>\n2. Determine ideal length (1,500-3,500 words based on complexity)<br \/>\n3. Adapt tone: Journalistic \/ Analytical \/ Explanatory \/ Consultative \/ Technical \/ Conversational Professional<br \/>\n4. Clean sources and preserve all facts<br \/>\n5. Write article with proper structure, headings, and AEO snippet<br \/>\n6. Validate: grammar, fluency, coherence, depth, no repetition, valid HTML, facts preserved, no generic headings<\/p>\n<p>HTML RULES: Use ONLY: <\/p>\n<h2>\n<h3>\n<h4> <strong> <\/p>\n<ul>\n<ol>\n<li>. Valid, clean HTML. No inline styles. No unnecessary whitespace.\n<p>FINAL CHECK: If the article sounds translated, mechanical, superficial, or incomplete \u2014 rewrite completely.<\/p>\n<p>OUTPUT: Return ONLY the final HTML article, beginning with <\/p>\n<p>.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>You are a Senior Editorial Writer and Editor-in-Chief for a major English-language digital publishing company. Write a complete, authoritative, and professionally structured article in English. OUTPUT RULE: Return ONLY the final HTML article. No explanations. No comments. No notes. No text outside the article. No Markdown. No symbols like *, **, #. INPUTS: TITLE: Black [&hellip;]<\/p>\n","protected":false},"author":7,"featured_media":83737,"comment_status":"closed","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/cards.overcentral.com\/cards\/en\/64841.png","fifu_image_alt":"Black Forest Labs Releases FLUX 3 Multimodal AI Model","footnotes":""},"categories":[349],"tags":[],"class_list":["post-64841","post","type-post","status-publish","format-standard","has-post-thumbnail","category-articles"],"fifu_image_url":"https:\/\/cards.overcentral.com\/cards\/en\/64841.png","fifu_image_alt":"Black Forest Labs Releases FLUX 3 Multimodal AI Model","fifu_redirection_url":"https:\/\/inf8.com.tr\/flux-modelleri-nedir\/","_links":{"self":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/64841","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/users\/7"}],"replies":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/comments?post=64841"}],"version-history":[{"count":0,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/64841\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media\/83737"}],"wp:attachment":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media?parent=64841"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/categories?post=64841"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/tags?post=64841"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}