Rendered at 19:55:19 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
vunderba 1 days ago [-]
One of the things they seem to be emphasizing here is the UX around being able to place specific elements where you want them in an image. If the positions of the components in the overall composition are very important, this seems to make that a lot easier and kind of reminds me of InvokeAI.
Ideogram V4, an open-weight model released back in June can also do this [1], but you have to use a relatively cumbersome JSON structure to describe all the different bounding boxes. So it’s definitely a bit of a hassle.
I'll probably be waiting until it goes open-weight (hopefully soon) like they did with Flux.2 / Klein.
You just get an LLM to do the bounding box stuff or use the ComfyUI node that provides a GUI for bounding box generation
CuriouslyC 1 days ago [-]
I always hated Comfy's node based UI, but agents make it tolerable. Now I just have them set up a workflow and I go in and tweak it manually if the results aren't where I want them. I even have agents cherry doing multiple runs and cherry picking the best outputs, models have gotten good enough that it's a real time saver, assuming you have references they can and a rubric to check against.
rf15 5 hours ago [-]
As someone with general experience with some generative AI tools I started exploring Comfy last week and it was the easiest to pick up by far. This is not meant in a troll way, it just reflects how you do it programmatically the most, which I've seen the most before. The other UIs are always opinionated layouts on top of the actual logic where you constantly have to look up/sift through menus for what you want to do. In Comfy, you just use the searchbar and get the right node.
That being said: defining composition made me immediately think that someone probably made a gui like this with easy to move bounding boxes, and I'm happy they did.
swiftcoder 24 hours ago [-]
> I always hated Comfy's node based UI
It's one of the most uniquely hostile user experiences I've ever had the (dis)pleasure of working with
bavell 18 hours ago [-]
Awhile back I built a project which reads comfy's API and builds a typescript sdk from it. Much nicer to work with and still benefits from node caching and the comfy ecosystem. Perhaps I should clean it up and open source it, though I haven't looked around and there may already be other projects doing this out there.
user43928 8 hours ago [-]
Using AI for the ComfyUI API or workflows worked pretty well for me even early this year.
However, today I don't see a reason to use ComfyUI at all.
For Qwen Image 2.1, I had Opus 5.5 create a backend outside of ComfyUI and it was able to make generation take 20% less time with some optimizations.
The optimizations it implemented were caching the text computation in Qwen Image 2.1 rather than including it in every step, fusing projections into a larger matrix multiplication, and decoding the VAE in horizontal bands or something like that.
If there was anything interesting in ComfyUI nodes, I imagine I could just have the AI adopt the relevant code instead of dealing with ComfyUI or custom nodes.
jarjoura 22 hours ago [-]
It's definitely not for me.
From where I'm sitting, it's just turning python functions into boxes and instead of write the function yourself, you drag from the output of one box to the input of another. For 2 or 3 boxes, this is cool, but I opened up a professional workflow and was taken into a view with 100s of boxes and wires all over the place. Uhh, ok?
For myself, I'd rather just create my own python environment, write some quick pytorch or mlx calls, wire up some cli to it and share that in a GitHub.
sorenjan 19 hours ago [-]
I recently found a project[0] that takes a ComfyUI workflow and turns it into a simple UI. I have no affiliation with it and haven't tested it myself, but it looks like it might be handy once you have a finished workflow you want to use.
ComfyUI has a feature called “App” which is basically this. You pick the inputs and outputs in a workflow and then they become a kind of a subgraph.
Disclaimer: never used it for actual generation so I don’t know if it’s doing anything special other than being a subgraph with a different name.
hdjrudni 24 hours ago [-]
How do you get agents to set up a workflow? You just get them to modify the JSON directly and then import it, or do you have a tighter integration (e.g. in the UI)?
CuriouslyC 24 hours ago [-]
The agents can interact with Comfy via API pretty well, which afaik ends up being directly with JSON.
vunderba 1 days ago [-]
Yeah, given how much better Ideogram v4 outputs are when you use the proper structured JSON (background, elements, etc.) I think most users probably have stuck some kind of Qwen/Gemma-based LLM between their raw prompt and the CLIP encoder.
popalchemist 1 days ago [-]
Ideogram 4.5 released 2 days ago with more features along these lines
Adobe has had something like this for a while with Generative Fill. Drag a box, specify contents.
woadwarrior01 12 hours ago [-]
Isn't that just inpainting?
armcat 1 days ago [-]
Does anyone know if it can be used to generate accurate frame-by-frame sprite sequences? I found that no image model can do this well (with sufficient fidelity) - neither with one shot (full spritesheet), nor single frame conditioning. It would be great if an imagegen model could do this. What I do now (I use my own tool https://github.com/acatovic/ai-game-studio) is basically generate a reference image, then condition on that image to generate a very short video, then extract and prune frames. Then I get indie-level sprite fidelity about 90% of the time.
nicerice 21 hours ago [-]
I've seen a rather impressive example of what sounds just like this in Qwen recently
An (apparently, as I haven't tried it) ready-to-go ComfyUI graph for the above. In theory you should be able to drop these into ComfyUI and have it work:
Wow the quality of those is great, reminds me of ghibli anime
Lucasoato 9 hours ago [-]
I’m looking for the same thing, on one side making sprites should be easier because of the lower complexity of pixel art; on the other hand making something with a specific style or with sprite frame-by-frame coherence seems harder.
Imagine online procedural MMO with old gen final fantasy / chrono trigger styles :)
avereveard 13 hours ago [-]
I've stumbled upon a reliablish pipeline you create a reference sheet and a single image pose then trellis v2 for the body and unirig for tigging, then you generate the poses you need in a sheet and send the pose sheet plus the reference images, and use these as 2d anymation cycles and you compose on top of the scene with lanes, this focus all attention of the model to fidelity instead of background integration
popalchemist 1 days ago [-]
The task you're describing is a video model task, not an image model task. It's inherently temporal.
Generate a sprite in an image editor, then use a video model to make the loop you want; then turn the resulting video back into individual sprite images.
armcat 1 days ago [-]
Sure and that's what I do, but a video can be seen as a causal generation on discreet sequence of images, each image conditioned on the one before it. It can also be seen as a series of image editing tasks. It would be cool to get this working in imagegen because of the amount of control you would get. Right now with video generation you can at best specify start and end frame and hope for the best.
popalchemist 1 days ago [-]
Image edit models can probably do a grid, but the temporal accuracy / coherence will never match what a video model, which is really a world model, can do.
armcat 23 hours ago [-]
Regarding your world model statement. This is completely FALSE. Learning the visual statistics of a physical world is NOT the same thing as learning its causal dynamics. The difference is observational likelihood versus intervention-dependent dynamics. There have been great studies disproving video models as world models, like this ICML paper: https://proceedings.mlr.press/v267/kang25g.html. Unfortunately lot of people treat them as world models, mostly because of their ability to reproduce increasingly convincing physical behaviour without ever discovering the underlying physical laws. This is due to many things that I could write an essay about, but better conditioning, latent space represtnation, scaling etc, all make them look awesome.
I can still get absolutely insane results with MiniMax H3 - insane in the sense that it would not make sense at all and would make your head spin.
That paper sets up a task where generating a correct video requires correctly modeling physical laws. From the failure to always generate the correct video, they infer that the model has failed to correctly model the physical laws. The whole premise of the experiment is that learning visual statistics is equivalent to learning causal dynamics, such that failure at one implies failure at the other.
The main difference in applications is that the bar for entertainment is lower, so that even a very bad world model may be acceptable.
popalchemist 23 hours ago [-]
They are proto world models (lots written about this - flux being an example of a video model whose weights also power world-action-engines used in robots) in that they attempt to model causality in time, the thing that is required for what OP is asking for and which image models will never do because it is out of domain.
bobcatsmith 14 hours ago [-]
They are not world models at all, you clearly don’t know what you are talking about and are going around in circles with incorrect statements. Just stop.
The idea to use predictive models of future observations for decision making has a long history, including early demonstrations of robot control which combined action-conditioned video prediction with model-predictive control.
Clearly I'm the uneducated one here eh
numlock86 14 hours ago [-]
The "print on a shirt" example looks so bad that I can't imagine a human looked at this and said "Let's put it on the showcase page!".
The other examples range from mostly good to okay'ish at least.
grobibi 3 hours ago [-]
Not trying to be cynical here, just morbidly curious. How much of their total usage do you think it spent on making porn?
psunavy03 2 hours ago [-]
Supposedly, back in the day, that's one of the reasons VHS became the dominant VCR format over Betamax.
The more things change, the more they stay the same.
arnaudsm 1 days ago [-]
The UX looks amazing and very steerable, congrats to the team for focusing on the interface.
Chats can be awful user interfaces.
vergessenmir 1 days ago [-]
I think we are all waiting for the open weights or local model releases.
pixelesque 1 days ago [-]
Has that been announced?
The website mentions:
> FLUX 3 Image is available under a commercial weights license for companies running image generation at scale. Fine-tune and deploy it on your own infrastructure. Reach out to us to learn more.
I guess the open ones would be non-commercial?
JimDabell 1 days ago [-]
> Open Weights version of FLUX 3 Image is launching in the coming weeks.
There's been an increasing trend of previously-open-weight models going closed-source once they reach a certain size (and size is proportional to capital investment). That the latest release will be open-weight is not necessarily a given just because the previous one was.
vunderba 24 hours ago [-]
It was literally in the announcement from Black Forest Labs:
"Open Weights version of FLUX 3 Image is launching in the coming weeks."
What trend? BFL was always non commercial. The "only" trend would be qwen image not being apache anymore.
spaceman_2020 15 hours ago [-]
Man, AI images are solved
No signal in digital images anymore. If you didn’t see it with your own eyes, it likely never happened
adammarples 10 hours ago [-]
Even in their example "turn the surfboard red", with a bounding box around the surfboard, it turned the whole surfer's wetsuit red too. Things are never quite going to get there.
spiderfarmer 10 hours ago [-]
> Things are never quite going to get there.
Nothing in life is ever perfect. Doesn't mean imperfect stuff can't have a lot of impact.
skybrian 1 days ago [-]
This isn't much of a test, but I bought $10 in credits on their playground and generated a test image. Not bad, but it didn't get the accordion keyboard right. Haven't tried editing yet.
I wish they would work on fixing the "studio lighting" sheen that all these AI-generated humans have
k12sosse 18 hours ago [-]
It works. You have to prompt for it. It helps to have an LLM help build your prompt with you, while you learn what the models need to read to do what you need. Especially a vision one, you can supply images and ask it to give you details on what you want help with.
imgbenchdude 23 hours ago [-]
[dead]
Doohickey-d 24 hours ago [-]
I also don't think there is a park that looks like that, in that location relative ton the Eiffel Tower (although I could be wrong).
cyanmagenta 6 hours ago [-]
The Champ de Mars park has that general vibe and is presumably the primary training reference source here, but it definitely doesn’t look exactly as depicted here.
hmstx 4 hours ago [-]
Let's see. Yeah, the lighting is good
- Location isn't quite right. The side gardens on the Champs de Mars are much wider than this. They're not in this style either :)
- Melty coins, and keyboard.
What else?
- Bunch of visible shape inconsistencies: building window sizes and alignment, fence railing, rivets on the bench, coins, accordion features, perspective violations etc...
- Metal clips on the bottom and top edges of the carrying case are on different sides. Said clips look inconsistent.
- One of the case straps originates inside (from under the velvet, even), the further one originates outside. The carrying handle is oddly mis-centered (as with most things).
- I'm not sure that accordion, compressed, fits in that case. She can just about stand in it (diagonally, on one foot).
- It has weird folds that disappear partway through and turn into adjacent folds going the other way. Between this, the keyboards and the dots, I'm sure there are other things wrong with accordion anatomy.
- Awkward left hand: not sure what's going on with the metacarpal joints, and index finger appears to begin forking into two tips pressing into two (misshapen) buttons. Her right hand is getting there too.
- Weird hair midline. I haven't seen a "sub-midline" like this in the wild :)
- The burgundy coat clips through the bench under her right leg. It's like the front slat goes through a hole in the side of the coat.
- We don't usually have that side curling bar on each side of the sitting area. It looks like it's clipping through the flats next to the leather handbag
- Ill-placed buckle on that bag.
- The green metal fence on the right begins behind her for one vertical bar, next one disappears into the plants, no further vertical bars, no top horizontal railing. Actually, it appears to only have fencing on one side, that's not something I'd see often. On the side where it does have a top, its top rail has very irregular width.
- Sand looks like... breadcrumbs? Very repetitive coarse shapes, like it was poorly done with a clone stamp.
- Something weird about the leaves.
- The foremost lamp isn't sure whether it has flat sides or not.
- The other lamp emerges out of nowhere from behind one of the blob-people in the alley.
- A bunch of unfortunate tangents and occlusions which make you wonder whether the diffusion hallucinated detail out of bigger shapes, or if it's been trained on photos where the tangents were done on purpose.
ie. center bottom bar of the bench being exactly aligned with the gray edge tiles, side curly railing-not-present-on-our-real-benches fully formed on her right, almost completely foreshortened out of existence where the bench's perspective forbids it...
Most people are fooled by much cruder fakes though. This is a lot better but keeps plenty of tells for observant people (and most people aren't)
motoxpro 3 hours ago [-]
Give it 8-9 months and all that stuff will be fixed. Given this is 100x better than what we had 8-9 months ago.
Or just put in more than 10 words of effort, e.g. a better prompt and bounding boxed prompt edits, and all that stuff would be fixed now.
mkl 21 hours ago [-]
The accordion folds and the coins are a bit messed up too.
htx619 1 days ago [-]
[dead]
gAI 1 days ago [-]
Love to see an AI lab outside of US/China releasing good models.
soundworlds 14 hours ago [-]
This is incredible. I tried it for quick UI element replacements in a screenshot, and it nailed 3/4 elements I highlighted first attempt.
As I'm finding with the best GenAIs, this allows for granular iteration, which is where it becomes useful in an industry-wide manner.
mdp2021 8 hours ago [-]
What are the best benchmarks to compare these kind (image, video etc.) of generative models - difficult to compare?
That would be to compare e.g. Qwen Image 3.0 with FLUX 3, with Midjourney etc.
I had seen some attempts - but I do not know well how they try to approach objectivity.
vunderba 4 hours ago [-]
I wouldn’t say that mine, the GenAI Showdown, is necessarily any more or less objective than any others but I definitely do a lot of manual curation.
Prompt results are graded based a weighted calculation which includes: adherence to the prompt, image fidelity, and steerability.
My comparison benchmark also tends to favor prompt adherence, which a lot of others don’t. Most of ones that I've seen tend towards rather simplistic prompts (e.g. "neon-lit city facing a robotic uprising, with high-tech battles, in anime style"), whereas the prompts I've created try to test high specificity.
I’ve been running them all the way back to SDXL.
You can compare specific models using the "View All Models" so if you want to see the progression of open-weight models, or model X vs model Y, you can do so.
Just a heads up - I haven't added Flux 3 as I'm waiting until BFL drops the open-weights version.
This is great. Looking forward to your Flux 3 results.
KazaNLP 1 days ago [-]
Agreed with other comments about the UX. I'm more interested in that than the model itself. Would like to start seeing UI like this where you get to choose the model and compare different models. Can't jump all over the internet to each model developers sandbox just to test their models. Doing it from one place would be nice.
I think that’s less a product of OpenRouter’s generosity and more a result of promotional pricing coming directly from BFL, since other third-party vendors have it as well (Fal.ai, etc.).
Porque no los dos? Presumably you need a model conditioned on the bounding box input to make effective use of the new UI
Zufriedenheit 9 hours ago [-]
When editing an image, they still scramble up small text anybody found a solution to this yet?
pks016 17 hours ago [-]
I tried precise editing with a photo of horse with fence in front of it. One shot didn't work. Tried with multiple smaller regions. Still didn't work
Tried adjusting exposure; Didn't work as well.
Trufa 1 days ago [-]
So much negativity as usual and so little talk about the product, this is pretty impressive, well done, it seems to be filling decently a gap that everyone that has worked enough generating images with AI has faced.
matthew-wegner 7 hours ago [-]
By "this is pretty impressive", do you mean you tried using it? Or do you mean the landing page seems impressive?
doctorpangloss 1 days ago [-]
there basically isn't any authentic use for image generation, i would hardly say the negativity is unfounded
mdp2021 2 hours ago [-]
> authentic use
If you do not consider contemplation, for which we have filled our homes with works of art for centuries,
and if you do not consider those job that involve graphics production e.g. in the Madison Avenue "Mad Men" business (advertising - the payments are authentic),
you could consider "authentic" that when we do video production through generative models some start from stills (image generation) and then have other models animate those stills.
vlyan 13 hours ago [-]
yes, just like emails and chats are inauthentic compared to perfumed letters written in cursive and delivered by a courier on a horse. they are a bit more convenient and affordable for most of us, though.
reilly3000 2 days ago [-]
That sort of steering ability that has been possible with the latest Gemini releases has been nice to work with over previous generations. It’s great to see this improve on the platform with declarative controls built into the API and coming soon as an open model.
> We will open up an early access phase for FLUX 3 Image in the following weeks.
Not sure if there was a separate post for early access or if they just skipped to this.
mromanuk 1 days ago [-]
For a moment I was confused that this was a release of a new stable diffusion. What happened with Stable Diffusion?
yorwba 1 days ago [-]
Some of the original developers left Stability AI to found Black Forest Labs (so in a sense this is a successor to Stable Diffusion) and Stability AI pivoted to audio and milking their existing models.
assimpleaspossi 1 days ago [-]
I shouldn't have to scroll all the way down, click to use the thing, then fumble around to figure out what this does. It should be clear at the top of the first page (so I know right away that I don't need this).
thorum 1 days ago [-]
Not sure what you mean. The top of the linked page is a video showing the product being used. Immediately below are the words “Compose images from scratch” and more examples, followed by “Text-to-image with strong prompt following and a native understanding of composition.” This is all within the first 100 words of content on the page.
bonoboTP 24 hours ago [-]
I agree that it's clear for me and most techies, but maybe for a more general audience they could have added "AI image generation" or some variant of "generative AI" or "image generation model" or something like that. You can also compose images from scratch in Photoshop and Blender.
9dev 7 hours ago [-]
Why should the announcement page for an image generation model and a technical description of its capabilities and interface be accessible for or catering toward a more general audience? What's next, should my API docs explain what an API is?
richardfulop 1 days ago [-]
It is a version 3 of a product.If you actually used or needed these, you would have known right away what this is and how it works.
nazgulsenpai 1 days ago [-]
Unless the version 3 announcement is your onboarding point, like it is for me just now. I can go research FLUX from here, but the original poster's point still holds.
assimpleaspossi 1 days ago [-]
But I don't use it or need it and wasted my time trying to figure out what it did. Clarity should be up front.
richardfulop 21 hours ago [-]
insufferable
Mashimo 1 days ago [-]
> Just write a prompt.
Text-to-image with strong prompt following and a native understanding of composition.
How is this not clear?
amelius 1 days ago [-]
Yes, you can do that with AI now.
Grimblewald 2 days ago [-]
looks cool, eternally greatful these models are marked for open weight releases. Pretty excited
imgbenchdude 1 days ago [-]
Tried Flux 3 on a tiny subjective image benchmark I’m calling One Knee Wonder.
Exact prompt:
Generate a photorealistic image of M81 urban BDU camouflage cargo trousers, shown by themselves. One trouser leg should be posed with the knee lifted 30° from vertical.
Accurate reproduction of the M81 urban camouflage pattern is critical. Match its colors, shapes, scale, distribution, and overall appearance as faithfully as possible.
Flux 3 gets greyscale urban-ish trousers and a lifted knee, but the blotches aren’t real M81 Urban — softer / wrong geometry vs the swatch. Not the worst I’ve seen on this prompt; clearly behind the Gemini 3 Pro Image example above on pattern.
Curious what other models do on the same prompt.
thomasikzelf 11 hours ago [-]
In the context of knee bending: I was creating a simple walking animation with gemini 3 pro and could animate all the poses instead of one, left leg bend right leg straight. Right leg bend it had no problem but this one it just did not seem to understand.
fuzzythrowaway 1 days ago [-]
I like the interface; very useful for some use cases that would otherwise be quite frustrating. Dislike that it's yet another platform held back by arbitrary moderation. You can't make a bicycle for the mind that locks if you try to ride it in the wrong direction.
myself248 1 days ago [-]
Unrelated to the flux images used for floppy disk archiving? Sigh.
Ideogram V4, an open-weight model released back in June can also do this [1], but you have to use a relatively cumbersome JSON structure to describe all the different bounding boxes. So it’s definitely a bit of a hassle.
I'll probably be waiting until it goes open-weight (hopefully soon) like they did with Flux.2 / Klein.
[1] - https://docs.ideogram.ai/using-ideogram/getting-started/prom...
That being said: defining composition made me immediately think that someone probably made a gui like this with easy to move bounding boxes, and I'm happy they did.
It's one of the most uniquely hostile user experiences I've ever had the (dis)pleasure of working with
However, today I don't see a reason to use ComfyUI at all.
For Qwen Image 2.1, I had Opus 5.5 create a backend outside of ComfyUI and it was able to make generation take 20% less time with some optimizations.
The optimizations it implemented were caching the text computation in Qwen Image 2.1 rather than including it in every step, fusing projections into a larger matrix multiplication, and decoding the VAE in horizontal bands or something like that.
If there was anything interesting in ComfyUI nodes, I imagine I could just have the AI adopt the relevant code instead of dealing with ComfyUI or custom nodes.
From where I'm sitting, it's just turning python functions into boxes and instead of write the function yourself, you drag from the output of one box to the input of another. For 2 or 3 boxes, this is cool, but I opened up a professional workflow and was taken into a view with 100s of boxes and wires all over the place. Uhh, ok?
For myself, I'd rather just create my own python environment, write some quick pytorch or mlx calls, wire up some cli to it and share that in a GitHub.
[0] https://github.com/saintbrodie/Orange
Disclaimer: never used it for actual generation so I don’t know if it’s doing anything special other than being a subgraph with a different name.
https://www.youtube.com/watch?v=2mecWZgbaEg
https://media.discordapp.net/attachments/1401891025970008154...
I don't know what went into making it, but their twitter is @araminta_k if you're curious
* https://alvdansen.github.io/animating-on-twos/
* https://github.com/alvdansen/animating-on-twos
* https://huggingface.co/alvdansen/h3-keyframe-animation
## Quick Start
An (apparently, as I haven't tried it) ready-to-go ComfyUI graph for the above. In theory you should be able to drop these into ComfyUI and have it work:
https://huggingface.co/alvdansen/h3-keyframe-animation#quick...
Imagine online procedural MMO with old gen final fantasy / chrono trigger styles :)
Generate a sprite in an image editor, then use a video model to make the loop you want; then turn the resulting video back into individual sprite images.
I can still get absolutely insane results with MiniMax H3 - insane in the sense that it would not make sense at all and would make your head spin.
That paper sets up a task where generating a correct video requires correctly modeling physical laws. From the failure to always generate the correct video, they infer that the model has failed to correctly model the physical laws. The whole premise of the experiment is that learning visual statistics is equivalent to learning causal dynamics, such that failure at one implies failure at the other.
The main difference in applications is that the bar for entertainment is lower, so that even a very bad world model may be acceptable.
Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms
https://bfl.ai/models/flux-3-action
The idea to use predictive models of future observations for decision making has a long history, including early demonstrations of robot control which combined action-conditioned video prediction with model-predictive control.
Clearly I'm the uneducated one here eh
The other examples range from mostly good to okay'ish at least.
The more things change, the more they stay the same.
Chats can be awful user interfaces.
The website mentions:
> FLUX 3 Image is available under a commercial weights license for companies running image generation at scale. Fine-tune and deploy it on your own infrastructure. Reach out to us to learn more.
I guess the open ones would be non-commercial?
— https://x.com/bfl_ai/status/2105734605621825738
https://bfl.ai/legal/non-commercial-license-terms
"Open Weights version of FLUX 3 Image is launching in the coming weeks."
https://nitter.cf/bfl_ai/status/2105734605621825738
No signal in digital images anymore. If you didn’t see it with your own eyes, it likely never happened
Nothing in life is ever perfect. Doesn't mean imperfect stuff can't have a lot of impact.
https://pages.skybrian.com/flux3-image-test/
What else?
- Bunch of visible shape inconsistencies: building window sizes and alignment, fence railing, rivets on the bench, coins, accordion features, perspective violations etc...
- Metal clips on the bottom and top edges of the carrying case are on different sides. Said clips look inconsistent.
- One of the case straps originates inside (from under the velvet, even), the further one originates outside. The carrying handle is oddly mis-centered (as with most things).
- I'm not sure that accordion, compressed, fits in that case. She can just about stand in it (diagonally, on one foot).
- It has weird folds that disappear partway through and turn into adjacent folds going the other way. Between this, the keyboards and the dots, I'm sure there are other things wrong with accordion anatomy.
- Awkward left hand: not sure what's going on with the metacarpal joints, and index finger appears to begin forking into two tips pressing into two (misshapen) buttons. Her right hand is getting there too.
- Weird hair midline. I haven't seen a "sub-midline" like this in the wild :)
- The burgundy coat clips through the bench under her right leg. It's like the front slat goes through a hole in the side of the coat.
- We don't usually have that side curling bar on each side of the sitting area. It looks like it's clipping through the flats next to the leather handbag
- Ill-placed buckle on that bag.
- The green metal fence on the right begins behind her for one vertical bar, next one disappears into the plants, no further vertical bars, no top horizontal railing. Actually, it appears to only have fencing on one side, that's not something I'd see often. On the side where it does have a top, its top rail has very irregular width.
- Sand looks like... breadcrumbs? Very repetitive coarse shapes, like it was poorly done with a clone stamp.
- Something weird about the leaves.
- The foremost lamp isn't sure whether it has flat sides or not.
- The other lamp emerges out of nowhere from behind one of the blob-people in the alley.
- A bunch of unfortunate tangents and occlusions which make you wonder whether the diffusion hallucinated detail out of bigger shapes, or if it's been trained on photos where the tangents were done on purpose.
ie. center bottom bar of the bench being exactly aligned with the gray edge tiles, side curly railing-not-present-on-our-real-benches fully formed on her right, almost completely foreshortened out of existence where the bench's perspective forbids it...
Most people are fooled by much cruder fakes though. This is a lot better but keeps plenty of tells for observant people (and most people aren't)
Or just put in more than 10 words of effort, e.g. a better prompt and bounding boxed prompt edits, and all that stuff would be fixed now.
As I'm finding with the best GenAIs, this allows for granular iteration, which is where it becomes useful in an industry-wide manner.
That would be to compare e.g. Qwen Image 3.0 with FLUX 3, with Midjourney etc.
I had seen some attempts - but I do not know well how they try to approach objectivity.
Prompt results are graded based a weighted calculation which includes: adherence to the prompt, image fidelity, and steerability.
My comparison benchmark also tends to favor prompt adherence, which a lot of others don’t. Most of ones that I've seen tend towards rather simplistic prompts (e.g. "neon-lit city facing a robotic uprising, with high-tech battles, in anime style"), whereas the prompts I've created try to test high specificity.
I’ve been running them all the way back to SDXL.
You can compare specific models using the "View All Models" so if you want to see the progression of open-weight models, or model X vs model Y, you can do so.
Just a heads up - I haven't added Flux 3 as I'm waiting until BFL drops the open-weights version.
Generative Comparisons:
https://genai-showdown.specr.net
Editing Comparisons:
https://genai-showdown.specr.net/image-editing
https://bfl.ai/pricing
Tried adjusting exposure; Didn't work as well.
If you do not consider contemplation, for which we have filled our homes with works of art for centuries,
and if you do not consider those job that involve graphics production e.g. in the Madison Avenue "Mad Men" business (advertising - the payments are authentic),
you could consider "authentic" that when we do video production through generative models some start from stills (image generation) and then have other models animate those stills.
What's new from the last post? GA?
> We will open up an early access phase for FLUX 3 Image in the following weeks.
Not sure if there was a separate post for early access or if they just skipped to this.
How is this not clear?
Exact prompt:
Generate a photorealistic image of M81 urban BDU camouflage cargo trousers, shown by themselves. One trouser leg should be posed with the knee lifted 30° from vertical.
Accurate reproduction of the M81 urban camouflage pattern is critical. Match its colors, shapes, scale, distribution, and overall appearance as faithfully as possible.
No person, other clothing, or props.
Ground truth swatch: https://commons.wikimedia.org/wiki/File:US_City_Camo_(M81_Ur...
Gemini 3 Pro Image (stronger pattern): https://i.postimg.cc/bZNQYYjx/2026-10-02-google-gemini-3-pro... Flux 3 (this run): https://i.postimg.cc/Xr7wNN0H/2026-10-02-black-forest-labs-f...
Flux 3 gets greyscale urban-ish trousers and a lifted knee, but the blotches aren’t real M81 Urban — softer / wrong geometry vs the swatch. Not the worst I’ve seen on this prompt; clearly behind the Gemini 3 Pro Image example above on pattern.
Curious what other models do on the same prompt.
Flux is in the top 9000 of the most common words.