An image model test?

Image model competition is heating up

I find the wallpaper a constant source of joy. Looking at my phone after 10AM and seeing what interesting new take on the daily news appears. The only thing I think I’d like to change is to shift the sources of news inspiration away from the AI bro news which is, frankly, all a bit one-note, is it not?

I digress, what I wanted to talk about is image models in general. I just added support for Muse Image and I was pleasantly surprised with the test generation. I’ll tilt the 10AM gen to use Muse Image so you can see for yourself. Strangely, this one uses the image API endpoint rather than the chat endpoint like all the other image gen models use, so it needed a tiny bit of extra plumbing.

This has got me to thinking. It might be nice to do an image spread as a test. My appetite for spending more on inference is limited, so I was thinking of a weekly job, something that looks at the non-AI news for the week perhaps. I can’t automate adding new image models, I think, since they tend to take different parameters, but they are infrequent enough I try to stay on top of them. The test would take our current image registry and the same model generated prompt.

I’m thinking landscape for this. It’d even be fun to automate a little video with some cross fades between them, I could find a screen at work to put that on.

Some issues to think through though. My registry classifies image generators by their capabilities for text generation, separating image specialists from those with text prowess good enough they are suitable for infographics.

This also makes me think about the v1 deep dive which had an endearing prompting weakness that had it tend to use the deterministic charting tool quite a lot. I liked the idea of setting up a research piece which would be focused on data, with the output as an infographic. I think perhaps I’m talking about a different kind of experiment there and the biggest challenge with that is the automated discovery of material with sufficient data to yield an infographic.

There I go on a tangent again. The weekly image comparison sounds found though right? The same news theme rendered out through a series of visual models. In brainstorming this, Claude reminds me that the shelf’s goal is not-benchmarks. Dammit, that’s right. So what if we do vary the prompting across the models but they’re focusing on the same topic. Would we constrain the visual style so the crossfades are less jarring (like the daily wallpapers)?

Hmm. Perhaps we do not provide the style. Is it not interesting to see what style an image model chooses when a scene is described? Probably the styles would fit better, and be at least thematically on point between the prompts. I quite like this.

You know what this means though. This is an easy half day job to cook up, which means it’s going to put off the monster Juggler concept for a while. That fits with the nerd-home balance though, our new kitchen renovation will be completed today, which means we need to move all this kitchen stuff back into the new kitchen-palace. Snack sized weekend experiment is just what is required.

← all meatspace posts