> In this world view, nano banana is a first early hint of what that might look ...

simonw · 2025-12-20T01:48:18 1766195298

What's interesting about Nano Banana (and even more so video models like Veo 3) is that they act as a weird kind of world model when you consider that they accept images as input and return images as output.

Give it an image of a maze, it can output that same image with the maze completed (maybe).

There's a fantastic article about that for image-to-video models here: https://video-zero-shot.github.io/

> We demonstrate that Veo 3 can zero-shot solve a broad variety of tasks it wasn't explicitly trained for: segmenting objects, detecting edges, editing images, understanding physical properties, recognizing object affordances, simulating tool use, and much more.

dragonwriter · 2025-12-20T01:01:42 1766192502

I think he is referring to capability, not architecture, and say that NB is at the point that it is suggestive of the near-future capability of using GenAI models to create their own UI as needed.

NB (Gemini 2.5 Flash Image) isn't the first major-vendor LLM-based image gen model, after all; GPT Image 1 was first.