The setup is easy to state. Take an image. Have a vision model write descriptions of it at fixed word budgets, ten words up to five hundred. Then have a different model, one that has never seen the image, answer thirty questions about it using only the description. Grade those answers against what someone looking at the picture said. The distance between what a reader recovers from words and what they recover from the image is the thing you are actually measuring.
Fifteen images, deliberately spread out. A red sphere on white as a control. Lincoln's portrait, Starry Night, the Declaration of Independence, Minard's chart of Napoleon's march, a hand drawn boiler schematic, Times Square at night, a Bruegel crowded with proverbs, pollen under an electron microscope. Eight conditions each, thirty questions each, three thousand six hundred graded answers.
The first version failed, and the way it failed was the most useful thing I learned all day.
I had included a control where the answering model got nothing at all. No description, no category, no title, just the thirty questions. That was supposed to be the floor. It scored 74 percent.
The questions were leaking. One of them read "how has the snake's body been divided", and once you have read thirty questions like that you have effectively read a description of the picture. A quiz written about an image is itself a compressed version of that image. I had to rewrite every question so that no stem named anything visible in the frame, referring to things only by position and role, then rerun the whole pipeline.
That is worth sitting with, because it is not only a bug in my setup. Any benchmark that asks a model questions about an image is partly measuring how much the questions gave away.
With the leak closed, the numbers get interesting.
Knowing nothing gets you 66 percent. Seeing the actual image gets you 87 percent. That span, from ignorance to sight, is what a description has to buy back. Five hundred words buys back 74 percent of it. Not 74 percent of the image, 74 percent of the distance to the image.
Share of the gap between knowing nothing and seeing the image, recovered at each word budget. Shaded band is a 95 percent bootstrap over images. Tap to enlarge.The shape matters more than any single point on it. The first twenty five words deliver more than half of everything five hundred words achieves. Between ten and twenty five words each additional word is worth roughly forty times what a word is worth between two hundred and fifty and five hundred. Description is enormously front loaded and then it flattens.
Fit that curve and extend it and a thousand words lands at about 84 percent recovery. Ninety percent takes around sixteen hundred words. Ninety five takes something like twenty four hundred. The confidence intervals on those extrapolations are wide enough that I would not defend the exact figures, because this is fifteen images and not a hundred thousand. The direction is clear though, and it is quietly funny that the proverb lands roughly where it does. A thousand words gets you most of a picture. It does not get you the picture.
The second half of the experiment fed each description straight into GPT Image 2 and had a judge score the reconstruction against the original.
Same shape. Fidelity climbs from 71 at ten words to 88 at five hundred, and the last two hundred and fifty words move it less than a point.
The sub scores split in a way I did not expect. Getting the subject right is nearly free. At ten words the reconstructions already score 91 on whether this is the same kind of scene. Detail is what you are paying for, and it starts at 57 and only reaches 81 after five hundred words. Words are cheap for identifying and expensive for specifying.
Then there is the result I keep thinking about.
I had tagged eight images as famous enough that a model might recognise them by name. At ten words those recovered 45 percent of the gap. The unfamiliar ones recovered 12. Four words of Van Gogh's Starry Night reproduce the painting almost exactly, because those words are not describing anything. They are an address for something the model already holds.
The same word budgets, reconstructed by GPT Image 2 and scored against the original. Tap to enlarge.The food still life next to it has no name to invoke, so ten words scores 35 and every word after that does real descriptive labour. By fifty words the unfamiliar images have caught up and passed the famous ones, which is what you would expect if naming is a shortcut that runs out.
So the honest answer is that a picture is not worth a fixed number of words. It is worth however many it takes to close the gap between what your reader already knows and what is in the frame, and for anything genuinely new that number is large and climbing slowly.
Where it stops working altogether is text. The Declaration of Independence scored 62 on reconstruction, the worst in the set, and five hundred words could not carry it. The honest description of a document is the document.
What I want to run next is the inverse. Rather than asking how many words an image needs, ask which words do the most work. Add ten at a time, measure what each addition buys, and find out whether visual information concentrates in a few high value terms the way this suggests it does.
