Sunstone North Lab · all-HTML version

A picture retrieves its own words

A read-out head trained on text only, attached to the frozen multimodal google/gemma-4-31B-it, reads images with zero image training. Across CIFAR-10 the image-to-word retrieval scores 0.93 recall at 1 against 0.10 chance. Everything below was captured offline in one prefill pass per image and replayed as pure geometry. Some categories are sharp and some miss; the misses are shown because the hits are only meaningful next to them.

bicycle (cifar100)

bicycle thumbnail

Image → word retrieval: bicycle 0.81, wheel 0.75, motorcycle 0.74, car 0.62, engine 0.52

Top captions from an open pool of 10,000 (never trained on):

Nearest discourse communities (learned from prose): cycling 0.58, cars 0.42, fitness 0.00

bird (cifar10)

bird thumbnail

Image → word retrieval: bird 0.82, frog 0.71, insect 0.68, wildlife 0.64, deer 0.63

Top captions from an open pool of 10,000 (never trained on):

Nearest discourse communities (learned from prose): gardening 0.69, knitting 0.23, Mommit 0.04

car (cifar10)

car thumbnail

Image → word retrieval: car 0.84, motorcycle 0.81, bicycle 0.80, wheel 0.69, vehicle 0.66

Top captions from an open pool of 10,000 (never trained on):

Nearest discourse communities (learned from prose): cars 0.66, cycling 0.30, financialindependence 0.02

castle (cifar100)

castle thumbnail

Image → word retrieval: castle 0.81, sea 0.60, clock 0.54, ship 0.52, building 0.49

Top captions from an open pool of 10,000 (never trained on):

Nearest discourse communities (learned from prose): pagan 0.41, Christianity 0.19, DebateReligion 0.14

bottle (cifar100)

bottle thumbnail

Image → word retrieval: cup 0.71, kitchen 0.66, food 0.64, mushroom 0.64, bottle 0.61

Top captions from an open pool of 10,000 (never trained on):

Nearest discourse communities (learned from prose): EatCheapAndHealthy 0.92, knitting 0.03, fitness 0.02

cat (cifar10)

cat thumbnail

Image → word retrieval: frog 0.73, bird 0.68, pet 0.62, flower 0.60, cat 0.60

Top captions from an open pool of 10,000 (never trained on):

Nearest discourse communities (learned from prose): Mommit 0.29, knitting 0.25, relationship_advice 0.18


0 images in head training · 0.93 CIFAR-10 R@1 · 10,000-caption open pool · one 22 MB head, the whole bridge.