FIELD NOTES01
How vision-language models learn where things are
Spatial grounding depends on preserving coordinates, local visual tokens, and training the decoder to speak in regions.
Essays, notebooks, and field notes connected by spatial reasoning.
Spatial grounding depends on preserving coordinates, local visual tokens, and training the decoder to speak in regions.