Károly Zsolnai-Fehér examines research into how language models answer spatial questions even though their inputs do not directly contain character counts or page widths. The study probes a model performing a line-fitting task and finds internal features that track position along a line and proximity to its boundary.
The researchers compare those features with biological place cells and boundary cells, which activate according to an animal's location. In the model, related sparse features form curved geometric structures that let it approximate character counts from token positions and recognize when text is near the end of a page.
Károly Zsolnai-Fehér highlights that nobody explicitly programmed this internal representation. The model developed it during training because it improved reliability, suggesting that interpretability research can uncover compact computational tools that models invent for unfamiliar tasks. The result does not prove human-like understanding, but it offers a concrete view of how useful abstractions can emerge inside a neural network.
Watch on YouTube


