sarankumar.space

The Many Minds of Transformers

August 31, 2025

What each attention head is really learning, and why running several in parallel gives models richer representations.

Originally published on Medium (~920 words). Read the full piece here: The Many Minds of Transformers.

PreviousHow Transformers Generate TextNextAttention Is All You Need to Understand