Topological Data Analysis of Large Language Models
Topological Data Analysis of Large Language Models: A Persistent-Homology Framework for Neural Representation Geometry Abstract Large Language Models (LLMs) have demonstrated remarkable capabilities across a broad spectrum of natural language processing tasks, yet the internal mechanisms governing their representations remain fundamentally opaque. We propose a rigorous methodological framework utilizing Topological Data Analysis (TDA), specifically persistent homology, to characterize the hidden geometric and topological structures of these neural activations. Rather than presenting experimental findings, this paper serves as a comprehensive methodology proposal designed to transition the analysis of LLM embeddings from heuristic geometric approximations to formalized topological invariants. By treating the outputs of self-attention heads and feed-forward networks as dynamic metric spaces, we construct Vietoris-Rips filtrations to trace the birth, persistence, and death of to...