Same function, different implementations
I studied how Transformers compute multiplication in S₅. Representation analysis and activation interventions identified three families of product representations. Models computed these representations through different MLP, attention, and residual paths.
Evidence and scope
The study compares a small set of trained models. Optimizer preferences varied across seeds. These observations do not establish a deterministic mapping from optimizer to circuit or a general statistical law.
The figure summarizes the representation and intervention results from my S₅ study. It is a schematic summary, not an additional experiment.