TransformerLens

TransformerLens

Mechanistic interpretability for GPT-style models

Description

Understanding how an LLM thinks internally, like what one attention head does, is hard to dig into with ordinary frameworks. TransformerLens is a library for mechanistic interpretability of GPT-style language models.

It loads dozens of open models and makes reading and editing activations at any layer easy, a standard interpretability tool.

Features



Activations:Any layer.

Hooks:Intervene mid-run.

Models:Dozens supported.

Standard:Community tool.