One subplot per dataset; each curve is averaged across layers. The plotted metric is Normalized MSE. Linear methods use the requested top-k budgets 1, 3, 10, 30, 100, 300. Linear curves extend from the largest requested budget to their measured complete-basis endpoints. SAE points are placed at their measured mean number of nonzero features; repeated budgets above native SAE activity collapse to one point. Stars mark complete-basis reconstruction for ICA, PCA, and Random, and native sparse reconstruction for SAE. For GPT-2, the dashed SAE curve is a training-context control evaluated only at token positions 0–63; solid curves use the full 1,024-position context.
