Browse / Data Science Ml / MechInterp Glossary and Constraints

MechInterp Glossary and Constraints

Provides domain-specific knowledge and experimental constraints for mechanistic interpretability research on Splatoon data models.

SkillData Science MlCode Search

Key features

  • Reference of ability family short codes and AP rung values
  • Guidelines for handling binary tokens and activation deltas
  • Standardized schemas for token parsing and build validation
  • Validation rules for gear-specific and main-only abilities
  • Identification of experimental pitfalls like ReLU floor quantization

Use cases

  • Planning mechanistic interpretability sweeps while avoiding out-of-distribution errors
  • Verifying the validity of ability combinations in experimental builds
  • Understanding the relationship between AP investment and model activations

FAQ

What does the MechInterp Glossary and Constraints skill do?

This skill acts as a domain-specific knowledge base for mechanistic interpretability (MechInterp) research. It provides Claude with precise references for Splatoon ability families, AP rungs, and the technical constraints needed to design valid experiments on SplatNLP models.

How does this skill improve my AI research workflow?

It prevents 'hallucinated' or invalid experiments by providing Claude with hard constraints. For example, it ensures Claude doesn't attempt to sweep multiple AP rungs for the same ability family or place 'Main-Only' abilities in incorrect gear slots, saving time on failed runs.

When should I use this skill?

Invoke this skill when you need to look up ability codes (like SCU or SSU), verify if a specific gear/ability combination is valid, or understand the token format for binary abilities like 'Ninja Squid' before running activation sweeps.

What specific mechanistic interpretability pitfalls does it address?

The skill identifies common experimental errors such as ReLU floor quantization (where low base activations make delta measurements unreliable) and OOD (Out of Distribution) token combinations that could lead to misleading feature analysis.

How does it handle binary vs. stackable abilities?

It provides strict token formatting rules and evaluation strategies for binary abilities (which have no AP suffix) versus stackable ones, ensuring Claude uses the correct statistical methods like PageRank or manual 2D analysis for each type.