nanointerpret
LLM Interpretability Playground
Last verified:
What is nanointerpret?
NanoInterpret is an interactive tool for exploring and manipulating features learned by Sparse Autoencoders (SAEs) in language models. Users can force-activate specific neuron features at variable intensities and observe how they affect generated text, enabling both model interpretability research and controlled text generation experiments.
nanointerpret pricing
Pricing model: Freemium
nanointerpret pros
- Interactive dual-mode interface for both feature intervention and exploration
- Comprehensive feature browser with 3,000+ learned concepts and descriptive titles
- Granular control over feature activation strength (50%, 100%, or 200%)
- Real-time text generation with adjustable parameters (temperature, top-p, repetition penalty)
nanointerpret cons
- Limited documentation on SAE methodology or which model is being interpreted
- 200% feature intensity may produce degraded or unpredictable outputs (acknowledged as 'may break the model')
- No information about access, usage limits, or integration with other tools
Frequently asked questions about nanointerpret
What happens if I set a feature to 200% activation?
The feature is activated at twice the maximum observed level, which may cause unpredictable or broken model behavior
How do I explore what the model has learned?
Use the feature dropdown to browse the concepts the SAE has discovered and see descriptions of what triggers each feature
Can I adjust how the text is generated?
Yes, you can configure Maximum tokens, Temperature, Top P, Top K, and Repetition penalty settings