AI Research
Turning Raw LLM Activation Vectors into Readable Prose
Anthropic's Natural Language Autoencoders (NLAs) bridge the gap between raw activation vectors and human-readable prose to decode LLM internal states.
Anthropic's Natural Language Autoencoders (NLAs) bridge the gap between raw activation vectors and human-readable prose to decode LLM internal states.