site link: https://www.ai-transparency.org/

paper link: https://arxiv.org/pdf/2310.01405.pdf

In Neuroscience or ‘cognitive neuroscience’, there’s a Sherringtonian view and a Hopfieldian view.

The Sherringtonian view aims to understand cognition as a result of node-to-node connections, implemented by neurons as a part of circuits in the brain.

The Hopfieldian view sees cognition as a product of ‘representational spaces’, implemented by patterns of activity across populations of neurons.

In AI transparency, Mechanistic interpretability aligns with the Sherringtonian view but there isn’t anything analogous to the Hopfieldian view.

Untitled

RepE can address many of the safety concerns that MechInterp was intended to address.

This work looks at RepE for LLMs in 2 main areas: Reading and Control.

Representational Reading

Seeks to locate representations for high-level concepts within a network.

They try to extract concepts for truthfulness, morality (harmful, helpful) and emotions (anger, fear, sadness).

Linear Artificial Tomography

a LAT scan is made up of 3 steps