AI safety

2 articles

Common Obligations

Six shared obligations for AI builders, and a proposal for making them hold across companies and borders without putting the largest labs in charge.

AI governance AI safety accountability international cooperation
Read more

When Outputs Lie

Your AI agent's outputs look composed. Its internal state is desperate. Anthropic's emotion vectors research reveals a second axis of agent drift that output evals can't catch.

autonomous agents interpretability AI safety agent reliability
Read more