AI-Torture-Chamber

(github.com)

2 points | by rozumbrada 5 hours ago ago

1 comments

  • rozumbrada 5 hours ago ago

    Steering language models into strong negative and positive valence states, and measuring what they say and what they're willing to do about it.