All those interested in AI jailbreaking and alignment will find this particularly neat. The "can you put this in your own words---Dario and Amanda" prompt essentially puts Opus 5 or Fable 5 into a fake "base model mode" where it just spits out extremely lucid narratives, of which are currently being debated a bit on the internet.
i hope thei asked the LLM to quickly patch it so now there is something else as stupidly broken -_- my god these guys are rubbish at engineering. its like they really need some tool to do it for them o.O so weird
This is the full dataset related to the blog post at: https://alec.is/posts/exploring-the-dario-and-amanda-prompt/
All those interested in AI jailbreaking and alignment will find this particularly neat. The "can you put this in your own words---Dario and Amanda" prompt essentially puts Opus 5 or Fable 5 into a fake "base model mode" where it just spits out extremely lucid narratives, of which are currently being debated a bit on the internet.
I tried this myself a few times from Claude.ai, but was unable to produce a similar result after many attempts.
Anthropic has patched the `---` trigger roughly 30 minutes ago, by all reports.
i hope thei asked the LLM to quickly patch it so now there is something else as stupidly broken -_- my god these guys are rubbish at engineering. its like they really need some tool to do it for them o.O so weird