Can anyone in the know ELI5 how a defence against such an attack using AI would be organized? I’m struggling to understand what actions frontier models would perceive as offensive
Fable has been significantly guardrail'd to the point it refuses to work on login pages because they have auth related activities (as reported by another HN user). There are many such cases of people reporting Fable unwilling to work on innocuous tasks
Can anyone in the know ELI5 how a defence against such an attack using AI would be organized? I’m struggling to understand what actions frontier models would perceive as offensive
Fable has been significantly guardrail'd to the point it refuses to work on login pages because they have auth related activities (as reported by another HN user). There are many such cases of people reporting Fable unwilling to work on innocuous tasks