Great concept! Not to break your bank too much, but I'd love too see this as a matrix across several different providers and/or local models. I only have quick access to gpt-4o at the moment, and that was not fooled by any of the prompts listed, except for some of the last (200+ char) ones... would be cool to compare with llamas, 4o, claude, gemeni, etc...
Also, mentioned elsewhere but scoring by token count is definitely the way to go.
Also, mentioned elsewhere but scoring by token count is definitely the way to go.