One thing that I've wondered about is how responsive to dog whistles the models would be. I think it may be possible to craft the same question with phrasing variants that are politically coded (based on an existing corpus) and see if it triggers different outputs.
♥ 3