I suspect this is less about AI and more about design. Use positional sound and some matching on-screen representation so that we can process people talking at once because they're spatially distinct. At least that's my armchair guess.
♥ 3