Wiki / concepts / wiki

Superhuman blind spots

A system can be superhuman almost everywhere and still hold "tricky lumps of knowledge" where it is confidently wrong. alphago would sometimes judge a dead group alive (or the reverse), and the team couldn't predict when it would wander into one of these lumps or fix them before the match. fan-hui found them in practice games; lee-sedol's game-4 move 78 steered AlphaGo into one. It then played visibly nonsensical moves while its value estimate collapsed ("it's on tilt"), and resigned.

Lessons: aggregate strength doesn't guarantee robustness; when it fails it turns delusional, as in game 4; and a determined adversary (or a creative human) can find the holes. It's the game-playing ancestor of the "Swiss-cheese capabilities" framing for LLMs (andrej-karpathy) and of verify-like-a-junior-analyst.

Source: report

Linked from

AlphaGoFan HuiLee Sedol