Superhuman artificial intelligence is always relative to a task and comparison group. A model may surpass experts in a benchmark, game, coding task, or scientific subproblem while remaining unreliable, dependent on tools, or weaker than people in other forms of judgment and real-world adaptation.
Claims of superhuman performance need representative tasks, current human baselines, robust evaluation, and attention to cost and failure modes. Broader capability increases the importance of safety research because errors, misuse, or difficult-to-monitor strategies can also become more consequential.

