Daniel Walsh reviews ten Claude sessions and ten Codex sessions from the same application using a retrospective skill. He explains how analyzing an agent's path to a result can reveal repeated commands, missing context and wasted work that are not obvious from the final code.
Daniel Walsh identifies four recurring problems: unclear Windows startup instructions, test accounts that do not reproduce real users' saved configuration, long-standing failing tests that obscure regressions, and scripted edits that corrupt special characters. He distinguishes improvements to task context and test coverage from failures that need automatic protections.
Daniel Walsh argues that a memory reminder is insufficient when agents repeatedly damage files. He describes checking edits before accepting them and preserving originals on failure, positioning session reviews as a complement to linters, type checks, tests and dependency scanning rather than a replacement.
Watch on YouTube




