Netflix staff software engineer Rajat Shah describes using AI agents to reduce the manual work of performance engineering. Instead of asking a model to optimize arbitrary code, the workflow starts with structured CPU profiling data, identifies expensive paths and traces them to the exact revision deployed in production.
The talk presents a quadratic-time pattern identified from a call stack and a repeated metrics-allocation pattern found across several services. These examples motivate a shared catalog of patterns and anti-patterns, stored as readable Markdown in Git rather than requiring a complex vector-memory system.
Shah keeps the agent's output at the level of a proposed code review. Functional tests and canary deployments check CPU usage, latency and errors under comparable traffic, while an engineer makes the eventual approval decision. A plausible optimization is not enough when it could change business behavior.
The longer-term approach moves learned patterns into code review and eventually code authoring, so inefficient implementations are caught earlier. The talk recommends building reliable profiling, integration and validation foundations first, then adding bounded orchestration. Greater autonomy requires stronger evaluations, sandboxing and protection against prompt injection.
Watch on YouTube




