Tejas Bhakta on Autoresearch for Faster GPU Inference

AI Engineer7m 30s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Tejas Bhakta describes autoresearch as a loop that proposes changes, tests correctness and speedAn autonomous research loop is a workflow in which an AI system repeatedly proposes, runs, evaluates, and updates research with limited human intervention., and keeps or reverts the results. GPU kernelsA graphics processing unit kernel is a function compiled to run many parallel operations on a graphics processor or related accelerator. suit this process because both outcomes can be measured, but the human still supplies architectural ideas while the agent searches implementation parameters.

    Tejas Bhakta recommends profiling compute, memory and launch overheadPerformance profiling measures where a program spends time or resources so that AI system improvements target actual bottlenecks., then giving the agent precise hardware and model context. Optimizations need realistic workloads: an agent may improve one kernel while slowing the whole system, disable CUDA graphs, or test only short contexts. A faster replacement may therefore need a limited operating range rather than becoming a universal default.

    Tejas Bhakta reports that combining custom kernels and hardware-level tuning delivered roughly threefold inference accelerationAI inference is the process of running a trained model on new input to produce a prediction, classification, generated response or action. at Morph. The result comes from accumulating useful improvements while rejecting failures, not assuming each experiment succeeds; Tejas Bhakta estimates that most attempted changes are poor and require strict checks.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Tejas Bhakta in a blue top beside the blue and white headline AUTORESEARCH FASTER INFERENCE on black. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 26 September 2026 and duration 7m 30s.

    Tejas Bhakta explains how human-guided autoresearch can accelerate GPU inference by testing kernels against correctness and speed while guarding against misleading benchmark gains.