0xSero and David Ondrej compare practical hardware paths for running open AI models locally. Their discussion ranges from paired RTX 3090 systems and DGX Sparks to RTX Pro 6000 clusters, with attention to memory capacity, memory bandwidth, context length, concurrency and the active parameters in mixture-of-experts models.
0xSero describes compressing GLM 5.2 and other large open models, then demonstrates local agents performing coding and file-management tasks. He reports how quantization changes memory requirements and explains why model size alone does not determine useful inference speed or task performance.
The conversation also treats owned compute as an alternative to recurring API spend. It covers incremental upgrade strategies, private inference for organizations, and the physical constraints that appear as systems grow, including heat, noise, household circuits and power caps. Speculative political claims and promotional segments are omitted from this catalogue summary.
Watch on YouTube

