Bijan Bowen tests Ling 3.0 Flash, a mixture-of-experts model with 124 billion total parameters and about 5.1 billion active per token. Its size and 256K context make the promised open-weight release potentially practical for high-memory local systems, although quantized performance remains untested.
The model repaired an initially broken browser-style operating system and produced a surprisingly complete C++ skateboarding game. It also improved a first-person game over several feedback rounds, showing that it could locate errors, revise large files and continue working deep into a long context.
Results were less consistent on spatial and unfamiliar tasks. A generated watch had orientation problems, the printable engine did not resemble the requested object and a rally game remained stuck on a black screen despite repeated rewrites. Several successful scenes were visually simple even when their underlying interactions worked.
Bijan Bowen concludes that Ling 3.0 Flash has strong foundational coding behavior for its active size, especially as a local companion to a separate creative model. The main open questions are how much capability survives quantization and whether the announced weights become available as expected.
Watch the original on YouTube