Sangwu Lee frames Krea 2 as an image foundation model designed for fast creative exploration and stylistic diversity. He contrasts that goal with slower systems that produce highly consistent results but can converge on a narrower visual distribution.
Sangwu Lee says the architecture becomes relatively fixed early, while most of the continuing work goes into the data. Krea deduplicates and rebalances concepts, removes synthetic images that could import another model's aesthetic, preserves unconventional styles, and builds detailed captions from OCR, metadata, and vision language models.
Sangwu Lee describes a staged training pipeline that progresses from low to high resolution pretraining through mid training, supervised fine tuning, preference optimization, reinforcement learning, and prompt expansion. His practical priorities are fast infrastructure, durable data quality, simple scalable methods, and future architectures that reduce separate encoders while using richer spatial conditioning such as bounding boxes and scene graphs.
Watch on YouTube



