The news: a video engine grafted on an image model’s ‘spatial attention.’ It is the eye that knows where everything is. Give it to video and the sets get richer, the textures come alive, and a character keeps the same face across dozens of shots. The old ‘sharpness creep,’ where every new shot blurred a little more, is gone as well. Where such a model runs, though, is a different story.

Nothing Blurs But the Invoice

Source: RT @C_of_Creativity: MiniMax-H3とZ-Imageの融合モデルでたー!! · twitter

▶ Animated edition: the characters speak for themselves (Korean audio, English subtitles included)

Download video

What this means for ThakiCloud

From ThakiCloud’s seat, this news asks the question again: where does the model run? A fusion model may be downloadable, but if the GPUs live in someone else’s building, per-frame billing follows. Metis serves models on your own GPUs, so serving cost ends as an electricity bill instead of an external invoice. Paxis runs the workflows on top, and because it is on-prem, nothing depends on another landlord’s meter. Just as spatial attention keeps a face consistent across shots, sovereignty over the runtime keeps the cost structure consistent.


An auto-generated comic riffing on this week’s industry news.

Tags: fusion-model, onprem, open-weights, serving-cost, sharpness-creep, spatial-attention

Categories:

Updated: