Evoke brings persistent memory to world models
Evoke is a 14B, three-step open-source world model that generates interactive 384×640 video at 24 fps while responding to camera controls and mid-session prompts. Its external, camera-indexed world-state memory targets longer, more coherent simulations without unbounded context growth.
Evoke’s real breakthrough is architectural: persistent world state sits outside the denoiser, making interactive video generation more practical for extended sessions. The release is promising for research, though its H200 requirements and restrictive depth-model licenses limit immediate production use.
- –Three-step, CFG-free generation improves responsiveness compared with many-step video models
- –External geometric memory helps preserve spatial continuity as users explore
- –Mid-rollout prompt changes enable controllable synthetic environments and scenario branching
- –The released stack supports text-to-video, image-to-video, video-to-video, and segmented prompt switching
- –Reported performance is strong on WBench, but hour-scale coherence still needs broader independent validation
DISCOVERED
2h ago
2026-08-23
PUBLISHED
2h ago
2026-08-23
RELEVANCE
AUTHOR
AI Search
