GPT-6 Astra Stumbles on Viral Canva Test
An independent creator tested GPT-6 Astra Extra High in Canva using computer use, without image generation. It beat Grok, but fell far short of the viral portrait demos, raising questions about reproducibility and cherry-picking.
This is a useful reality check, not a definitive model benchmark: computer-use demos can be highly sensitive to prompts, reference images, app state, retries, and editing.
- –Astra’s achievement is GUI control—placing shapes and strokes in Canva—not native image generation
- –A failed replication does not disprove the original demo, but weakens claims that the workflow is reliably push-button
- –Fair comparisons require identical prompts, source images, browser state, permissions, time budgets, and retry limits
- –Developers should evaluate full action trajectories and intermediate failures, not only polished final clips
- –Strong desktop-use benchmark scores do not guarantee consistent creative performance in consumer design apps
DISCOVERED
1h ago
2026-09-11
PUBLISHED
1h ago
2026-09-11
RELEVANCE
AUTHOR
SKatalystAI
