GPT-5.6 Sol leads TryAI Canvas Arena drawing benchmark
TryAI benchmarked four frontier AI vision models—GPT-5.6 Sol, Claude Fable 5, Grok 4.5, and Gemini 3.6 Flash—across 28 drawing tasks using an open-source colored pencil toolset on a blank digital canvas. The models executed commands to adjust stroke attributes, lay down colored marks, smudge, erase, and self-review their canvas across target image reproductions (the Mona Lisa and Starry Night) and open-ended text prompts. The evaluation highlighted stark differences in tool-calling behavior, efficiency, and cost: GPT-5.6 Sol produced the best visual details for $7.74 across seven drawings, while Claude Fable 5 cost $160.58 due to deep reasoning overhead and excessive steps. Notably, across all models, structural similarity scores peaked mid-run and declined during later iterations, demonstrating that continuous self-reviewing and over-editing degraded output quality.
Interactive stroke-by-stroke drawing benchmarks expose significant cognitive and execution differences between frontier vision models that static image benchmarks miss.
- –**GPT-5.6 Sol Leads in Quality & Cost-Efficiency**: Sol set stroke parameters inline, took fewer average steps (29 per drawing), and delivered superior visual detail for under $8 across seven drawings.
- –**Claude Fable 5 Incurs Extreme Cost Overhead**: Fable 5 accumulated an estimated $160.58 bill (~20x higher than peers) due to long reasoning loops and heavy tool usage without matching Sol's visual results.
- –**Over-Editing Causes Quality Degradation**: In every target reproduction run, models hit their peak Structural Similarity (SSIM) score mid-task and subsequently eroded fidelity with further edits.
- –**SSIM Scores Disconnect from Human Perceptual Artistry**: While Gemini 3.6 Flash recorded the highest raw SSIM scores due to macro color placement, human reviewers consistently preferred Sol's fine line work and shading.
DISCOVERED
2h ago
2026-07-22
PUBLISHED
5h ago
2026-07-21
RELEVANCE
AUTHOR
hershyb_