Prompts
Jon → Hermes (spoken input)
Hermes → ComfyUI (KREA2 turbo t2i)
Generated Gallery Image
tldraw Canvas — GUI Export
What Happened
Generation: The gallery image was generated in 38 seconds using KREA2 turbo text-to-image. The prompt called for a Jeff Koons balloon dog, modern art sculptures, and interesting architecture — the model delivered a convincing gallery space with a skylight, track lighting, a reflective balloon dog, a bronze figure on a pedestal, a colorful star sculpture, and a large mirror sphere.
Vision analysis: Before drawing, the agent used a vision model to analyze the generated image in precise spatial detail — identifying the room shape, vanishing point, each sculpture's position and size, the skylight angle, the light beam direction, and the architectural cutout in the back wall. This analysis guided the drawing.
First drawing attempt — flat: 34 geometric shapes (rectangles, ellipses, stars) placed as flat 2D blocks. Jon immediately said: "You gotta draw the perspective. Do the perspective lines so we see the floor, the far wall."
Second drawing attempt — perspective: 42 shapes with actual depth. The floor became a trapezoid narrowing toward the back wall. Left and right walls receded to a vanishing point. Pedestals were drawn as trapezoids (narrower at the back) to show perspective. Freehand draw shapes used compressLegacySegments for precise polyline paths — the skylight, light beam, track lights, and sculptures all placed in 3D space.
The result reads as a pen sketch — not photorealistic, but spatially coherent. The limitation: tldraw's draw shapes are uniform-stroke polylines with no fills or smooth curves, so the balloon dog (which has tube-like curves) is the weakest element, represented as overlapping circles. But the room perspective, the converging lines, and the sculpture placement clearly convey the 3D space.