Product news
Why AI ads get hands and text wrong, and how we fix it
Chrisphine Otieno · 11 Oct 2026 · 2 min read

Two mistakes give AI ads away at a glance: hands with the wrong number of fingers, and words that are almost, but not quite, spelled right. On a poster with your price and phone number, a single wrong letter can cost you the sale.
Why it happens
Image and video models learn from pictures, not from rules. Hands appear in countless poses and are often half hidden, so models never learn "five fingers" as a fact. Letters are similar: the model learns what text looks like, not how to spell. The longer and more detailed the text, the more likely it drifts.
What the Studio does on every request
1. The prompt director asks for it explicitly. Before your request reaches the model, the director rewrites it into a production prompt with fixed quality rules: every visible hand natural, with five fingers, doing a simple clear action (holding, pointing, tapping a phone); no close-ups of complicated finger poses.
2. Only the words you asked for. The director quotes on-screen words exactly as you gave them: large, bold, high-contrast lettering, fully inside the frame. No made-up words, no extra signs, no logos you did not provide. When there is nothing to write, it leaves space for you to add text later.
3. A safety net without the director. When the director is not available, every composed prompt still ends with the same ask for natural hands and exact, legible words. (If you choose Use my exact words, we send your words as written.)
4. Negative prompts where the model takes them. Models that accept a "what to avoid" field are told to avoid extra or fused fingers, distorted hands, garbled text and watermarks.
Pick the right model for text
For posters and price cards, pick Ideogram 4 (clear text) on the quote. It is built for sharp, correctly spelled headlines, prices and contacts. For videos, keep on-screen text short: a headline, a price and a contact.
Check before you generate
In scripts, press Edit words and read every on-screen line letter by letter. What you type there is what the model is asked to show.