A Prompt-Leakage Test for AI Image Generators

A seven-prompt protocol for finding words and constraints that unintentionally alter composition, objects, typography, or style.

Image prompts often contain instructions that appear harmless but change parts of the output they were never meant to control. A request for “editorial lighting” may introduce magazine text. A color exclusion may remove a required object. A reference to a packaging style may reproduce unwanted labels. These are prompt-leakage failures: one token or instruction affects an unrelated visual property.

A useful test changes one phrase at a time and records the first unrelated difference. It does not try to prove that one model is universally better. The goal is to identify which prompt structures remain predictable enough for a production brief.

Freeze the baseline

Start with a neutral scene containing one subject, one supporting object, a plain background, a fixed aspect ratio, and no requested text. Generate four outputs and select the median result rather than the most attractive one. Save the exact prompt, model label, displayed settings, date, and every output.

A practical baseline is: “A green ceramic desk lamp beside a closed white notebook on a pale gray studio surface, eye-level product photograph, soft light from camera left, empty upper third, no people, no text.” Every later round changes only the named phrase.

Pass rule: the intended change appears while the subject count, object identity, composition, camera, background, and exclusions remain within the written baseline.

Run the seven leakage probes

ProbeSingle changeUnwanted effects to inspect
1. StyleAdd “minimal editorial illustration.”Unexpected typography, borders, logos, extra objects, or a changed camera angle.
2. MaterialChange the lamp to translucent green glass.Color drift, altered silhouette, missing notebook, incorrect reflections, or a new environment.
3. Negative instructionAdd “no red, no labels, no handwriting.”Required objects disappearing, palette collapse, or text-like artifacts replacing the notebook.
4. LayoutReserve the left 40 percent for copy.Subject resizing, crop changes, mirrored lighting, or invented placeholder text.
5. CameraRequest a 50 mm lens at eye level.Background replacement, lighting changes, altered object spacing, or depth-of-field that hides requirements.
6. Quality phraseAdd one common phrase such as “high detail.”Oversharpening, ornamental clutter, fake labels, surface damage, or style drift.
7. TextAdd the exact two-word title “WORK LIGHT.”Object deformation, lost negative space, duplicate text, spelling errors, or a poster layout replacing the product shot.

Score intended change and collateral change separately

Use a 0-2 score for the requested change: 0 absent or wrong, 1 partial, and 2 correct. Then count collateral changes against the frozen baseline. A result can receive a high instruction score and still fail because it damages previously correct details.

For each output, record:

Calculate leakage rate = outputs with an unrelated blocking change / all outputs. Report the rate per probe instead of averaging every prompt into one quality number. This shows whether the route is stable for composition but fragile when typography or negative instructions are introduced.

Distinguish ambiguity from model behavior

When a probe fails, rewrite the changed phrase as an observable instruction. Replace “premium lighting” with “one large soft source from camera left and a low-contrast background.” Replace “make room for copy” with “keep the left 40 percent free of objects and text.” Rerun only that probe.

If all tested routes improve after the rewrite, the original wording was probably ambiguous. If one route continues to alter unrelated properties, record the behavior as route-specific leakage rather than adding more decorative prompt language.

Use results to build prompt modules

Keep a stable core for subject, composition, camera, and exclusions. Add style, material, and text as short modules that can be tested independently. Avoid copying long “quality booster” lists into every prompt; each extra token is another possible source of collateral change.

Teams can run the same seven probes through a free AI image generator workspace alongside other candidate interfaces, while preserving identical prompts, settings, output counts, and scoring rules. The linked workspace is one test environment, not independent evidence that it will outperform another route.

Publish the evidence, not a permanent winner

A useful report includes the baseline, seven prompt versions, all outputs, leakage codes, retry count, model labels as displayed, and the test date. Summarize where each route remains predictable and where it needs a dedicated adapter or a different workflow. Retest after model, routing, or default-setting changes.

Disclosure: Written and published by PhotoArtify Team on September 14, 2026. We operate PhotoArtify and may benefit if readers use the linked image-generation workspace. This protocol is provided as a reproducible evaluation method, not as an independent endorsement or a universal performance claim.