Every AI image tool - Midjourney, ChatGPT, Nano Banana, Sora - is a creative translator, not a creative partner. It executes whatever vision you hand it. Hand it a vague sentence, you get a vague, generic image. Hand it a specific, layered description, and you get something that looks art-directed.
The 6-layer framework is the structure behind that difference. It is the same methodology whether you're prompting for a portrait, a product shot, or a lifestyle scene, and it works identically across every AI image tool on the market today.
What elements should be in an AI image prompt
A strong AI image prompt includes six elements, written in this order: concept, subject, colors and materials, composition, lighting, and camera and lens. This is the direct answer to "how do I write better AI image prompts" - most people only ever describe the subject, which is layer 2 of 6. The other five layers are what separate a professional-looking result from a flat, average one.
Think of it like a photographer building a scene before ever pressing the shutter. Concept sets the type of shot. Subject fills in who or what is in it. Colors and materials give it texture. Composition frames it. Lighting sets the mood. Camera and lens locks in the photographic feel. Skip any layer, and the AI fills the gap with a generic default.
Why most AI image prompts fail
Generic prompt in, generic result out. That is the entire failure pattern, and it happens for three specific reasons:
- The prompt describes only the subject - "a woman in a red dress" - with no direction on composition, lighting, or camera.
- The prompt relies on vague adjectives like "nice," "beautiful," or "professional," which carry no visual instruction an AI can act on.
- The prompt skips descriptive visual language entirely. "Blue dress" tells an AI almost nothing. "Deep cerulean silk midi dress with cool purple undertones" tells it everything.
Every AI image tool defaults to the same safe, average output when given vague input - centered subject, flat even lighting, medium framing. The 6-layer framework exists to override those defaults, one layer at a time.
Generic prompt in, generic result out.
The 6-layer AI prompt structure
Here is each layer in the order you should write them.
Layer 1 - Concept
Concept defines what kind of image this is before you describe anything else - portrait, product still life, flat lay, lifestyle scene, editorial. This single decision frames how the AI interprets every layer that follows.
Without this
'a photo of a candle'
With this
'product still life of a candle'
Layer 2 - Subject
Now describe the main subject in full detail - material, color, texture, size relative to its surroundings, position, and any distinguishing detail. Every detail you skip is a detail the AI invents on its own.
Without this
'a nice candle'
With this
'a hand-poured soy wax candle in a frosted glass jar with a natural cotton wick'
Layer 3 - Colors & Materials
Hex codes mean nothing to an AI image generator. Broad color names like "blue" or "gold" are too generic to be useful. Describe temperature, undertone, and finish instead, and name the material along with how it behaves under light.
Without this
'gold color'
With this
'warm champagne gold with a subtle satin sheen'
Layer 4 - Composition
Composition tells the AI where to place the subject and how to frame the shot - subject position (centered, rule of thirds, off-center), framing (full body, close-up, three-quarter), and negative space. Without this direction, most AI tools default to a centered subject with medium framing every time.
Without this
(no composition direction given)
With this
'candle positioned on the left third, generous negative space to the right'
Layer 5 - Lighting
Lighting is the single most impactful layer in the framework and the most commonly skipped. Describe the source, its direction, and its quality. We cover this layer in full depth - including 30 ready-to-use phrases - in How to Describe Lighting in AI Prompts, so we'll keep it brief here.
Without this
'good lighting'
With this
'soft diffused window light from the left, warm golden undertones'
Layer 6 - Camera & Lens
The final layer locks in the photographic feel. Specify a focal length and aperture - an 85mm lens at f/2.0 for flattering compression and background blur, a 100mm macro lens for extreme material detail, a 24mm wide angle for environmental context. This is what separates an image that looks AI-generated from one that looks professionally photographed.
Without this
'realistic photo'
With this
'shot on 100mm macro lens at f/4.0'