The 6-Layer AI Prompt Framework: How to Structure Any AI Image Prompt

6-Layer AI Prompt Structure 1. Concept 2. Subject 3. Colors & Materials 4. Composition 5. Lighting 6. Camera & Lens

Key takeaways

  • Generic prompts fail because they skip layers - describing only the subject with no lighting, composition, or camera direction
  • Order matters - concept frames everything, and each following layer adds specificity on top of the last
  • The framework is platform-agnostic - it works identically for Midjourney, ChatGPT, Nano Banana, and Sora
  • Each layer should compress to roughly one sentence or phrase in the final prompt - not a paragraph
  • The best editing check is simple: could an 8-year-old visualize your prompt out loud?

Quick Answer

Structure any AI image prompt in 6 layers, in this order: concept (the content type), subject (full descriptive detail), colors and materials (specific language, not hex codes), composition (subject position and framing), lighting (source, direction, quality), and camera and lens (focal length and aperture). Each layer adds a level of specificity that generic prompts skip - which is exactly why they produce generic results.

In this guide

  1. What elements should be in an AI image prompt
  2. Why most AI image prompts fail
  3. The 6-layer AI prompt structure
  4. Putting all 6 layers together
  5. Prompt structure for Midjourney vs ChatGPT vs Sora
  6. The 8-year-old test
  7. Where this framework applies

Every AI image tool - Midjourney, ChatGPT, Nano Banana, Sora - is a creative translator, not a creative partner. It executes whatever vision you hand it. Hand it a vague sentence, you get a vague, generic image. Hand it a specific, layered description, and you get something that looks art-directed.

The 6-layer framework is the structure behind that difference. It is the same methodology whether you're prompting for a portrait, a product shot, or a lifestyle scene, and it works identically across every AI image tool on the market today.

What elements should be in an AI image prompt

A strong AI image prompt includes six elements, written in this order: concept, subject, colors and materials, composition, lighting, and camera and lens. This is the direct answer to "how do I write better AI image prompts" - most people only ever describe the subject, which is layer 2 of 6. The other five layers are what separate a professional-looking result from a flat, average one.

Think of it like a photographer building a scene before ever pressing the shutter. Concept sets the type of shot. Subject fills in who or what is in it. Colors and materials give it texture. Composition frames it. Lighting sets the mood. Camera and lens locks in the photographic feel. Skip any layer, and the AI fills the gap with a generic default.

Why most AI image prompts fail

Generic prompt in, generic result out. That is the entire failure pattern, and it happens for three specific reasons:

  • The prompt describes only the subject - "a woman in a red dress" - with no direction on composition, lighting, or camera.
  • The prompt relies on vague adjectives like "nice," "beautiful," or "professional," which carry no visual instruction an AI can act on.
  • The prompt skips descriptive visual language entirely. "Blue dress" tells an AI almost nothing. "Deep cerulean silk midi dress with cool purple undertones" tells it everything.

Every AI image tool defaults to the same safe, average output when given vague input - centered subject, flat even lighting, medium framing. The 6-layer framework exists to override those defaults, one layer at a time.

Generic prompt in, generic result out.

The 6-layer AI prompt structure

Here is each layer in the order you should write them.

Layer 1 - Concept

Concept defines what kind of image this is before you describe anything else - portrait, product still life, flat lay, lifestyle scene, editorial. This single decision frames how the AI interprets every layer that follows.

Without this

'a photo of a candle'

With this

'product still life of a candle'

Layer 2 - Subject

Now describe the main subject in full detail - material, color, texture, size relative to its surroundings, position, and any distinguishing detail. Every detail you skip is a detail the AI invents on its own.

Without this

'a nice candle'

With this

'a hand-poured soy wax candle in a frosted glass jar with a natural cotton wick'

Layer 3 - Colors & Materials

Hex codes mean nothing to an AI image generator. Broad color names like "blue" or "gold" are too generic to be useful. Describe temperature, undertone, and finish instead, and name the material along with how it behaves under light.

Without this

'gold color'

With this

'warm champagne gold with a subtle satin sheen'

Layer 4 - Composition

Composition tells the AI where to place the subject and how to frame the shot - subject position (centered, rule of thirds, off-center), framing (full body, close-up, three-quarter), and negative space. Without this direction, most AI tools default to a centered subject with medium framing every time.

Without this

(no composition direction given)

With this

'candle positioned on the left third, generous negative space to the right'

Layer 5 - Lighting

Lighting is the single most impactful layer in the framework and the most commonly skipped. Describe the source, its direction, and its quality. We cover this layer in full depth - including 30 ready-to-use phrases - in How to Describe Lighting in AI Prompts, so we'll keep it brief here.

Without this

'good lighting'

With this

'soft diffused window light from the left, warm golden undertones'

Layer 6 - Camera & Lens

The final layer locks in the photographic feel. Specify a focal length and aperture - an 85mm lens at f/2.0 for flattering compression and background blur, a 100mm macro lens for extreme material detail, a 24mm wide angle for environmental context. This is what separates an image that looks AI-generated from one that looks professionally photographed.

Without this

'realistic photo'

With this

'shot on 100mm macro lens at f/4.0'

Putting all 6 layers together

Here is a single prompt built one layer at a time, so you can see exactly how each layer stacks on top of the last.

Layer 1 - Concept

'Product still life of a candle'

+ Layer 2 - Subject

'Product still life of a hand-poured soy wax candle in a frosted glass jar with a natural cotton wick'

+ Layer 3 - Colors & Materials

'...in a frosted glass jar with warm champagne gold satin-finish lid, set against a soft cream linen surface'

+ Layer 4 - Composition

'...positioned on the left third of the frame, generous negative space to the right'

+ Layer 5 - Lighting

'...soft diffused natural window light from the left, warm golden undertones, gentle shadow falling to the right of the jar'

+ Layer 6 - Camera & Lens (final prompt)

'Product still life of a hand-poured soy wax candle in a frosted glass jar with warm champagne gold satin-finish lid, set against a soft cream linen surface, positioned on the left third of the frame with generous negative space to the right, soft diffused natural window light from the left with warm golden undertones and a gentle shadow falling to the right of the jar, shot on 100mm macro lens at f/4.0'

Notice the final prompt reads as a single flowing description, not six disconnected fragments. Each layer folds into the next.

Prompt structure for Midjourney vs ChatGPT vs Sora

Midjourney wants the 6 layers written as structured descriptive phrases separated by commas, with platform parameters appended at the end - --ar for aspect ratio, --v 8 for model version, --stylize for interpretation strength. See How to Write Midjourney Prompts That Actually Work for the full parameter breakdown and Midjourney-specific examples.

ChatGPT and Nano Banana prefer the same six layers delivered as natural, flowing sentences rather than comma-separated phrases. The information is identical - concept, subject, colors, composition, lighting, camera - just written the way you'd describe a scene to another person.

Sora and other video tools need everything above plus a 7th consideration: motion and duration. A still image prompt describes a frozen moment; a video prompt has to describe how that moment moves and for how long. See How to Write AI Video Prompts for the complete breakdown of that 7th layer.

The 8-year-old test

Before you run any prompt, read it out loud and ask: could an 8-year-old visualize this scene from your words alone? If there are gaps - a missing light source, an undescribed color, no sense of framing - that's exactly where to add more specific detail.

Could an 8-year-old visualize it? That's the test.

This isn't a trick question. It's a genuine editing checklist. Read your prompt, picture what a child with zero context would imagine, and compare that to what you actually want. The gap between those two things is your next edit.

Where this framework applies

The 6-layer framework isn't limited to any one use case. Here's where it shows up across the site:

Let a tool build all 6 layers for you.

The Midjourney Prompt Builder walks you through each layer with dropdowns and assembles the full prompt in real time, then hands you 3 AI-enhanced variations - subtle, standard, and cinematic.

Open the Prompt Builder

Frequently asked questions about the 6-layer AI prompt framework

Six elements, written in order: concept (the content type), subject (full descriptive detail), colors and materials (specific descriptive language, not hex codes), composition (subject position and framing), lighting (source, direction, and quality), and camera and lens (focal length and aperture). Each layer adds specificity the AI needs to produce a professional-looking result.

Yes - the 6-layer structure works as a repeatable formula: concept, subject, colors and materials, composition, lighting, camera and lens, written in that order. It applies whether you're generating a portrait, a product shot, or a lifestyle scene.

Yes. The 6 layers themselves are platform-agnostic - only the delivery format changes. Midjourney appends -- parameters after the phrase-based prompt, ChatGPT and Nano Banana prefer the same information delivered as a natural sentence, and Sora or other video tools need a 7th consideration for motion and duration.

Generic results almost always come from skipping layers - describing only the subject with no composition, lighting, or camera direction, or using vague adjectives like "nice" instead of specific visual language. See our full breakdown in Why Your AI Images Look Generic for all 8 specific causes and fixes.