Mastered knowledge · Stable Diffusion 1.5
Cinematic POV prompts in Stable Diffusion 1.5: the complete guide
First-person shots are the fastest way to make an SD1.5 image feel cinematic — and the fastest way to get warped geometry, six-fingered hands and a fisheye lens you never asked for. This guide is the distilled prompt logic that fixes it: keyword mechanics, the categorized positive/negative dictionary, the 6-step hierarchy and the LoRA weights that hold it together.
Why SD1.5 behaves this way
- SD1.5 is a latent text-to-image diffusion model: CLIP text encoder -> U-Net denoising in latent space -> VAE decode to pixels.
- The CLIP text encoder (ViT-L/14) has a hard 77-token limit; anything beyond is silently truncated. Concise prompts beat verbose 'word salads'.
- SD1.5 has no innate concept of perspective or anatomy; it activates statistical associations between words and visual features learned from training data.
- A keyword only works if it was strongly correlated with the desired visual outcome in the training data (e.g. 'POV' works because such images were captioned).
- The negative prompt is processed like a positive prompt; CFG pushes the image away from the negative description.
The practical consequence: prompt engineering here is not giving orders, it is activating the right learned associations — and staying inside the 77-token CLIP budget.
The POV keywords — and their side effects
These terms are the primary directive that tells the model to build the scene from a first-person standpoint:
POV · point of view · first person view · first person perspective · FPV
For Danbooru-tag anime models, add the perspective tags directly:
POV HANDS · POV ARMS · 1st person view · view from behind
The side effects you must counter
A bare POV token drags these artifacts in with it. Every one of them has a negative-prompt countermeasure:
| Side effect | Countermeasure |
|---|---|
| wide-angle / fisheye lens distortion | wide angle lens, fisheye lens, distorted perspective, warped |
| stretched or warped geometry | wide angle lens, fisheye lens, distorted perspective, warped |
| disproportionately large hands | wide angle lens, fisheye lens, distorted perspective, warped |
The categorized positive / negative dictionary
Six groups. Use the positive column to build the scene, the negative column to fence it.
| Group | Positive | Negative |
|---|---|---|
| Perspective & View | POV, point of view, first person view, eye-level shot, subjective camera, immersive, camera locked at eye level | (full body shot), (character's body), (viewer's hands/arms) |
| Anatomy & Foreground | pov hands, hands in frame, arms reaching towards camera, pov hands grabbing, visible hands, extended arm, hand holding [object] | extra limbs, floating limbs, missing limb, bad anatomy, malformed hands, long neck, long body |
| Depth & Focus | shallow depth of field, cinematic bokeh, out of focus background, sharp focus, depth of field, vignette | deep depth of field, inaccurate skin color, muddy background |
| Composition & Framing | close-up, medium shot, cowboy shot, low-angle shot, high-angle shot, over-the-shoulder view | cropped head, cut off, out of frame |
| Lighting & Cinematics | cinematic lighting, dramatic lighting, moody atmosphere, low-key lighting, volumetric lighting, film grain, high budget | flat lighting, harsh shadows, overexposed, underexposed, noisy |
| Realism & Quality | hyper-detailed, ultra-realistic, photorealistic, 8k resolution, best quality, masterpiece, intricate details | drawing, painting, sketch, impressionist, blurry, ugly, deformed, text, watermark |
The 6-step prompt hierarchy
Assemble broad → specific. This is the order that keeps the important tokens inside the CLIP budget before truncation kicks in.
- 1. Quality & Style Foundation
- 2. Subject & Action
- 3. POV & Composition
- 4. Foreground Elements
- 5. Background & Environment
- 6. Lighting & Mood
A worked example
Built from the hierarchy above — a "handing a glass of water" POV scene.
Positive
cinematic still frame, hyper-detailed, ultra-realistic, 8k resolution, sharp focus, masterpiece, a woman handing a glass of water to the viewer, POV shot, eye-level close-up, shallow depth of field, her right arm is extended towards the viewer, perfect 5-fingered hands visible, softly blurred living room background with a window, natural window lighting, volumetric sun rays, subtle film grain Negative
bad anatomy, disfigured, extra limbs, mutated hands and fingers, poorly drawn hands, deformed, ugly, cloned face, duplicate, text, floating limbs, missing limb, malformed hands, long neck, long body, extra fingers, deformed fingers, wide angle lens, fisheye lens, distorted perspective, warped, cropped head, cut off, out of frame, (full body shot), (character's body), flat lighting, harsh shadows, overexposed, underexposed, noisy, deep depth of field, inaccurate skin color, muddy background Hands and anatomy: what actually shows up
| Goal | Positive tokens | Negative tokens |
|---|---|---|
| Hands in frame | pov hands, hands in frame, arms reaching towards camera, arm extended towards the viewer, hand holding {object}, perfect 5-fingered hands, photorealistic hands, manicured nails, visible hands, extended arm | extra limbs, floating limbs, missing limb, bad anatomy, malformed hands, mutated hands and fingers, poorly drawn hands, long neck, long body, extra fingers, deformed fingers |
| Depth & focus | shallow depth of field, cinematic bokeh, out of focus background, sharp focus, depth of field, vignette | deep depth of field, inaccurate skin color, muddy background |
| Framing | close-up framing, medium shot framing, cowboy shot framing, low-angle shot, high-angle shot, over-the-shoulder view, eye-level shot, camera locked at eye level | cropped head, cut off, out of frame, (full body shot), (character's body) |
Lighting & cinematic vocabulary
Mood Tokens dramatic dramatic lighting, moody atmosphere, low-key lighting natural natural window lighting, practical lights only, soft shadows volumetric volumetric lighting, god rays color teal shadows, warm highlights, cinematic color grading cinematic cinematic lighting, high budget, film grain, halation
Useful negatives: flat lighting, harsh shadows, overexposed, underexposed, noisy. Baseline
quality/realism tokens: cinematic still frame, hyper-detailed, ultra-realistic, photorealistic, 8k resolution, sharp focus, best quality, masterpiece, intricate details, film grain, halation, vignette, high ISO.
Negative-prompt packs (concise beats verbose)
generic
bad anatomy, disfigured, extra limbs, mutated hands and fingers, poorly drawn hands, deformed, ugly, cloned face, duplicate, text
anime
lowres, bad anatomy, bad hands, error, missing fingers, extra digit, fewer digits, cropped, worst quality, low quality, jpeg artifacts, signature, watermark, username, blurry
photoreal
cartoon, 3d render, painting, drawing, illustration, anime, airbrushed, plastic skin, overexposed, oversaturated, jpeg artifacts
pov
wide angle lens, fisheye lens, distorted perspective, warped, full body shot, (character's body), extra limbs, malformed hands, cropped head, cut off
Overloading the negative prompt degrades quality. Start focused, test, then add
only the terms that fix a defect you actually observed.
Weighting and LoRA syntax
Weighted tokens (A1111 / SD.Next compatible)
Syntax Meaning (token) 1.1x emphasis (token:1.3) explicit emphasis weight [token] 0.9x de-emphasis [token:0.8] explicit de-emphasis weight (token:1.1:0.5) emphasis + start/stop blend ratio token1 AND token2 per-token conditional switching (wildcard-like) token1 | token2 random choice between alternatives BREAK shift attention to what follows (A1111) token1 token2 plain adjacency (default, usually best) underscored_tag underscores become spaces; used in Danbooru style {{token}} nested emphasis (multiplies weight) [from:to:step] schedule: swap 'from' for 'to' at step N
Guidance: emphasis 1.1 - 1.2 (subtle boost, generally enough); strong 1.3 - 1.5 (use sparingly; >1.5 often degrades composition); de-emphasis 0.8 - 0.95 (suppress a term without removing it).
LoRA
<lora:lora_name:weight> — examples: <lora:add_detail:0.6> · <lora:char_style_v1:0.8> · <lora:cinematic_light:0.5>, <lora:char_style_v1:0.7>
Item Guidance weight range 0.4 - 1.0 is the practical band; start at 0.6-0.8 style LoRAs 0.4 - 0.7 (prevent style bleed / over-saturation) character / concept LoRAs 0.7 - 1.0 (raise until identity holds) detail / quality LoRAs 0.3 - 0.6 (add_detail, detail_slider) trigger words append the LoRA's documented trigger token(s) to the positive prompt stacking 2-3 LoRAs max; keep total weight sum <= ~1.5 to avoid collapse negative prompt add model-specific negatives (e.g. 'extra digits') when a LoRA degrades anatomy clip skip style LoRAs often prefer Clip skip 2 CFG interplay high CFG (>=9) magnifies LoRA artifacts; prefer 6-8 with LoRAs
- Check the LoRA card for its exact trigger token; without it the LoRA is often inert.
- Put the trigger token near the subject tokens, not at the very end.
- If the LoRA has no trigger word, rely on the training subject name in natural language.
Other syntax: embedding (embedding:embedding_name:1.0) · hypernet <hypernet:hypernet_name:1.0>. Keep LoRA and checkpoint names generic — these
templates work in A1111, ComfyUI and SD.Next.
Generation parameters that hold the result together
Parameter Proven setting cfg 4-8 typical; 7 is the proven all-round default; >10 risks oversaturation & artifacts steps 20-30 typical; 28 is a proven default; hires pass: 12-20 extra steps samplers DPM++ 2M Karras (default), Euler a (artistic), DDIM (fast), DPM++ SDE Karras (fine detail) seed fix the seed to reproduce; random seeds to explore hires fix upscale 1.5-2x at 0.4-0.6 denoising strength after the base pass resolution native 512x512/512x768/768x512 for SD1.5; use hires fix for larger outputs clip skip 2 for many anime/style models; 1 for photorealistic checkpoints negative prompt weight keep negative focused and concise; overloading it degrades quality vae use the checkpoint's matching VAE (or a universal fp16 fix VAE) to avoid grey images
Aspect ratios: 1:1 → 512x512 · 2:3 portrait → 512x768 · 3:2 landscape → 768x512 · 9:16 vertical → 512x896 (via hires fix) · 16:9 wide → 896x512 (via hires fix)
POV prompt FAQ
Why do POV prompts in Stable Diffusion 1.5 produce distorted images?
Bare "POV" activates wide-angle and fisheye associations from the training captions, so the model stretches geometry and enlarges the hands. The fix is not to remove POV but to counter it: keep the keyword and add wide angle lens, fisheye lens, distorted perspective and warped to the negative prompt.
How do I get hands into a POV prompt without getting a full body?
Anchor the frame in the viewer position with phrases like "pov hands", "hands in frame" and "her right arm is extended towards the viewer", then exclude the rest of the character with "(full body shot)" and "(character's body)" in the negative prompt.
What order should a Stable Diffusion 1.5 prompt be written in?
Broad to specific: quality and style foundation, subject and action, POV and composition, foreground elements, background and environment, then lighting and mood. SD1.5 has a hard 77-token CLIP limit, so anything past that is silently truncated — put what matters first.
What LoRA weight should I use for a style LoRA?
Style LoRAs work in the 0.4–0.7 band to prevent style bleed and oversaturation, character or concept LoRAs in the 0.7–1.0 band, and detail LoRAs at 0.3–0.6. Stack two to three LoRAs at most, keep the total weight sum around 1.5, and prefer CFG 6–8 when LoRAs are loaded.