
AI Prompts That Work the First Time: A Practical Guide
A prompt is a text brief for a model. And almost every time a generation comes out nothing like the picture in your head, the model is not the problem. The brief was vague.
Below is a working method: how to build an image prompt, how to rewrite it for video, what to do with negatives, and how to avoid burning your whole generation budget in one evening.
Why it fails on the first try
Three reasons come up over and over.
- The model cannot tell what matters. Give it twenty equally weighted details and it will pick the lead itself. Probably not the one you wanted.
- The prompt contradicts itself. «Soft diffused light» and «hard contrasty shadows» in the same line is an argument the model settles at random.
- Everything changes at once. You dislike the shot, so you adjust light, angle and wardrobe in one go. It gets worse. Which change did it? No idea.
A prompt is not a spell or a pile of tags. It is a short brief read by someone who cannot ask you a follow-up question.
How an image prompt is built
Order matters: most models give more weight to what comes first. This sequence works.
- Subject. Who or what is in frame, as one clear noun with a couple of qualifiers.
- Action or pose. What they are doing right now. One action, not three.
- Environment. Where it happens and at what time of day.
- Light. One source and its character: soft window light from the left, hard backlight, warm tungsten.
- Optics and framing. Shot size, focal length, depth of field.
- Style and mood. Palette, texture, emotion.
- Technical parameters. Aspect ratio and other flags, always last.
Recent models such as Google's Nano Banana Pro read plain human sentences and do not need tag spam like «8k, masterpiece, trending on artstation». Write in full sentences, the way you would brief a photographer. We covered that model in detail in the Nano Banana Pro guide.
Midjourney works on similar logic: describe the scene, then append parameters at the end of the line for aspect ratio, stylization strength and exclusions. The exact flag syntax shifts between versions, so check the official docs rather than a year-old blog post.
Example 1: a shot from scratch
Close-up portrait of a middle-aged barista in a small city coffee shop, holding a warm ceramic cup with both hands, soft window light from the left, shallow depth of field, warm amber and charcoal palette, 50mm lens, subtle film grain, calm and slightly tired expression.
Why it works: one subject, one action, one light source, one palette, one lens. Mood sits at the end, not the front. Nothing in the line contradicts anything else. There is simply no room left for the model to improvise.
Example 2: editing an existing image
Replace the plain grey background with a softly blurred workshop interior at dusk. Keep the person, their pose, clothing, facial features and the direction of light exactly as they are. Match the new background colour temperature to the existing warm key light.
Why it works: it contains the part people forget, the preservation clause. «Change the background» gives the model permission to redo everything. «Change the background, keep the face, pose, clothing and light direction unchanged» narrows the job to a single operation. When editing, stick to one change per pass.

A video prompt follows different logic
The classic beginner mistake is taking a good image prompt and handing it to a video model. It will not work, because a video model is not asking what is in the frame. It is asking what changes.
This matters most in image-to-video, where you hand the model a finished still. Appearance, lighting and composition are already locked by the picture. Repeat them in text and you create a second set of instructions that argues with the first, and the character drifts. In image-to-video, describe motion only: what the subject does, what the camera does, and what stays still.
Kling recommends its own five-part structure: subject, subject motion, scene, camera language, and lighting with atmosphere. ByteDance describes a similar six-step formula for Seedance: subject, action, environment, camera, style, constraints, aiming at roughly 60 to 100 words. MiniMax H3 generates picture and sound in a single pass, which adds an audio layer to the prompt: lines in quotes, action sounds, ambience and music. We break that model down in the MiniMax H3 guide, and compare the models against each other in our video model comparison.
Example 3: animating a finished still
Slow push-in on the barista. She lifts the cup, takes one sip, then lowers it. Steam drifts upward. The background stays soft and static. Handheld micro-drift, no cuts.
Why it works: one camera move instead of four, one meaningful action with a beginning and an end, and an explicit statement of what must stay still. Not a word about appearance or lighting, because the source image already carries both.
If the model handles sound, add the audio layer as its own line: «quiet cafe ambience, one ceramic clink, no dialogue, no music». Name the sounds you want, or the mix is left to chance.
What to do with negatives
A negative prompt lists what must not appear. Some services give it a dedicated field, others take it as a flag inside the line. It works best on artefacts, not on meaning.
- Good at removing: blur, distortion, extra fingers and limbs, watermarks, stray text, jitter.
- Bad at removing: things you never asked for. Writing «no cars» when there were no cars in the prompt only puts the word «cars» in front of the model.
- Not a styling tool. A negative will not make an image prettier, it only trims defects.
The base set for video tends to be the same every time: blur, distortion, warped fingers, extra limbs, text, watermark. Save it once and paste it into every generation.
The one-variable rule
This is the most underrated habit in the whole workflow. Change one thing per pass. Light. Or angle. Or palette. Never all three.
The reason is simple: generation has a random component. Change three parameters, get a better result, and you still do not know which one helped, so you cannot repeat it. Move one at a time and five iterations leave you with five known-good settings you can apply to any shot afterwards.
Practical tip: keep prompts in a text file rather than in the service input box. You keep the edit history and can roll back to a version that worked.
How to stop wasting money on generations
Video generation costs far more than an image and takes longer to render. That dictates the order of work.
- First get the still frame right in an image model: composition, subject, light, palette.
- Only then feed that frame into a video model in image-to-video mode.
- In the video prompt, describe motion and camera only.
- If the motion misses, fix the motion. You never have to revisit the image.
This way you need far fewer expensive attempts, and results are steadier because the video model is not reinventing the character on every run. The full pipeline is in how to make an AI video, and free quotas are available in Kling, among others.
Common mistakes
- Tag lists instead of description. Current models read connected prose better than a garland of commas.
- Junk boosters. «8k, ultra detailed, masterpiece, award winning» add almost nothing and steal the model's attention.
- Two conflicting requirements. Soft light and hard shadows, wide shot and tight portrait, locked-off camera and an orbit.
- Three camera moves in one short clip. A push-in, an orbit and a pan inside five seconds is guaranteed mush.
- Describing appearance in image-to-video. It duplicates what the frame already holds and causes face drift.
- Ten edits in one sentence. Especially painful in image editing, where you cannot tell which edit broke the result.
- No preservation clause. Without «keep the face and pose unchanged», the model assumes it may redo everything.
- Length for its own sake. A 300-word prompt usually performs worse than a dense 60 to 100.
Short checklist
- One subject, one action, one light source.
- Most important first, technical parameters last.
- For video, describe what changes, not what is there.
- Keep negatives to artefacts only.
- One edit per iteration.
- Still frame first, motion second.
If you would rather learn this systematically than in fragments, we run AI video training built around your own projects.
Frequently asked questions
What language should I write prompts in?
Large models understand many languages, but they were trained mostly on English descriptions, so English tends to give more predictable results, especially for optics and lighting terms. A practical approach is to think in your own language and translate the final prompt, double-checking the technical vocabulary.
How long should a prompt be?
For video, roughly 60 to 100 words is a good target. Images allow a little more room, but the principle holds: a dense paragraph beats a wall of text. Every extra requirement dilutes the weight of the rest.
Do I always need a negative prompt?
No. It earns its place where a model repeatedly produces the same defect: warped fingers, jitter, watermarks. With no specific problem to solve, a long negative will not improve the shot and can get in the way.
Why does the same prompt give different results?
Generation has a random component, so two runs of identical text produce different frames. That is exactly why you change one parameter at a time: otherwise you cannot separate «got lucky» from «the prompt got better».
Can I reuse a prompt across different models?
The meaningful part travels fine: subject, action, light, camera. Technical flags and service-specific syntax do not, and have to be rewritten. An image prompt does not transfer to a video model at all, since the two are answering different questions.
Where do I start when the output looks nothing like the idea?
Cut the prompt down to three lines: who, what they are doing, what the light is. Get an acceptable base, then add details one at a time. Repairing an overloaded prompt is almost always slower than rebuilding it from a short core. You can practise for free in Seedance, among other options.
Need an AI system or a video for your business?
Describe your case — we will come back with a proposal and an estimate within a day.
Discuss your case