Prompt Like THIS to Master Making AI Videos (5 Levels)
At a glance
- Length
- 22 min
- Channel
- Youri van Hofwegen
- Video from
- Aug 2026
- Rating
- ⭐⭐ Great video · 2/2
- Best for
- AI video creators seeking repeatable, structured prompting workflows
What this video answers
- What is the difference between each level of prompting?
- Do I need to use Higgsfield specifically to apply these techniques?
- Can beginners jump straight to the advanced levels?
- How does character consistency actually work in practice?
- What's included in the Cinema Studio controls section?
Understanding AI Video Prompting Across Five Levels
Youri van Hofwegen's tutorial breaks down the craft of prompting for AI video generation into a structured five-level framework, using Higgsfield as the primary tool. Rather than treating AI video creation as a one-step process, the video demonstrates how deliberate prompt engineering can transform raw output into predictable, cinematic results. By working through the same scenes and settings at each level, van Hofwegen shows viewers exactly how much control they gain as prompting sophistication increases.
The core insight is that AI video quality isn't just about having a good idea—it's about communicating that idea with precision. The progression from simple prompts to structured multi-shot workflows reveals that beginners often underestimate how much detail and structure can influence the final output. This approach appeals equally to creators experimenting for the first time and to those already familiar with AI tools but struggling to achieve consistent, intentional results.
Key Strengths in This Five-Level Prompting Framework
- Progressive complexity: Each level builds logically on the last, so viewers can apply techniques immediately rather than waiting for the "expert" section
- Practical consistency: Using identical scenes across all five levels lets viewers isolate exactly what each prompting strategy changes, removing guesswork
- Character control via Elements: The tutorial addresses a common pain point—how to keep characters consistent across multiple shots—with a dedicated tool demonstration
- Cinema Studio controls: Advanced cinematic control options are shown in context, making professional-grade features accessible rather than intimidating
- Workflow structure: Multi-shot prompt templates replace ad-hoc guessing, giving creators a replicable system rather than isolated techniques
- Free resource included: Access to a prompt generator removes the barrier of starting from scratch and provides a safety net for new users

Who Benefits Most From This Prompting Guide
This tutorial suits creators at any experience level with AI video tools, but especially those frustrated by inconsistent results or unsure how to scale from simple clips to full projects. If you've generated a few AI videos but felt like you were relying on luck rather than skill, this framework fills that gap. Video editors, content creators, and marketing professionals who want faster, more predictable AI workflows without learning to code will find immediate value.
The verdict: worth watching if you're serious about using AI for video, whether you're just starting or refining existing skills. It's not for casual experimenters who only need a single video, but anyone building a repeatable process should invest the time.
Frequently Asked Questions About AI Video Prompting Levels
What is the difference between each level of prompting?
The video progresses from basic one-line prompts (minimal control) through increasingly detailed and structured approaches, ultimately reaching multi-shot workflows with character consistency and cinematic controls. Each level gives you more predictability and intentionality over the output.
Do I need to use Higgsfield specifically to apply these techniques?
The tutorial uses Higgsfield as the demonstration platform, so the exact interface will differ if you use another AI video tool. However, the prompting principles—specificity, structure, character consistency—transfer across most AI video generators, even if the controls are named differently.
Can beginners jump straight to the advanced levels?
The tutorial is designed sequentially, so you'll understand concepts better by starting at level one. That said, if you're already comfortable with AI tools, you can scan the early levels and focus on techniques you haven't tried. The progression is logical but not rigid.
How does character consistency actually work in practice?
The video shows how to use Elements within Higgsfield to maintain character appearance across multiple shots. Rather than describing each character repeatedly in every prompt, you reference a saved element, reducing errors and saving time on longer projects.
What's included in the Cinema Studio controls section?
Cinema Studio features are shown in the context of applying cinematic techniques to your AI videos—camera movement, framing, and visual refinement that make output look more intentional and professional rather than obviously AI-generated.

Key Terms
- Prompting
- Writing detailed instructions to guide an AI tool toward a specific creative output rather than accepting default results.
- Elements
- A feature that saves and references character or visual details across multiple shots so they remain consistent without rewriting descriptions each time.
- Cinema Studio
- Advanced controls within Higgsfield for adding cinematic techniques like camera movement and framing to AI-generated video.
- Multi-shot workflow
- A structured approach to creating videos with multiple scenes or clips, using consistent prompting rules and character references across all segments.
- Higgsfield
- An AI video generation platform that supports detailed prompting, character consistency, and cinematic controls for creating longer, more intentional videos.
Sources: Prompting · Elements · Cinema Studio · Multi-shot workflow · Higgsfield — definitions cross-referenced with Wikipedia
Video by Youri van Hofwegen on YouTube. If you enjoyed it, please subscribe to their channel and show your support for the great video.
Description
Create your own AI Videos with Higgsfield 👉 https://youricreates.com/Higgsfield
In this video, I walk through the five levels of AI video prompting, showing how each step gives you more control over your results using the same scenes, settings, and model inside Higgsfield. I cover everything from simple one-line prompts to structured multi-shot workflows, character consistency with Elements, and Cinema Studio controls so you can create more cinematic AI videos with predictable results.
Get The 5 Levels of AI Video Prompting 👉 https://youri-van-hofwegen.kit.com/92eb5a883f
✅ Generate top tier AI video prompts for free: https://www.videoprompt.studio/
for inquiries: yvh (at) youripartnerships.com
Video transcript Accessibility
A full written transcript of this video, provided for accessibility. Select any timestamp to jump the video to that moment.
You're probably wondering why your AI videos keep coming out mediocre, even when you're using all the best models. And the biggest thing I've learned is that the reason you're not creating cinematic outputs, no matter what tool you use, always comes down to how you prompt. So, after working with these
tools for months, I've been able to come up with five levels of prompting. And in this video, I'll walk you through all of them, so you can finally go from a beginner to a pro who's mastered AI video creation. We're starting at level one, which is exactly where most people are. And it's important to go through
this stage to show you what works and what doesn't. For this whole video, I'll be working inside Hexfield, which is an all-in-one platform that keeps every AI tool I need in one place. So, if you want to follow along, I've left a link for it in the description below. Once you log in, everything's laid out along
the top navigation [music] bar, and since we're making video, you just click on video to open the video generation workspace. From the model selector, I'll pick Seed N 2.0, because right now, it's one of [music] the best video models out there. And it's the one model we'll use throughout this video, so the only
difference will come down to the prompt. And to test this properly, I'll run the exact same ideas through all five levels with the same settings. So, I'll set the duration to eight seconds, the resolution to 1080p, the aspect ratio to 16x9, and turn the sound on. Level one prompting is just [music] describing
your idea with a simple sentence. So, for the first idea, I'll write a boxer walking to the ring and generate. And this exact prompt, along with every other level, is in the description completely for free, so you can follow along and try each level yourself. >> [cheering] >> And honestly, for a few words, it looks
good. A fighter is walking down a tunnel with the ring glowing at the end, and it even added its own crowd noise underneath. So, I'll keep every setting the same and just swap the prompt for a busy Tokyo street at night. >> [music] >> And again, it looks the part. Wet neon streets with the signs stacked
everywhere. Then, I'll change it to a car drifting around the corner. And I get a car sliding through a bend with smoke pouring off the tires. And the last one, a woman making coffee in her kitchen, comes out as this calm [music] warm morning scene. >> [music]
>> None of these look bad on their own. SeeDance 2 [music] is good enough where it can understand what a scene needs when you describe something and it fills the gaps. But, that's also what you're sacrificing. With a lazy prompt like that, you hand over all the control to the model. So, the shots, the camera
angle, the environment, and the characters will always end up looking different every time you generate. So, if I run each prompt two more times with the exact same words, you'll notice that they look like completely different videos. The boxer's a different guy each time in a different tunnel with a
completely different feel. The Tokyo one does something different with the camera every run, whether it's speeding through the city, walking along with the people, or coming in from above. And the city itself looks different, too. The car comes out with a different design every time, in a different setting, and even
the time of day changes. And the coffee one, the order she does things is completely different. [music] And between the three videos, these outputs have completely different shots. So, if you've ever made a video that worked for you and tried to build upon it, but everything kept falling apart, this is
exactly why. And that's level one. You hand over a rough idea and the model actually designs everything for you. So, you just end up burning through your credits on videos you didn't actually control, which is fine when you just want something quick. But, the second you want to work on an actual project, a
one-liner can't do it. And the first fix is what you actually say to the model. So, level two is where you stop giving it a sentence and actually start describing the shot. And there are five things you add here: the subject, the action, the setting, the lighting, and the mood. So, instead of a boxer, I'll
describe who he actually is, a lean welterweight with cropped black hair and worn oxblood gloves. And instead of just walking to the ring, I'll say what he does the whole way, moving down a narrow concrete tunnel with the ring light growing at the end. Then, I'll name where the light comes from, the caged
fluorescents overhead, instead of leaving it up to the model. And the mood is the one everyone gets wrong, because if you just write the word tense, you get nothing. The model can't turn a word like that into anything on screen. So, instead of naming the mood, I show it. His jaw is set, his breathing is short
through his nose, his gloves tap together as he walks. And that's what actually comes across as tension. [music] So, I'll paste that in and generate. And right away, it's completely different from level one. I get exactly the fighter I described, in the tunnel I described, with the light coming from
where I put it. So, I'll keep the settings the same and do the same for the others, describing the Tokyo street down to the wet asphalt and the neon, and the car down to the model, and the smoke coming off the tires.
And now, every one of them sticks to what I wrote, the subject, the setting, and the lighting, so it feels a lot more intentional. And for the coffee one, I'll go even further and describe her actual movements step by step. She lifts the jar and tips the grounds
into the basket, levels them off with her thumb, locks it into the machine, and presses the switch, all in that exact order. But, we've already established that for a one-off video. Seedance 2 is already pretty good, even with lazy prompts. So, obviously with these, it would work better. The real
test, though, is running them again to see what actually holds up. Now, the boxer stays right, the same fighter in the same tunnel, but the camera doesn't hold at all. One run stays behind him the whole way, and the next swings around to his face. And the shots, the angle, and the movement change every
single time. The Tokyo street and the car do the exact same thing. They're a lot more similar than before with lots of elements actually holding up, but the camera angle changes between videos, and the model controls it for me. But, the coffee is the exception, and this is the part that matters, because I described
her steps so tightly that it comes back almost the same on every single run. And that's the whole lesson of level 2 right there. Describe it loosely, and the model picks the camera for you. Describe it tightly like the coffee video, and your description leaves it almost nothing to decide. But, either way, you
still never actually control the camera yourself. And once you start doing that, you're beginning to work a lot more like a director rather than someone who randomly generates videos. And that's exactly what level 3 is about. You're at the stage where you're beginning to understand exactly what the generator
needs and try taking control of those elements, like the camera control. And that comes down to four things: the shot type, the angle, the movement, and the lens. The shot type is how much you actually fit in the frame. The angle is where you put the camera and how the subject comes across. For example, a low
one makes it look powerful, and a high one makes it look small. The movement is whether the camera pushes in, tracks along, or just stays still, which is exactly what the model kept making up on its own back in level 1. And the lens is the one almost nobody writes, where a wide one stretches the space out, and a
long one compresses everything flat. But, the important part is that the moment you start directing the camera, you have to structure the whole prompt around it. And every shot now locks to a single continuous take. So, let me go through each one. For the boxer, I'll describe the camera as a medium shot
with a wide lens, low, just below his shoulder, traveling alongside him as the ring light rises up his face. >> [snorts] >> That low angle is doing all the work because it actually makes him feel intimidating, like a real fighter. And for the first time, it looks like a real shot from a film instead of a random
generation. The Tokyo one, I'll rewrite completely because now the camera pushes slowly forward down the middle of the street. So, I describe everything in the order the camera actually reaches it. From the wet asphalt right in front to the neon signs stacking up further ahead. And it comes back gliding through the
street exactly the way I laid it out. And that's the whole point. The camera isn't something you add on at the end. It changes how you write the entire prompt. The car is all about the lens, so I put it on a wide shot with a long lens, low at road level, locked off on the outside of the corner.
That long lens is the thing hardly anyone thinks to write, and it completely changes the shot. It compresses the depth and stacks the tire smoke into a solid wall against the trees. And that compression you're seeing is nothing but the lens. It's the difference between a flat video and
something that looks shot on a real camera. And for the coffee one, I'll name a locked wide from across the kitchen with all the things on the counter sitting between the camera and her. And on the first go, it actually pulls back and gives me the width I asked for.
So, now I run each of these again to see how much of the camera really held. And here's the catch. Since you're just describing what you want, it's a request, not a setting. Some things will hold while others won't. The lens is the hardest one to pull off. For example, with the second generation of the car
video, it's a lot closer in and it doesn't give that zoomed out feeling the first one did. Or in the second boxer video, the camera angle is now slightly above his shoulders, not below. The coffee one is still the most accurate one and it maintains that locked wide shot in both videos. This is the level
where you start understanding the core values of prompting and it's the start of being advanced. But that gap between the shot you asked for and the shot you got back is the whole ceiling of this stage. But there's a way to actually control [music] the camera a lot better, as well as create multi-shot videos from
just one prompt, which is what Seed ends 2 is made for, instead of being locked in a single shot. And here we're starting to reach the pro levels where we control every single shot inside the prompt. So, you go from directing a single scene to directing a whole sequence in one generation. And there's
a lot of new things at this level. The biggest one is the structure itself. Instead of one flowing paragraph, you break the prompt into named blocks, the subject, the location, the light, the camera, the movement, and the audio, each on its own line, so the model never has to guess how they fit together. Then
there's timestamps where you stop describing what happens and start saying exactly when it happens, second by second, and that is the single biggest jump in this whole video. I also mentioned before that you get multiple shots inside one prompt. Now, with intentional cuts that you place yourself
instead of the random ones the model threw in back at level one. Then there are the locks, a subject lock that pins down exactly who your character is, and cross frame rules that spell out what has to stay the same across every cut. [music] And finally, the negative prompt, which is just a list of
everything you want the model to avoid. And most people skip this completely. And since these are full sequences now, instead of single shots, I'll bump the duration up so there's room for every cut. So, let me go through each one. For the boxer, I'll build the whole thing as a shot list with three shots on one
continuous walk and the location change on the final cut where he enters the arena. >> [panting] >> And when it generates, it actually pulls it off. The cut lands exactly where I timed it and the whole location switches from the tunnel to the arena on that one frame, which is something you could
never do with a single paragraph. The Tokyo one is all about timing dialogue. So, I'll give the two businessmen a line each in Japanese. And I timed them down to the exact second. One speaking around 7 seconds in and the other answering right after. So, it feels like a real conversation. I'll also lock their
positions so one stays on the left of the frame and the other on the right the whole way through. And it comes back with both of them talking right on the timing that I set and holding their spots across every cut, which is the kind of control you just can't get by describing it loosely.
The car is where I'll get specific with the speed. I'll write a ramp into the middle of the shot so the second it breaks traction, the whole thing drops into slow motion and then snaps back to full speed for that cinematic effect. I'll also specify that the audio should follow the car so the engine note
stretches out low through the slow part and goes back up when the speed returns. Woo! That was pretty close. >> And that's exactly what I get. The slow motion hits right on the slide and the sound stretches right along with it, which is what really sells it. And the coffee one is the most complicated yet
because I'm building five shots in 15 seconds. The main anchor will be the toaster, which will work like a countdown. The bread goes in at zero seconds and pops back out at 15. I'll also specify that I want the last cut to focus on the toast instead of the woman.
And it nails it. All five cuts are timed correctly, and the little radio in the kitchen keeps playing the same song straight through every single cut without ever restarting. Every one of these does exactly what I structured it to do. The cuts, the timing, the dialogue, the slow motion, all of it
lands perfectly. [music] And if we test them and run them again, you'll notice that they're the closest ones yet. If you actually structure your prompts [music] like this, you direct every shot so it's hard for the model to come up with something totally different. But you're probably thinking
that learning to write your prompts like that is a totally new skill set, and it's going to take hours to learn. The good news is you don't [music] have to. That's because there's a tool that does all of this structuring for you. It's called video prompt.studio, and I've left a link for it in the description
below. The whole thing is basically three steps. First, you paste in your base prompt, which is just a plain description of what you want. But here's the part people miss. The better your base prompt is, the better the result you get back. It will work with one sentence bases, but it'll be even better
if you explain what you want in each scene. So I'll break my idea into three shots based on the Tokyo scene from before. Then it hands me back the full structured version with every single field filled in. The subject, the camera, the lighting, all of it laid out, so you don't have to do anything
manually. But as good as it is, it's not perfect. What this tool does is structure your base prompt. It doesn't come up with parts that you haven't mentioned, so things like the dialogue or the position locks aren't present. So, if you just take what it gives you and generate, you're right back at level
one where you're letting the model fill in everything you left out. That's why you still have to go in and fix those parts yourself, which is what keeps you the one in control. The tool gives you a huge head start and it saves you hours of typing, but it doesn't do the thinking for you. So, I'll add my
dialogue back in, tighten up the continuity between the shots, and finish the prompt straight in Scene dance. And it comes back even better than before with a structure similar to what I wrote by hand, except this time I got there in a fraction of the effort. This
tool will make your workflow a lot faster while still keeping your videos looking professional. But, there is one main issue at this level, and that is character consistency. Because even with the cleanest structured prompt, you can't actually control how a character looks. So, the moment the video cuts,
their appearance changes. It's most obvious on the boxer where his face shifts a little every single time the shot cuts. So, it's still never quite the same guy twice. And taking control of that one last thing so that you direct the entire project is what the final level is all about. This is how
the pros actually prompt their AI videos. And we'll start with locking in our characters so that they look the same in every single shot, no matter how many times the video cuts or we regenerate. And the reason character consistency kept breaking comes down to one thing. Every level up to now, you've
been describing a character type and the model just built someone who fits that description, which ends up being a slightly different person every single time. Level five is where you stop giving it a description and start giving it an actual instance. To do that, I'm using a feature in Higgs Field called
elements, which lets you save a character or a location once and then reuse it in any video just by tagging its name. The first step is to actually create your character. So, click on image from the top navigation bar and select GPT image 2 as the model because it's currently one of the best when it
comes to hyperrealism. Now, instead of describing a scene, I'll prompt for a single character sheet with three panels of the same man. One from the front, one from the back, and a close-up of his face all on a plain gray background. That plain background actually matters because the less there is going on
around him, the easier it is for the model to lock his face down. And the front panel is headless on purpose. So, the only face in the whole sheet is that one clean close-up, which gives the model a single face to copy instead of averaging a few slightly different ones. So, now using that character sheet, I'll
make three more versions of that same person in a different outfit. For the first one, I want him in a suit. So, I'll upload the original character sheet as a reference and prompt for an outfit swap while keeping his facial identity locked in. He still looks exactly the same, but he's now fully suited up. I'll
save him as an element, set the category to character, and name him man suit. And I also made the other two versions the exact same way. One where he's wearing his boxing outfit and one where he's in an everyday casual fit. All of these are saved up and ready to be used. Then, I'll do the same for the locations. I'll
generate each one empty with nobody in it and save them as elements, too. The arena, the Tokyo street, and the car interior. Now, for the actual videos, I'll switch over to Cinema Studio 3.5. And this is the real jump at level five because all that camera and lighting control I had to type out by hand back
in level four now lives in panels right in the interface. And this time, they're actual settings, so the model is forced to generate with these. I'll also generate everything here in a wider 21 by 9 for that proper cinematic look. So, for the first shot, I'll tag man boxer walking through arena aisle, and then
I'll set up the panels one by one. First is the genre, which sets the whole mood and pacing of the shot. And I'll put it on epic so it carries that big dramatic movie feeling. Then the style, which I switch to manual so I can write the lighting myself. A hard white light straight down from overhead with deep
crushed blacks and high contrast because that's exactly the look of a real fight walkout. Then I open the camera settings, which give me [music] four separate dials to work with. The camera itself sets the base look, so I'll set it to fine film for that filmic quality. The lens controls how sharp everything
is and how much the [music] background blurs. So I'll put it on anamorphic for that wide cinematic feel. The focal length is how wide or tight the shot is framed and I'll leave it at 50 mm so he's framed naturally without any distortion. And the aperture is how much of the scene stays in focus. So I'll
keep it at F4 to hold the whole shot sharp from front to back. >> Time to end it all. >> And when it generates, the difference is clear. The character remains consistent throughout, the atmosphere is genuinely cinematic, and the whole video looks like a scene out of a movie. Then I'll
take that same person, put him in man suit, drop him into Tokyo street, and change the panels to match the new mood. I'll set the genre to noir this time, switch the camera to clean digital for a sharper modern look, and open the aperture right up to F1.4 so the background blurs out behind him.
>> It's surely cold today. >> And it's the same man in a totally different scene with his outfit and the location matching their references. And for the last one, I'll switch him to man cool and put him inside car interior driving through the mountains. Here, I'll set the genre to general, put the
lens on clinical sharp so every detail is highlighted, and lock the camera completely still so it feels mounted right there on the dashboard. >> Seems like I will be there on time. >> And again, it's the same person in a completely different environment. Compare the first boxer video we made
with this one, and the difference should be clear. That's the level of quality this [music] stage operates at, and the whole thing scales. Each level is building on the previous one, so by the end, you can use all your knowledge along with the best tools to create something truly cinematic. With
optimized prompting, controlled JSON format, locked character consistency, and the director's controls of Cinema Studio, where everything is set, not only can you create AI videos on a professional level, but you can do it in a fraction of the time and create these at scale. So, if you want to master
these five levels and create cinematic videos like these yourself, use the link in the description to sign up to Higgsfield. Thanks for watching, and I'll see you in the next one.
How videos are chosen here
Every video on Helicopterstour.com is hand-picked and reviewed by Justin — nothing is added automatically. Each one gets an original written guide and an honest rating: ⭐ 1 out of 2 means a good video worth your time, and ⭐⭐ 2 out of 2 means a great one we would recommend to anyone. The videos belong to their creators — every page links back to the original channel so you can subscribe and support them.
