How to Use Sora AI: A Video Prompting Guide

How To, Insights | Published on | Last updated on

10 min

Update (July 2026): Sora 2 and Sora 2 Pro are no longer available on Leonardo.Ai following a provider-side change by OpenAI. Veo 3.1 Fast and Veo 3.1 are the recommended replacements — see our Veo 3 prompting guide — and if you're exploring beyond the Veo family, Seedance 2.5 is our newest and most popular video model. The techniques below still transfer well to other video models.

It’s not always simple to describe in words what you already see with your mind. You might have a perfect shot visualized, but finding the exact words to describe it to an AI can be a challenge. That’s why we created this prompting guide for Sora 2.

We’ll equip you with a reliable vocabulary and a set of proven workflows to improve your prompt engineering skills. To do this, we’ve combined official guidance from OpenAI with our own extensive internal experiments to share what we’ve found works best for creating high-quality videos with Sora 2.

While the guide is focused on Sora 2, most of these principles apply to other video generation models as well.

The Core Components of a Sora 2 Prompt

An effective Sora prompt is built from a few core components that tell the model what to create. Leaving any of these out forces the model to improvise, which can lead to generic or unpredictable results (we’ll discuss balancing creativity and control in a moment). Here are the four elements you should always specify in your prompt:

  1. Style: Think of the overall visual aesthetic as the foundation for your shot. Whether it’s a “high-end TV commercial,” “gritty realism,” or “16mm black-and-white film,” this choice provides the visual framework that shapes every other detail the model generates.
  2. Subject: Who or what is the main focus of the scene? Be specific. Instead of just “yogurt,” try “a spoonful of thick, creamy Greek yogurt.”
  3. Action: What is the subject doing? Use clear, descriptive verbs that point to visible results. For example, “the honey coiling beautifully as it lands” is more evocative than just “add honey.”
  4. Scene: Where and when is this taking place? Describing the location, environment, and time of day establishes the context. “A sun-drenched kitchen with a white marble countertop” provides more visual information than “a kitchen.”

Let’s put together the details we’ve outlined above into a prompt:

Prompt

In a sun-drenched kitchen, a spoonful of thick, creamy Greek yogurt is gently placed into a glass bowl on a white marble countertop. Fresh, vibrant blueberries and a drizzle of golden honey are added, the honey coiling beautifully as it lands. Style: High-end food commercial, bright and clean aesthetic.

Balancing Creativity and Control

Notice in the video above that because we didn’t specify the camera framing or movement, the model had to improvise, defaulting to a static, close-up shot. Similarly, because we didn’t mention sound, the model made its own creative choice and added some generic instrumental music.

This brings us to a fundamental concept in prompting: the balance between creative freedom and directorial control. Sora 2 can act as either a creative partner or a precise execution tool, and your prompt determines which role it plays.

Shorter, less descriptive prompts invite surprising interpretations. Our yogurt prompt is a good example of this. It defines the core elements (style, subject, action, scene) but leaves camera work and audio open. This is a great strategy when you’re brainstorming or want the AI to show you an interpretation you might not have thought of.

Longer, more detailed prompts restrict the model’s creativity to achieve a specific result. This can be particularly useful if you need to match a storyboard or a client’s brief. Let’s enhance our original prompt to lock in the camera movement and sound:

Prompt

In a sun-drenched kitchen, a spoonful of thick, creamy Greek yogurt is gently placed into a glass bowl on a white marble countertop. Fresh, vibrant blueberries and a drizzle of golden honey are added, the honey coiling beautifully as it lands. Style: High-end food commercial, bright and clean aesthetic. Camera: Single continuous shot without cuts; slow arcing macro of the bowl with shallow depth of field. Audio: Upbeat acoustic music, faint spoon clink.

You might notice that even though we specified a single continuous shot, we still got a two-shot video. As you’ll quickly discover, Sora 2 has a slight tendency to create videos with multiple cuts, likely because it was trained on many such examples. This can be an asset or a liability, depending on whether you need a single take.

That’s why it’s a good practice to alternate between models depending on your goal. On the Leonardo.Ai app, you can easily switch between models to take advantage of each one’s unique strengths.

If you’re evaluating multiple models, you might also want to review our article comparing Sora to Veo 3 in real-world workflows.

Now that we’ve covered the essential building blocks, we can move on to the advanced techniques that give you more control.

Camera Control

To gain more control over your output, you need to shift from vague, descriptive language to specific, technical direction. For example, a common mistake is asking for a “cinematic look.” The term is too subjective and forces the model to improvise, which may not match your vision.

Instead, strong prompts use industry-standard terms that function as clear commands. Let’s break down the three main elements of camera control:

  • Camera shots (framing): This defines how much of the subject and their environment is visible. Be specific with terms like wide shot, medium shot, close-up, and over-the-shoulder shot.
  • Camera angles: Specify the camera’s position relative to the subject. A low angle shot can make a subject seem powerful, while a high angle shot can make them appear small or vulnerable.
  • Camera movements: Describe how the camera moves. Use professional terms like dolly (moving forward/backward), track (moving sideways), pan (turning horizontally), tilt (turning vertically), crane (moving up/down), etc.

While Sora 2 can handle complex instructions, the most reliable approach is often the simplest: limit each shot to one clear camera move to ensure the best prompt adherence, like in this example:

Prompt

From a slightly low angle, a hand-held medium close-up focuses on a young man in his 20s with headphones on, sitting on a city bus at night. Rain streaks down the window as the city’s neon lights reflect across their thoughtful face. Style: Indie film, moody, contemplative, shallow depth of field.

How to Prompt for Dialogue and Sound With Sora 2

Sora 2’s ability to generate video with native synchronized audio is one of its most powerful features, allowing you to create scenes that feel alive and immersive. To get a great result, you need to prompt for sound with the same precision you use for visuals.

Diegetic vs. Non-Diegetic Sound

Before we get into prompting, it’s helpful to know the two main types of sound in film:

  • Diegetic sound: These are sounds that exist within the world of the story, which the characters can hear. This includes dialogue, footsteps, rain, or a car radio playing.
  • Non-diegetic sound: This is sound that only the audience can hear, added to enhance the mood. This includes the film’s musical score, background music in an ad, or a narrator’s voice-over.

Sora 2 can generate both, and being clear about which you want is key.

Best Practices for Dialogue

Getting dialogue to sync correctly requires following a few best practices:

Getting dialogue to sync correctly requires following a few best practices:

  • Use a separate block: This is the most important rule. Following OpenAI’s official documentation, dialogue must be placed in a separate block below your visual description, clearly labeled “Dialogue.”
  • Keep it concise: Long speeches are unlikely to sync well within a short clip. Keep lines brief and natural to match the video’s length.
  • Label speakers: For scenes with more than one character, label each speaker consistently (e.g., “Detective:”, “Suspect:”) to help the model associate the line with the correct person.

Let’s see an example:

Prompt

Scene 1 Waiting under a flickering neon canopy. Drones hum overhead as rain falls in streaks of blue light. Style: Futuristic noir, moody cyberpunk palette. Camera: Slow pan as droplets catch the neon reflections. Dialogue: Detective: “He’s late.” Partner: “Maybe his signal got jammed.” Scene 2 (5–10s) The detective crushes a glowing data chip under his boot. A distant hovercar beam cuts through the scene. Both look up as static crackles across their comms. Style: Futuristic noir, moody cyberpunk palette. Camera: Low-angle tracking shot, emphasizing reflections and shadows. Dialogue: Detective: “He’ll come. The system always sends replacements.” Partner: “Unless we’ve already been replaced.” Scene 3 (10–15s) A figure materializes from a shifting hologram, half-glitching, holding a pulsing cube. The alley lights flicker as time seems to warp around them. Style: Futuristic noir, moody cyberpunk palette. Camera: Slow dolly in toward the figure as distortion deepens. Dialogue: Detective (quietly): “Show time.” Partner (nervous): “Or the end of time.”

Best Practices for Sound Design and Music

Effective sound design makes your scene feel authentic and lived-in. Details are what will help you create a rich environment: a café isn’t just “noisy”; it’s “the hiss of an espresso machine, the clinking of ceramic cups, and a low murmur of conversation.”

Alternatively, use a single, powerful sound for dramatic effect. A prompt for a tense scene might not need a full score if it has “the loud, sharp creak of a floorboard overhead.” This grounds the scene in realism and can build suspense more effectively.

Music is your most direct tool for telling the audience how to feel. Specify genre, mood, and tempo. Instead of “happy music,” guide the model with descriptive language. Is it “a gentle, upbeat acoustic guitar track” for an inviting ad, or “a fast-paced electronic track with a driving beat” for an energetic one?

Suggest instrumentation for even more control. A request for “a slow-building thriller score with low, ominous strings” will produce a much more specific and effective result than simply asking for “suspenseful music”.

Let’s see an example:

Prompt

A hero in a dark tactical suit runs down a rain-slicked alleyway at night, splashing through puddles. Style: Gritty, high-contrast action film. Camera: Handheld tracking shot from behind the hero. Audio: The sound of heavy breathing and splashing footsteps, layered with a fast-paced, percussive action score.

How to Prompt for Multi-Shot Scenes With Sora 2

Often, the video you have in mind is not a single continuous shot, and it needs cuts, different angles, and a sequence of actions. With Sora 2, there are two main strategies for creating a multi-shot video.

Method 1: Generate Separately (When Consistency Isn’t Key)

If the video you have in mind is composed of 2-3 shots and character or spatial consistency isn’t a primary concern, it’s worth generating the shots independently and then stitching them together in video editing. You don’t need to be a professional editor to do this—you can just use a drag-and-drop tool like Canva’s free online video editor.

You’ll write a separate prompt for each shot and generate them individually. The advantage is having maximum control over each individual shot. Here, we’ll create a funny ad for a travel agency using this method:

Prompt 1 (4-second video):

Prompt

A squirrel in a tiny business suit sits at a miniature desk in a high-rise office, looking stressed. He adjusts his tiny tie and sighs at a huge stack of paperwork. Style: Whimsical and funny commercial, cinematic. Camera: Static medium shot at the squirrel’s eye-level. Dialogue: Squirrel: (sighing) “I’m nuts to be working this late.” Audio: Soft office hum, the rustle of tiny papers.

Prompt 2 (4-second video):

Prompt

A squirrel wearing tiny sunglasses sits on a miniature beach chair on a tropical beach, looking relaxed. He holds a small coconut drink with a tiny umbrella in it. Style: Bright, sunny, and relaxed commercial. Camera: Static wide shot. Dialogue: Squirrel: (contented sigh) “Ahh, much better.” Audio: Gentle ocean waves and a soft calypso music track.

Method 2: Timeline Prompting (For Character Consistency)

When you need the same character to appear across multiple cuts in a single scene, using timeline-specific instructions might be a better approach. Here’s an example:

Prompt

Style: TV Commercial [00:00-00:04] A graphic designer squints at his laptop in a brightly lit, busy café, struggling to see the screen because of the glare from a window. [00:04-00:08] Close-up on the designer’s hands as he quickly and easily applies the new anti-glare filter to his laptop screen. [00:08-00:12] Medium shot of the same designer, now smiling and working effortlessly. The laptop screen is perfectly clear and visible.

Despite a minor artefact on the laptop in the third shot, the character consistency is perfect. Note that because we didn’t specify any audio, the model improvised within the boundaries of the TV-commercial style we specified.

You can see more video examples of this method in this LinkedIn post by Mark Isle.

How to Prompt for Character Consistency With Sora 2

As we saw in the previous example, the timeline prompting method is great for keeping a character consistent within a single scene. Because the entire multi-shot sequence is generated from one prompt, the model has a better chance of remembering the character’s appearance from one cut to the next.

Another method you can use is the image-to-video (I2V) workflow, which allows you to maintain character consistency across multiple shots. This is particularly useful if you have a storyboard with 10-20 shots and all demand strong character and spatial consistency—you won’t be able to use the timeline method for this use case.

Let’s see an example of an image-to-video workflow in Leonardo.Ai, where we use:

  1. Lucid Origin to create a character
  2. Nano Banana to create a two-shot sequence
  3. Sora 2 to animate the images

At the moment of writing this blog, OpenAI’s API restricts the use of people in the image-to-video workflow, so we’ll demonstrate this method on an animation. We start by creating our hero using Lucid Origin:

Prompt

Character concept art of Sir Reginald, a large, friendly-looking knight with slightly oversized, dented armor. He is standing on a stylized New York City sidewalk, looking utterly confused. Style: 3D animation, bright and colorful, comedic tone.

Using Nano Banana, we’ll create two shots of the mighty Sir Reginald, a knight discovering that fighting dragons was somehow less complicated than navigating New York City:

Using the first image only (the one on the left), we’ll combine an image-to-video workflow with the timeline-specific prompting we’ve just learned earlier and build a video with two shots:

Prompt

[00:00 – 00:04] A knight in full armor points a finger at a hot dog menu, looking lost. Pedestrians on the busy NYC sidewalk stare in amused disbelief. Audio: The busy sounds of a New York City street. [00:04 – 00:08] Cut to a medium shot. The knight offers the hot dog vendor a large, shiny gold coin. Dialogue: Knight: “Hark, vendor! I offer this coin for one of your ‘hot dogs’.” Vendor: (Squinting at the coin, unimpressed) “Is that from a video game? Look, pal, it’s six bucks. Card or tap.”

Next, we use the second image to create a shot from a different scene (looks like Sir Reginald got the hot dog, but isn’t entirely happy about the whole situation):

Prompt

The knight is deep in his thoughts, he is confused. Dialogue: Knight: (in a deep, defeated voice) “I want to go home.”

One important thing to notice is that prompts used in an image-to-video workflow don’t need to follow all the rules we learned in the prompting essentials section. That’s because the image already provides most of the information regarding style, subject, and context. The best thing to do is to focus your prompt on these three elements:

  1. Action (what you want to happen in the video—this one is a must)
  2. Camera movement (optional but recommended)
  3. Audio (optional but recommended)

How to Prompt for Viral Social Media Videos With Sora 2

Prompting for virality is less about creating a perfect cinematic shot and more about crafting a concept that is instantly understandable, emotionally resonant, and surprising enough to make someone stop scrolling.

Analysis of high-performing AI videos suggests they activate four key triggers:

  1. Surprise (the unexpected)
  2. Clarity (instantly get the joke or the situation)
  3. Familiarity (a pop culture reference or a relatable situation)
  4. Emotion (joy, curiosity, nostalgia)

Your goal is to combine these elements into a short, punchy video concept. While there’s no magic formula for going viral, there are reliable structures that increase your chances.

  • Use the “Hook, Twist” formula. This is a proven framework for building a scroll-stopping video.
    Hook: A visual contradiction or absurd concept that grabs attention immediately (e.g., “a celebrity shoplifting”).
    Twist: Something unexpected happens that makes the video memorable and shareable (e.g., “the celebrity is shoplifting macaroni and cheese”).
  • Use proven viral archetypes. Try building your prompts around these archetypes :
    Contradictory context: Place a historical or fantasy figure in a mundane, modern situation. This creates instant surprise and humor.
    Anthropomorphism: Give human-like emotions and behaviors to animals or inanimate objects. This creates relatable and often whimsical narratives.

Most social media videos use a portrait (9:16) resolution, but you can also create them in landscape (16:9)—this won’t hurt the video’s performance, as long as the content is highly engaging. Let’s see an example:

Prompt

Style: Funny social media video A cat slowly pushes a glass of water off the edge of a table. The glass falls and breaks. The cat looks directly into the camera with a cold, calculated expression. Dialogue: Cat: (in a villainous voice) “It was… inevitable.”

Conclusion

Getting the most out of Sora 2 requires a clear vision, a specific vocabulary, and a strategic approach to bringing your ideas to life. The model is a powerful creative partner, but it relies on your direction to produce a high-fidelity result.

While the techniques might seem a bit complex at first, they provide a reliable framework for translating your creative vision into a finished video. The tools are more powerful than ever, but the story, the emotion, and the vision need to come from you. These prompting fundamentals carry straight over to Veo 3.1 and the other video models on the Leonardo.Ai app — practice what you’ve learnt in this blog there!