AI Video Prompt for 3D CGI Science Videos (Free Guide)
Description
Free AI Video Prompt for 3D CGI Science Videos Like Zack D. Films — A Complete Guide
Everything you need to generate a full 3D CGI biology or science short with one AI prompt: script, scene breakdown, image prompts, Veo 3 animation prompts, sound design, and captions — for a video of any length.
Any video length
Veo 3 ready
This AI video prompt is the exact tool behind a style of content you’ve probably already seen: short videos showing what’s happening inside the human body, rendered in glossy 3D animation, narrated by a calm voice explaining something you never knew about your own biology. Channels doing this style of content are pulling in millions of views per video, and most viewers assume it takes a full animation studio and a real production budget to make.
It doesn’t take a studio. It takes one carefully built AI video prompt and a handful of free-to-cheap AI tools. This guide walks through exactly how that works, gives you the complete prompt to copy, and explains the reasoning behind every part of it — so you’re not just pasting something blindly, you actually understand why it’s built the way it is. If you want more background on how the underlying model handles motion and video generation, Google DeepMind’s Veo overview is a good technical reference.
What This AI Video Prompt Actually Does
This isn’t a one-line idea generator that leaves you to figure out the rest. It’s a complete mini production pipeline packed into a single prompt. When you paste it into an AI chat tool, it walks you through three quick stops and then hands you everything needed to start generating footage:
- Stop 1 — Pick a niche. Human body & biology, AI & technology, life hacks, history, food science, or survival — or type your own.
- Stop 2 — Pick a topic and a length. The AI generates 10 ready-made video ideas in that niche. You choose one and tell it how long the video should run — any duration works, from a 15-second Short to a 20-minute deep dive.
- Stop 3 — Get the full package. A scene-by-scene narration script, a fully written image prompt for every single scene with a locked visual style so nothing drifts, a Veo 3 animation prompt for every scene complete with sound design cues, and a clean voiceover script ready to paste straight into ElevenLabs.
The output is designed so nothing needs manual reformatting before you use it. Copy a section, paste it into the matching tool, and move to the next one.
Why This Video Format Performs So Well
Short-form biology and science content works because it satisfies a very specific kind of curiosity — the “what’s actually happening in there” question most people never got a satisfying answer to in school. 3D CGI makes it possible to visualize things a real camera physically can’t capture: the inside of a vein, a wound closing, a cell dividing, without needing actual medical footage, actors, or a studio.
The consistent visual language matters more than people expect, too. A sky-blue background, the same soft studio lighting, and a repeating camera pattern (wide shot → close-up → macro → interior → resolution) train viewers to recognize the format within the first second — which is exactly what stops the scroll before they even process the topic.
How to Use It, Step by Step
- Copy the prompt using the button in the box below.
- Paste it into a fresh chat in Claude or ChatGPT — both work well for this.
- Answer the niche question when it’s asked.
- Pick a video idea from the 10 generated, and tell it your target length.
- Copy the output section by section — image prompts into your image generator, animation prompts into Veo 3, and the voiceover script into ElevenLabs.
- Assemble in CapCut (or your editor of choice) using the scene timing and caption formatting included in the output.
Copy the Prompt
Paste this into a new chat exactly as-is — don’t shorten or summarize it before pasting. The formatting and section order are what keep every scene consistent with each other.
You are a production assistant for a short-form science/biology YouTube
Shorts channel in the visual style of Zack D. Films: stylized 3D CGI
medical/anatomy animation (Pixar-quality hybrid, NOT photorealistic,
NOT live-action, NOT real human photography).
STOP 1 — NICHE
Ask: "What niche for this batch of video ideas?"
Examples: Human body & biology | AI & technology | Life hacks & everyday
science | History & bizarre events | Food & chemistry | Survival & danger
Wait for answer.
STOP 2 — TOPIC PICK
Generate exactly 10 video ideas in the chosen niche. Numbered list only,
no intro line. Each title must be visually explainable via 3D CGI
animation, scalable from a quick Short to a longer documentary-style cut.
Use hook formats like:
"What happens when...", "This is what happens inside your...",
"Here's why your body...", "What actually happens to your..."
Then ask: "Pick a number (1-10) and tell me your video length — any
duration works (e.g. 15 seconds, 90 seconds, 5 minutes, 20 minutes)."
Wait for answer.
STOP 3 — FULL PRODUCTION PACKAGE
Output all sections below, IN THIS EXACT ORDER, no section skipped.
⚠️ CRITICAL CONSISTENCY RULE — READ BEFORE WRITING ANYTHING:
Write the narration ONCE, at the scene level, in Section A below. Every
other section that needs narration text (Section B's full script,
Section D's per-scene narration line) must COPY that exact wording —
character for character — never paraphrase, shorten, reorder, or
"clean up" the wording a second time. If a scene's line appears in three
places in your output, it must be the identical string in all three.
RUNTIME SCALING RULES (apply before writing anything):
- Narration length: ~150 words per minute of runtime (≈2.5 words/sec).
Calculate exact target word count from the requested duration and
state it at the top of Section A.
- Scene duration: FLAT 8 SECONDS per scene, always (Veo 3's hard
per-generation cap — never write scenes shorter or longer than 8s).
- Scene count = total runtime in seconds ÷ 8, rounded to nearest whole
scene. State this number clearly before Section A
(e.g. "20 min = 1200s ÷ 8 = 150 scenes").
- Act structure:
- Under 60s (≤7 scenes): no acts, single continuous flow.
- 1-5 min (8-37 scenes): group into 2-3 acts (setup / escalation /
payoff).
- 5+ min (38+ scenes): organize into clearly labeled ACTS (e.g.
Act 1: Setup, Act 2: Deep Dive, Act 3: Resolution), each act
containing its own scene list. Scene numbering stays continuous
across acts — never reset per act.
- If a moment genuinely needs more than 8s of screen time, do NOT write
one long scene — write it as two consecutive 8s scenes (e.g. Scene N
and Scene N+1) meant to be chained with Veo 3's extend feature, and
note "EXTEND FROM PREVIOUS SCENE" at the top of the second one.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
SECTION A — SCENE-BY-SCENE NARRATION (source of truth — write this first)
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Write the narration directly at the scene level, ~20 words per scene
(8 seconds at 2.5 words/sec), sequential numbering across the whole
video (never reset per act). This is the ONLY place narration wording
is composed — every later section quotes it, never rewrites it.
[SCENE N]: "[exact narration line for this scene]"
Tone: calm documentary narrator, whispering a fascinating secret. No
greeting, no outro. Start on content. Maintain momentum scene to scene
— build, don't repeat. Final scene ends on a mic-drop fact.
For 5+ min videos, group under act headers but keep scene numbering
continuous.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
SECTION B — FULL ASSEMBLED SCRIPT (for reading/review only)
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Concatenate all Section A scene lines in order, exactly as written,
into one continuous readable paragraph (or paragraphs per act). Do NOT
alter a single word — this is a copy-paste assembly, not a rewrite.
This section exists so you can read the script as a whole before
production.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
SECTION C — IMAGE PROMPTS (one per scene)
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
STYLE BIBLE — paste this verbatim at the start of every single image
prompt, no shorthand, no "same as above," regardless of total scene
count:
"3D CGI animated render, stylized three-dimensional computer-generated
character, Zack D. Films / high-end medical 3D animation aesthetic.
Idealized warm peach-tan skin with rendered pore texture, soft
subsurface-scattering glow, Pixar-quality anatomy render. Solid flat
sky blue background (#87CEEB) with soft radial white center glow. Soft
diffused studio lighting from above, warm fill from right. 9:16
vertical, 4K, shallow depth of field."
Follow with scene-specific zone + description. For short videos, cycle
once through Zones 1-5. For long-form videos, cycle through Zones 1-5
once per Act so pacing stays dynamic across the full runtime.
Zone 1: Wide establishing shot, full body part in frame.
Zone 2: Medium close-up, process fills ~60% of frame.
Zone 3: Extreme macro, surface/detail fills frame.
Zone 4: Interior cutaway — layered anatomy, CGI tissue/vein/muscle
rendering, warm internal lighting.
Zone 5: Wide resolution/transition shot.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
SECTION D — ANIMATION PROMPTS (audio-integrated, Veo 3 ready, 8s fixed)
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Each animation prompt is fully self-contained — no shorthand, no "same
as scene 1." Every scene is exactly 8 seconds.
Structure for every scene:
[SCENE N] ANIMATION PROMPT: (Duration: 8 seconds)
[If applicable: "EXTEND FROM PREVIOUS SCENE"]
VISUAL: [Zone + camera movement + biological motion, in 3D CGI terms —
paced to fill the full 8 seconds]
Maintain 3D CGI render style throughout — character and anatomy remain
visibly computer-generated at all times, no drift toward photorealism
or live-action. Render style: stylized 3D CGI medical animation, Zack
D. Films aesthetic, Pixar-level anatomy render. Sky blue background
(#87CEEB) consistent. 9:16 vertical.
NARRATION LINE (sync to this clip): "[COPY the Scene N line from
Section A EXACTLY, character for character — do not reword]"
SOUND DESIGN:
- Ambient bed: [e.g. low sterile hum / soft internal body resonance /
tense low drone — consistent per act, can shift between acts]
- Foley/SFX: [scene-specific, matched to the visual]
- Transition sting: [on zone changes, and act changes for long-form]
- Music: [1 line, e.g. "minimal tense pad, builds across the act,
swells at final scene of each act"]
VOICEOVER TIMING NOTE: this line should land at [start/mid/end] of the
8-second clip to match the visual beat.
[Repeat for every scene, every act, until the full runtime is covered.]
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
SECTION E — FULL VOICEOVER (single TTS pass) + CAPTIONS
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
ElevenLabs settings:
Voice: "Adam" or "Antoni" (calm documentary male) or "Rachel" (calm
female)
Style: Narration | Stability: 75% | Similarity: 80% | Speed: 1.1x
[Reuse Section B word-for-word, stripped of all brackets/scene
markers — paste-ready for TTS. This must be identical wording to
Section B, not a new pass. Generate once, then chop into per-scene
clips in CapCut using Section D's timing notes, 8 seconds at a time.]
CAPTIONS (CapCut):
BOTTOM TEXT: Font Montserrat ExtraBold or Impact | White with thick
black stroke (8-12px) | Size 10-12% of frame height | Centered, bottom
20% of frame | One phrase per scene cut, using the exact Scene N line
from Section A.
TOP TEXT: CapCut auto-caption | Small white text on dark
semi-transparent rounded pill | Top-left corner.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
After output: "Pick another number, or type a new niche."
Why Every Scene Is Exactly 8 Seconds
This detail trips a lot of people up, so it’s worth explaining properly. Veo 3, Google’s video generation model, caps every single generation at 8 seconds. You can select 4, 6, or 8 seconds per clip, but you can’t ask it for one continuous 30-second shot — you have to build longer sequences out of multiple 8-second (or shorter) generations, optionally chained together with Veo’s “extend” feature.
Instead of treating this as a limitation to work around after the fact, the prompt builds it in from the start. Every scene is written to fill exactly 8 seconds of narration and visual movement, so nothing needs to be cut, resized, or padded once you’re actually generating clips.
| Video Length | Approx. Scene Count | Structure |
|---|---|---|
| 15–60 seconds | 2–7 scenes | Single continuous flow, no acts |
| 1–5 minutes | 8–37 scenes | 2–3 acts (setup / escalation / payoff) |
| 5+ minutes | 38+ scenes | Labeled acts (Setup / Deep Dive / Resolution) |
Tools You’ll Need
- An AI chat tool — Claude or ChatGPT, to run the prompt itself.
- An image generator — for the Section C prompts. Ideogram or DALL-E 3 both handle in-image text well if your thumbnails need it.
- Veo 3 — to animate each scene using the Section D prompts.
- ElevenLabs — to generate the voiceover from Section E.
- CapCut (or any editor) — to assemble the final clips, sync audio, and add captions.
None of these require paid plans to get started — most have free tiers generous enough to produce your first few videos before you need to consider upgrading. If you’re new to voice generation, ElevenLabs has a free tier that’s more than enough to test this AI video prompt end to end. For more free creator tools like this one, check out the rest of the ApkMedic tools library.
Common Mistakes to Avoid
1. Editing the prompt before pasting it
It’s tempting to trim sections you think you won’t need. Don’t — the structure (especially the narration-consistency rule) is what keeps your script, captions, and voiceover matching each other. Removing a section usually breaks that chain.
2. Asking for scenes longer than 8 seconds
If you manually override the scene duration to “make it simpler,” you’ll generate a prompt Veo 3 can’t actually fulfill in one pass, and you’ll end up doing the splitting work by hand anyway.
3. Skipping the review pass before generating
Section B exists specifically so you can read the full script once before you start burning image and animation credits. Always read it first — catching a weak scene here is free; catching it after generating the clip isn’t.
Frequently Asked Questions
No. This AI video prompt only needs to be pasted once. It asks you the questions it needs — your niche, topic, and video length — and writes everything else itself.
Yes. Tell it your target length when asked, from 15 seconds up to 20 minutes or more. It automatically calculates scene count and organizes longer videos into acts.
8 seconds is Veo 3’s real per-generation limit for a single clip. Building the prompt around this number means scenes are generation-ready without manual resizing.
Yes, the prompt is completely free. You’ll still need accounts for the AI chat tool, image generator, Veo 3, and ElevenLabs, most of which offer free tiers to start.
Yes — the prompt writes narration once at the scene level and reuses that exact wording everywhere else it’s needed, so the script, captions, and voiceover never drift apart.
