Back to Blog
AI Tools
Oct 1, 2026
17 min read

How We Make Motion Graphics Videos With AI (Real Examples Inside)

AI makes our voice, music and 3D icons. Code makes every frame. Here is the full workflow behind our Instagram Reels, from script to final render, with two real examples.

I. M.

Editing timeline with microphone and globe clips and a golden waveform leads to a vertical video screen showing the liquid-gold MeetingsAI logo, with the app icon above.
#AI Tools#Tutorial#Productivity

Introduction

The best way we have found to make motion graphics videos with AI is to split the job. AI generates the raw ingredients: the voice, the music and the 3D icons. An AI coding agent writes the animation itself as code, so every logo, color and word on screen stays exact. Then we check the result with measurements instead of guesses.

That is how we made the two Reels in this post: a 30 second feature video and a 15 second liquid-gold logo animation. Below is the full workflow, what AI is good and bad at in this job, and how we packaged the process into a Claude Code skill called make-meetingsai-video, so each new video starts from a checklist.

One disclosure up front: these are our own videos for our own app, MeetingsAI. This post describes what worked for us. It is a workflow, not a law.

Key Takeaways

  • AI video generators redraw every frame, so logos and on-screen text can drift. Code draws the same frame the same way every time.
  • Our split: AI makes the voice, music, 3D icons and some sound effects. Code (Remotion and React) makes all the motion, text and app UI.
  • The voiceover is the clock. Word timestamps decide when each scene starts, so visuals land on the spoken word.
  • Every video gets checked with tools: still frames, loudness, color tags, and a speech-to-text pass on the final mix.
  • We saved the process as a Claude Code skill. It holds the steps, the brand rules and a growing list of mistakes not to repeat.

See the Result First: Two Reels Made This Way

Before the how, here is the what. Both videos are posted on our Instagram account, and both were built with the workflow below.

Example 1: The 30 second feature reel

A photo of messy handwritten notes gets crossed out with a marker stroke. The logo springs in. A phone rises and shows live transcription, live translation, Cheat Mode, the summary and chat, then a dark Private Mode screen with an airplane flying across it.

The phone screens are not recordings. They are React components, so every line, chip and checkbox can animate on its own. The voice, captions, music and sound effects are timed to the spoken words. The video is 1080 by 1920 pixels at 30 frames per second, 900 frames in total.

Example 1: the 30 second feature reel

Watch for three things: the captions switch on the exact spoken word, the marker scribble draws itself, and the phone UI moves element by element.

Example 2: The 15 second liquid-gold logo

No video model drew this one. A single WebGL shader paints the logo in every frame: a molten orb cools into honey-colored glass, the hexagon forms, gold lines flow out along the logo and close on an impact 8.1 seconds in, then cool to ink as a cream wave sweeps across.

React adds the paper, the wordmark and the tagline. The music, impact and shimmer sounds came from AI models. The drop, the riser and the sub-boom were synthesized with code, because the AI versions came back silent.

Example 2: the 15 second liquid-gold logo animation

The same shader renders in three sizes: 9:16 for Reels, 16:9 for wide screens and 1:1 for feed posts.

Why Not Just Use an AI Video Generator?

AI video generators are good at footage that never existed: a street at night, a person walking, a product turning on a table. They are weaker at brand-exact graphics. They redraw every frame from a prompt, so a logo can warp, a color can drift and on-screen words can turn into almost-words.

For a product video, that is the wrong trade. The app's screen has to show the app's real words.

Where generators fall short for product videos

  • Text changes. In one of our own image tests, a model turned the words "Spell it right" on a phone screen into "Spelli it right". A video model has to get every word right on every frame.
  • Screens get redrawn. In the same tests, a tilted phone came back bent, and the text on its screen was redrawn instead of copied.
  • No exact repeats. Fixing one word means generating the clip again and hoping everything else stays put. In code you change one string and render again.
  • Brand drift. A hex color, a font or a logo shape is a suggestion to a generator. In code it is a value.

What code gives you

  • Fonts, colors and the logo live in one theme file, so every scene uses the same ones.
  • Fixing a typo is a one-line edit and a short render.
  • One design can render in more than one size. Our logo animation ships in three.
  • Motion is repeatable. The same code gives the same video every time.

Remotion, the React framework for video that we use, now documents prompting videos with coding agents like Claude Code.

What AI is best at in this job

Two things. First, making assets: a voice, a music bed, a set of 3D icons, a photo of messy notes for the hook. Second, writing the code that turns those assets into motion. That is where AI saves us the most time, so that is where we use it.

How to Make Motion Graphics Videos With AI: Our 8-Step Workflow

Here is the workflow in the order we run it. Each step exists because something went wrong without it.

Step 1: Plan the beats before you generate anything

Start with the script, then a beat list with seconds: what is on screen while each line is spoken. The feature reel has nine beats: the hook, the logo, live transcription and translation, Cheat Mode, the summary, chat, Private Mode, the airplane-mode proof and the end card.

A handy rule: our voice speaks about 2.5 words per second, so a 30 second video holds about 70 words and leaves 2 to 3 seconds for the end card. We show the plan and get a yes before we generate anything, because redoing a voiceover after the scenes are built hurts.

Step 2: Check every claim against the real app

AI is happy to invent features, and a video makes an invented feature look real. So before a scene is built, every claim is checked in the app's code. Two examples:

  • We say "50+ languages", not a bigger number, because the claim has to survive a check against the translation engine's real list.
  • We never show Cheat Mode working in airplane mode, because it needs an internet connection.

Step 3: Let AI make the assets, and only the assets

  • Voiceover. A text-to-speech voice from ElevenLabs. Write "Meetings AI" as two words in the voice script, or the voice says something like "Meetings-eye". On screen it stays MeetingsAI.
  • Music. The AI music model we use follows tempo instructions poorly, so we make two takes, measure the beat in each, and keep the one with the clearest beat.
  • Images. GPT Image 2.5 makes the hook photo and the 3D icons. We ask for icons as a 2 by 2 sheet on a transparent background, then crop each one, so the whole set shares one style.
  • Sound effects. AI sound effect models can return silence or hiss. Every file goes through a script that flags silence and noise, and simple effects such as a whoosh, tick, ding or boom are synthesized with code instead.

One more rule lives here: real text never goes into a generated image. The model redraws any text it sees, so every word on screen is text in code.

Step 4: Sync every scene to the voice

The voiceover is the clock. We run speech-to-text (whisper.cpp) on the finished voice to get a timestamp for every word, then start each scene on its word. When the voice says "Cheat Mode", the Cheat Mode screen is already moving.

Raw timestamps have one flaw: they drift around pauses, so the word after a pause can land in the wrong place. A small script aligns the timestamps to the script text and snaps pauses to punctuation. Without it, cues feel slightly off and you cannot say why.

Step 5: Build the motion in code

This is the coding agent's job. In Remotion, a video is a React component that receives a frame number and returns what that frame looks like. A few choices made our videos look polished:

  • Real UI, rebuilt. Phone screens are React components inside a CSS iPhone frame, using the app's real layout and strings. Nothing is a recording, so every element can animate by itself.
  • One theme file. Cream graph paper, yellow marker highlights and the brand fonts live in one place.
  • Springs and hand-drawn lines. Elements settle with springs, and marker strokes draw themselves on.
  • Safe zones. Reels and TikTok cover the top, the right side and the bottom of a 9:16 frame, so key content stays inside a box from 60 to 940 pixels across and 160 to 1520 pixels down.
  • Big words outside the phone. Text inside a phone mockup is too small to read on a real phone. Key words are said big, in headlines and callout pills outside it.

Step 6: Look at frames all the time

A coding agent cannot watch a video, but it can look at still images. After each change it renders stills at key frames, reads them, and fixes what is off. For a full pass it builds a contact sheet: one image with a thumbnail every half second. That catches overlaps, clipped text and safe-zone mistakes before anyone watches the video.

Step 7: Render and package for phones

Two details cost us time.

First, color. Remotion's default output can look washed out on phones, so we re-encode it with BT.709 color tags in limited range.

Second, loudness. We aim for about -14 LUFS, a common level for online video, with true peaks under -1.5 dB. We tried ffmpeg's automatic loudness mode first, but the music started to pump, so we switched to a fixed gain plus a peak limiter.

Step 8: Verify with measurements, then listen

An agent cannot hear, so it measures. The checks:

  • Loudness and peak level on the final file.
  • Silence and noise checks on every audio file.
  • For a voiceover: speech-to-text on the final mix, compared with the script. If a word is missing or wrong, we find out before Instagram does.
  • For an impact: a beat-detection script confirms the hit lands within one frame of the visual. At 30 frames per second, one frame is about 33 milliseconds.
  • Code checks: a type-check and lint on the video project.

Then a person listens to the final mix. Always.

How the Liquid-Gold Logo Animation Works

We had a PNG of our logo and no vector file. So the agent rebuilt the logo as signed-distance textures: images where each pixel stores how far it is from the logo's edge. One WebGL fragment shader reads those textures and draws every frame: the molten orb, the honey glass, six lobes, the inner hexagon, the gold lines flowing out along the logo, the impact, and the cooling to ink.

Remotion calls the shader once per frame. One gotcha: rendering needs Remotion's ANGLE GPU option, or the canvas comes out blank.

The sound design has one trick worth stealing. About 0.3 seconds before the impact, the music dips to around 20 percent of its volume, and a synthesized sub-boom hits on the exact frame. The short dip makes the hit feel bigger than any volume bump would.

Turning the Workflow Into a Claude Code Skill

A Claude Code skill is a folder with a SKILL.md file: markdown instructions that Claude loads when your request matches the skill's description. You can also run one by name. Ours is called make-meetingsai-video. When we ask for a promo, a feature demo or a logo intro, Claude reads it first and follows the same process every time.

What is inside

  • The workflow, in order. The eight steps above, each with the command to run.
  • Reference notes. Where the project's assets live (so existing voice, music and icons are reused before new ones are generated), the brand look, layout positions for each video size, a map of reusable components, and notes on the logo shader.
  • A claims list. What is safe to show in a video, and what is not.
  • A common mistakes table. One line per mistake, one line per fix: a blank WebGL canvas, washed-out colors, a voice that mispronounces the name, music that pumps, cues that land late after a pause. It grows every time a video teaches us something.
  • Small scripts. Voice timing, audio checks, beat detection, synthetic sound effects, before and after still comparison, and render and package.

The skill does not replace judgement. It removes repeat mistakes. It also keeps the agent honest: the skill tells Claude it cannot hear audio, so it has to measure it and ask a person to listen.

It is private. It lives on our machines and holds brand details, so we are not publishing it. Building your own is a small job.

How to write one for your own videos

  1. Finish one video with your coding agent and keep notes on every problem.
  2. Write the steps in the order you ran them, with the exact command for each.
  3. Add a common mistakes table, one line per mistake and one line per fix.
  4. Add the checks that prove a video is done: loudness, frame checks, a transcript of the final mix.
  5. Keep the main file short, and move detail into separate reference files that the agent opens only when it needs them.

What AI Still Gets Wrong

AI output is a draft. Every problem in this list happened to us at least once.

  • It invents features. Check every claim in the app's code (Step 2).
  • It redraws text. Keep real text in code, never in a generated image.
  • AI sound effects can be silent or noisy. Measure every file, and synthesize simple effects.
  • Voices mispronounce names. Spell the name the way it sounds in the voice script.
  • It cannot see or hear motion. Use stills, contact sheets, measurements and a human listen.
  • Timing drifts around pauses. Align word timestamps to the script.

Your First Motion Graphics Video With AI: A Simple Way to Start

Start smaller than a 30 second reel. The feature section of our home page has four silent 8 second loops, and loops like them make a good first project, because a loop has no voice, no music and no timing file. You only learn the motion part.

  1. Pick one idea that fits in 8 seconds, such as one feature in action.
  2. Follow Remotion's guide for Claude Code. At the time of writing, it installs the Remotion Agent Skills with npx remotion skills add.
  3. Describe the scene in plain words: what appears, in what order, in which colors and fonts.
  4. Ask the agent to render stills, look at them and fix what is off before it renders the video.
  5. Make the first frame the finished "poster" state and the last frame lead back into it, so the page looks right before the video loads and the loop has no jump.
  6. Compress it. We encode AV1 first, with H.264 as a fallback. Our AV1 loops come out between 130 and 300 KB each.

Keep the background still. A static background is the cheapest way to keep a loop small.

Frequently Asked Questions

Can AI make motion graphics?

Yes, in two ways. Generative video tools draw each frame from a prompt, which suits cinematic footage but can warp logos and text. A coding agent can instead write the animation as code, for example in Remotion, and render it to a video file. We use the second way for brand-exact motion graphics, and AI models for the voice, music and images.

What is the best way to make motion graphics videos with AI?

It depends on the job. For graphics that must match your brand, have a coding agent write the animation in code, and use AI only to generate assets such as voice, music and icons. For lifelike footage of people or places, a generative video model fits better. Many teams combine both.

Do I need to know how to code to make motion graphics with AI?

Not to start. The coding agent writes the code, and you describe the scene in plain words. It helps to be comfortable in a terminal. You still judge the result yourself, by looking at rendered frames and listening to the audio. The agent fixes things, and you decide.

What is Remotion?

Remotion is a framework for making videos with React. A video is a component that receives a frame number and returns what to draw, and Remotion renders it to a video file. Its documentation includes a guide for prompting videos with coding agents such as Claude Code. Check its license terms for your team size before you use it commercially.

What is a Claude Code skill?

A skill is a folder with a SKILL.md file: markdown instructions, plus optional scripts and notes. Claude Code loads it when your request matches its description, or you can run it by name. Ours, make-meetingsai-video, holds our video workflow, brand rules and checks.

How do you keep AI-made videos on brand?

Keep brand decisions out of the AI's hands. Fonts, hex colors, the logo and every word on screen live in code or in one theme file, and AI only generates assets such as voice, music and icons. Then check frames, not just the final clip.

Conclusion

AI does its best work in motion graphics when its job is narrow: make a voice, a song or an icon, or write the code that draws the frame. The brand, the words and the timing stay in code that you can read, test and render again. That is how one small process gives us a 30 second feature video and a liquid-gold logo, and how the next video starts from a checklist.

Want to see the app behind the videos? Download MeetingsAI and try live transcription, translation and Private Mode on your next meeting.

Share this article

I. M.

Related Articles