
Project Overview
lanshu-create-ai-presenter-video is a general-purpose AI presenter video production Skill for Codex. Simply provide a topic or complete script, along with an authorized reference image that clearly shows an adult, and it can orchestrate the main stages of AI presenter video production.
The complete workflow covers script preparation, voiceover, presenter generation, lip-sync calibration, subtitles, keyword animations, editing, rendering, and quality assurance. The project uses capability-based selection, completing tasks with the tools available in the current environment rather than binding the source code to specific providers, model names, or proprietary interfaces.
This Skill is not a standalone video generation model. It must run in Codex or an Agent environment compatible with local Skills and relies on accessible video generation, speech generation, and lip-sync capabilities within that environment.
Key Features
- Provider-neutral: The workflow does not require a specific model, platform, or proprietary API.
- Complete production pipeline: Dedicated processes cover everything from input validation, scripting, and voiceover to editing, subtitles, cover images, export, and quality assurance.
- Unified time reference: The complete voiceover serves as the timeline for positioning presenter footage, subtitles, shots, keywords, and transitions.
- Stage-based reference loading: Codex reads relevant documentation only when it reaches the corresponding stage, reducing context usage.
- Cost control: The workflow starts with a low-cost presenter test clip and defines paid-generation notices, retry limits, and interruption recovery rules.
- Security and privacy safeguards: Authorization is confirmed before presenter images are uploaded or voices are cloned, and the repository does not store API keys, access tokens, or user assets.
- Standardized delivery: The workflow is designed to produce a master file, a share-ready version, and a QA report.
Prepare Inputs and the Runtime Environment
Minimum Inputs
- A video topic or complete script.
- An authorized reference image that clearly shows an adult.
You can also provide a voice sample, screen recording, images, B-roll, brand assets, target platform, video duration, horizontal or vertical aspect ratio, visual style, watermark, and closing call to action. A voice sample may be used for voice cloning only with the appropriate authorization.
Environment Requirements
- Codex or an Agent environment compatible with local Skills.
- Python
3.9+. FFmpegandffprobe.- Bash,
jq,awk, andsed. - At least one video generation, speech generation, and lip-sync capability accessible in the current environment.
Before starting a task, confirm that these commands and generation capabilities are available in the current environment. The project repository is available on GitHub.
Install the Skill
Run the following command to clone the repository into the Codex Skills directory:
git clone https://github.com/cclank/lanshu-create-ai-presenter-video.git \
~/.codex/skills/lanshu-create-ai-presenter-video
After installation, the path should be:
~/.codex/skills/lanshu-create-ai-presenter-video
Understand the Directory Structure
lanshu-create-ai-presenter-video/
├── SKILL.md
├── README.md
├── agents/
│ └── openai.yaml
├── assets/
│ └── job.template.json
├── references/
│ ├── generation.md
│ ├── editing.md
│ └── qa-recovery.md
└── scripts/
├── init_job.py
├── preflight.py
└── finalize_delivery.sh
The three reference documents support different production stages:
generation.md: Covers input validation, scripting, voice, capability selection, paid generation, presenter prompts, and consistency.editing.md: Covers the timeline, opening and closing, subtitle presets, keyword animations beside the presenter, cover images, and export.qa-recovery.md: Covers technical validation, manual review, and fixes for common issues.
Method 1: Create a Video Directly in a Conversation
The simplest approach is to explicitly invoke the Skill in a Codex conversation while providing the script, presenter image, and final video requirements. For example:
Use $lanshu-create-ai-presenter-video to turn this script and presenter image into a 30-second, 16:9 AI presenter explainer video with real-time subtitles.
When submitting the request, it is best to include the asset locations, target platform, duration, aspect ratio, subtitle format, style, watermark, and closing call to action. If no authorized voice sample is available, the workflow will use an appropriate stock voice by default rather than cloning a voice without permission.
Method 2: Initialize a Standard Task Directory
For projects that require configuration reviews, task records, or staged execution, use the repository scripts to create a standard task directory first.
Step 1: Set the Skill Path
SKILL_DIR=~/.codex/skills/lanshu-create-ai-presenter-video
Step 2: Initialize the Task
python3 "$SKILL_DIR/scripts/init_job.py" \
--job-dir ~/Videos/my-presenter-video \
--presenter-image ~/Pictures/presenter.png \
--topic "Video topic" \
--duration 60 \
--aspect 9:16 \
--rights-confirmed \
--adult-presenter-confirmed
This example creates a 60-second task with a 9:16 aspect ratio. --rights-confirmed indicates that the rights to use the presenter image have been confirmed, while --adult-presenter-confirmed indicates that the person in the reference image has been confirmed as an adult.
Use the corresponding confirmation parameters only after authorization and presenter status have genuinely been verified. The permission details in the task configuration should still be reviewed before any remote upload.
Step 3: Complete job.json
After initialization, open the task directory and review the generated job.json. Complete the manual review fields and remote upload permissions according to the actual task. Do not submit unauthorized presenter images, voice samples, or other assets to remote services.
Step 4: Run the Preflight Check
python3 "$SKILL_DIR/scripts/preflight.py" ~/Videos/my-presenter-video/job.json
The preflight check validates the task configuration before formal generation begins. Resolve any issues it identifies before proceeding to video, speech, or lip-sync stages that may incur charges.
Understand the Complete Production Workflow
- Receive the topic or script and presenter image: Confirm that the assets are usable and verify image authorization and the presenter's adult status.
- Finalize the script and complete voiceover: Confirm the final narration before generating the full voiceover.
- Create a low-cost presenter test clip: Validate the presenter result and technical approach before large-scale paid generation.
- Generate the continuous primary presenter footage: Produce the main visuals based on the approved voiceover and presenter settings.
- Edit against the same audio timeline: Ensure that presenter footage, subtitles, shots, keywords, and transitions share a common time reference.
- Add subtitles, keyword animations, and a cover image: Reinforce key information and complete the publishing package.
- Perform quality assurance: Check lip-sync, presenter consistency, audio, and visual quality.
- Complete delivery: Export the master file, share-ready version, and QA report.
The complete voiceover is the workflow's key reference. Finalizing the audio first and arranging video and subtitles around it helps reduce lip-sync drift and continuity problems between segments.
Default Output Settings
If no more specific requirements are provided, the workflow uses the following defaults:
- Vertical
9:16,1080×1920resolution, and30fpsframe rate. - Videos generated from a topic are typically kept between 45 and 75 seconds.
- An appropriate stock voice is used when no authorized voice sample is available.
- The structure usually includes a clear opening, two to four content beats, and a concise ending.
- Music and promotional calls to action are added as needed.
- The target loudness for publishing is approximately
-16 LUFS.
For horizontal video, explicitly specify 16:9, as shown in the conversation example. For short-form video platforms, keep the default 9:16 or set it explicitly during task initialization with --aspect 9:16.
Advanced Usage Tips
Finalize the Voiceover Before Arranging Visual Elements
Do not create unrelated timelines for subtitles, shots, and presenter segments. Generate and approve the complete voiceover first, then align subtitles, presenter footage, keyword animations, and transitions to the same audio to improve overall synchronization.
Create a Test Clip Before Paid Generation
Before entering paid generation for the first time, clearly specify the uploaded content, generation duration, pricing basis, test-clip plan, and retry limit. A low-cost test clip can reveal issues with presenter appearance, lip-sync quality, or tool compatibility early, avoiding the waste of generating a complete video immediately.
Resume Existing Tasks First After an Interruption
If a remote task is interrupted, do not immediately resubmit it. First query the existing task ID to determine whether the task is still running or has already produced a result, thereby avoiding duplicate charges. After three consecutive paid candidates fail, stop generating additional candidates and summarize the issues.
Separate Technical Validation from Manual Review
Technical checks can detect file- and media-level issues, while manual review should assess whether the lip-sync looks natural, the presenter remains consistent, the voice is appropriate, the subtitles are readable, and the visuals meet publishing requirements. The relevant validation and failure recovery rules are located in references/qa-recovery.md.
Protect Task Records and Asset Privacy
- Do not write API keys, access tokens, or signed download URLs to the repository.
- Remove credentials and temporary URLs before committing task-level request records.
- Preflight and delivery reports should retain filenames only, without including absolute paths from the developer's machine.
- Confirm the rights to use presenter images before uploading them remotely.
- Confirm voice authorization before using voice-cloning capabilities.
Conclusion
lanshu-create-ai-presenter-video divides AI presenter video production into standardized, reviewable, and recoverable stages, using a unified voiceover timeline to reduce lip-sync and editing continuity issues. After installing the Skill, you can start a task directly in a Codex conversation or initialize job.json and run the preflight check first for stricter control over authorization, costs, generation, and final quality.
The project is distributed under the MIT License, allowing free use, modification, and distribution. Feedback can be submitted through GitHub Issues, while improvements to the workflow, compatibility, and quality checks can be contributed through Pull Requests.
