One person can connect AI tools for text, images, voice, animation, and slides to make multimedia content that once required a small team.
Raymond’s explanation
In the past, “making a video” meant needing at least four people: a copywriter, designer, voice actor, and editor. With this toolchain, one person can start with a Markdown file and produce illustrations, voice-over, animation, and even a complete instructional video.
The key is not using the strongest tool at every step, but “mixing and matching”—using the tool best suited to each stage and connecting them with AI. It is like building with LEGO: no single brick is special; the combination is what makes the work.
Current toolchain combination (2026-04 snapshot)
| Stage | Tool | Reason for selection | Status |
|---|---|---|---|
| Chinese illustrations | Gemini (generate_image.py) | Most reliable Chinese text rendering | ✅ Stable |
| Scenario photos | Z-Image Turbo (HuggingFace) | Largest free allowance and fast | ✅ Stable |
| Speech synthesis | MiniMax speech-2.8-hd | Most natural Chinese; $1/month | ✅ Stable |
| Animated B-roll | Remotion (React) | Strongest control over branding | ⚠️ Templates continue to expand |
| Slides | md-to-slides / Slidev | Markdown → HTML slides | ⚠️ Voice and automation integration pending |
| Video editing | YouTube Clipper + FFmpeg | Subtitle segmentation and compositing | ✅ Stable |
| Transcription | mlx-whisper (local) | Offline, free, accelerated on M-series chips | ✅ Stable |
Mix-and-match principles
- Chinese text rendering → always Gemini (Chinese text in the FLUX series is completely garbled)
- Realistic photos → Z-Image (free), or Gemini when precise control is needed
- Voice → MiniMax (Traditional Chinese input needs OpenCC conversion to Simplified Chinese; otherwise it produces a Hong Kong accent)
- Do not aim for a single platform to do everything; aim for the best value at each stage
Immature stages
- Slides-to-video automation: the complete pipeline from md-to-slides → voice-clone → compositing into a video is not yet stably connected
- B-roll template library: the number of templates is currently limited and needs to expand as instructional videos are produced
- Automated video chapter cards: Remotion can already generate chapter cards, but integration with the editing process still needs improvement
Where it has been discussed
Articles
- Comparison of Cowork vs NotebookLM vs Code (2026-04)
- Teaching CLAUDE.md and SKILL for Claude Code (2026-04)
Implementation records
- 2026-04-06: Compared three HF models and established a Gemini + Z-Image mix-and-match strategy
- 2026-04-06: Completed development of voice-tool and launched MiniMax API
- 2026-04-04: ui-gallery CIS v5.0, systematizing the brand visual system
Related concept pages
- AI Tool Applications — This card is a core practical case for that topic page
- Automation — Connecting the toolchain itself demonstrates automation thinking
Related concepts
Tools Extend Thinking, Automation, The Capability Compound-Interest Flywheel