Videocaptioner

🎬 卡卡字幕助手 | VideoCaptioner - 基于 LLM 的智能字幕助手 - 视频字幕生成、断句、校正、字幕翻译全流程处理!- A powered tool for easy and efficient video subtitling.

Last verified:

Visit Videocaptioner

What is Videocaptioner?

VideoCaptioner (卡卡字幕助手) is an open-source AI-powered video subtitle tool that automates the entire subtitle creation workflow. It uses Whisper for speech recognition and large language models (LLM) for intelligent subtitle segmentation, correction, optimization, and translation. The tool supports 99 languages and can process a 14-minute video in just 4 minutes at a cost of less than $0.002.

Key features include: automatic speech-to-text transcription with multiple ASR engines (online free APIs and local Whisper models), LLM-powered intelligent subtitle segmentation for natural reading flow, AI subtitle correction for typos and terminology, high-quality contextual translation with reflection methodology, support for multiple subtitle formats (SRT, ASS, VTT, TXT), hard and soft subtitle embedding, batch video processing, VAD (voice activity detection), vocal separation, word-level timestamps, multi-platform video download (YouTube, Bilibili, TikTok, Douyin, Xiaohongshu), and an intuitive subtitle editing interface with real-time preview.

VideoCaptioner is designed for content creators on YouTube, Bilibili, and other video platforms who need professional subtitles without expensive software or high-end hardware. It works well for beginners (no GPU required, simple drag-and-drop interface) and experienced users (supports custom LLM APIs, local offline processing, advanced configuration options). The tool is ideal for multilingual video publishing, accessibility compliance, educational content, courses, tutorials, and anyone wanting to add professional subtitles efficiently.

Videocaptioner pricing

Pricing model: Freemium

VideoCaptioner is 100% free and open source under GPL-3.0 license. No paid plans or subscriptions exist. The software itself is completely free to download and use. Basic features like speech recognition using free online APIs (B接口, J接口) and Microsoft/Google translation require no API keys or configuration. For LLM features (subtitle optimization, correction, advanced translation), users need to configure their own LLM API key (OpenAI, DeepSeek, SiliconCloud, Ollama, etc.). Processing costs depend on user's API provider - approximately ¥0.01 (<$0.002) for a 14-minute video using OpenAI's official pricing. The project also offers an LLM API relay station (api.videocaptioner.cn) with high concurrency and cost-effective models, but using it is optional.

Videocaptioner pros

  • Open source and free to use with no subscription fees
  • Processes 14-minute video in just 4 minutes (lightning fast)
  • Ultra low cost - less than $0.002 per 14-minute video
  • No GPU required - runs on standard hardware
  • Supports 99 languages for speech recognition
  • Multiple ASR engines including free online APIs (B接口, J接口)
  • Local Whisper model support for privacy and offline use
  • LLM-powered intelligent subtitle segmentation for natural reading
  • AI subtitle correction for typos, punctuation, terminology
  • Contextual translation with reflection methodology for quality
  • Word-level timestamps for precise subtitle timing
  • VAD (voice activity detection) reduces hallucinations
  • Vocal separation improves audio quality in noisy videos
  • Batch video processing for multiple files
  • Multiple subtitle format exports (SRT, ASS, VTT, TXT)

Videocaptioner cons

  • Requires LLM API configuration for optimization and translation
  • Built-in public model is unstable (user should provide own API)
  • macOS users need to install Homebrew and Xcode tools first
  • WhisperCpp model is unstable according to documentation
  • Large-v3 Whisper model may have hallucination/repetition issues
  • Free online ASR APIs only support Chinese and English
  • Lower concurrency with some API providers requires thread adjustment
  • Soft subtitles require compatible player (like PotPlayer) to display
  • Translation quality depends on chosen service (Microsoft translation mediocre)

Frequently asked questions about Videocaptioner

What is VideoCaptioner?

VideoCaptioner (卡卡字幕助手) is an open-source AI-powered video subtitle processing tool based on large language models (LLM). It supports the complete workflow of speech recognition, subtitle segmentation, optimization, correction, translation, and video synthesis. The tool works on Windows, macOS, and Linux, requiring no high-end hardware or GPU for basic operation

Is VideoCaptioner free to use?

Yes, VideoCaptioner is completely free and open source under GPL-3.0 license. The software itself has no subscription fees or paid plans. Basic features like free online speech recognition APIs and basic translation work without any API keys. You only need to provide your own LLM API key if you want to use advanced features like LLM-based subtitle optimization and translation

What languages does VideoCaptioner support?

VideoCaptioner supports 99 languages for speech recognition using local Whisper models (WhisperCpp and fasterWhisper). Free online APIs (B接口, J接口) only support Chinese and English. For translation, it supports English, Simplified Chinese, Traditional Chinese, Japanese, Korean, Cantonese, French, German, Spanish, Russian, Turkish, and Portuguese

Do I need a GPU to use VideoCaptioner?

No, VideoCaptioner does not require a GPU. It can run on standard hardware without GPU. The tool supports both online API calling and local offline processing. While GPU support (CUDA) is available for faster Whisper transcription with fasterWhisper, it's optional and not required for basic operation

How fast is VideoCaptioner?

VideoCaptioner can process a 14-minute 1080P video in just 4 minutes. This includes full workflow: local Whisper transcription, GPT-4o-mini optimization and translation to Chinese, and video synthesis. The processing speed is lightning fast powered by Whisper + LLM stack

What subtitle formats does VideoCaptioner export?

VideoCaptioner supports exporting subtitles in multiple formats including SRT, ASS, VTT, and TXT. It also supports embedding both hard subtitles (burned into video) and soft subtitles (not burned, requires compatible player)

How do I configure LLM API for VideoCaptioner?

For LLM features (subtitle segmentation, optimization, translation), you need to configure an LLM API in settings. Options include SiliconCloud, DeepSeek (recommended deepseek-v3), Ollama (local), OpenAI-compatible interfaces, or the project's built-in public model (gpt-4o-mini, but unstable). You can also use the project's API relay station at api.videocaptioner.cn with BaseURL: https://api.videocaptioner.cn/v1 and your API key from the personal center

What is the difference between hard and soft subtitles?

Hard subtitles are burned directly into the video file and will display on any player. Soft subtitles are not burned into the video - they're separate subtitle tracks that require compatible players (like PotPlayer) to display. Soft subtitles processing is extremely fast but the styling will be the player's default white style, not the custom style set in VideoCaptioner

Does VideoCaptioner support batch processing?

Yes, VideoCaptioner supports batch video subtitle processing. This feature allows you to process multiple video files simultaneously, significantly improving efficiency. The latest version also includes batch subtitle functionality for faster results without compromising quality

What are the latest features in VideoCaptioner?

The latest version includes VAD (voice activity detection) to filter non-voice segments and reduce hallucinations, MDX-Net vocal separation for noise reduction and audio quality improvement, word-level timestamps for precise subtitle timing, and batch subtitle processing. These features significantly enhance the subtitle generation process

Categories

Use cases

Browse all AI tools on NeedAnAI