For the complete documentation index, see llms.txt. This page is also available as Markdown.

πŸ“Auto Captions

Select "Auto Captions" to instantly generate animated captions.

Characters per caption

Choose how many characters appear in each caption.

Unique settings are saved for each Sequence aspect ratio: Landscape, Portrait, and Square.

Language

Select Auto to auto-detect the spoken language or choose a transcription language.

Selecting a transcription language speeds up processing time.

Single-word captions

Check this box to force Auto Captions to always split by word.

Also generate a Caption Track

Check this box to also create a new Caption Track.

Model

Choose from 6 transcription models with varying speed & accuracy (default Tiny).

Name
Size
Description

Tiny

41.5MB

Fastest model with basic accuracy.

Base

78MB

Balanced model offering good accuracy and reasonable speed.

Small

252.2MB

High accuracy model with moderate speed.

Medium

785.2MB

Very high accuracy model with slower processing but excellent results.

Large V3 Turbo

833.7MB

Most accurate text transcription with slower processing time. Recommended for non-English languages.

Large V3

2.9GB

Highest accuracy model. Best for offline/batch transcription where speed is not critical.

Transcription models are saved locally to your machine. Delete and redownload models at any time.

VAD (Voice Activity Detection)

Choose from 4 presets for precise speech detection: Auto, Podcast, Monologue, Fast-paced. Choose Custom to dial in your own advanced controls.

Processing Details

Captioneer lists the available GPU and models on your machine. You can toggle between Automatic, GPU, or CPU processing.

Last updated