AI Mode
๐ Local (Browser) โ Free, private, works offline
โ๏ธ Cloud (OpenRouter) โ Higher quality, requires API key & internet
๐ Auto โ Cloud when available, local fallback
Controls both transcription and chat. Cloud mode uses your OpenRouter API key. Auto mode tries cloud first, falls back to local.
๐๏ธ Transcription Model
โญ whisper-medium (~1.7GB) โ Best Arabic quality (WebGPU)
๐ฏ whisper-small (~510MB) โ Fallback / no WebGPU
โญ whisper-medium (q4): 7/7 Arabic key words correct. ~680MB download (one-time). WebGPU required (Chrome/Edge 113+). On non-WebGPU browsers, falls back to whisper-small (510MB, fewer Arabic errors).
๐ฌ Chat LLM Model
๐ฑ Bonsai 1.7B (~200MB) โญ Default
๐ฟ Ternary Bonsai 1.7B (~300MB)
Qwen3 1.7B (~1.2GB) โ Fast
SmolLM2 1.7B (~1.1GB) โ Fastest
Llama 3.2 1B (~0.8GB) โ Basic
Qwen3 4B (~2.6GB) โ Best balance
Llama 3.2 3B (~2.1GB)
Qwen2.5 3B (~2.1GB)
Qwen3 8B (~5.2GB)
Llama 3.1 8B (~5.0GB)
Qwen3.5 4B (~2.8GB)
Runs in your browser via WebGPU. No API key needed. Cached after first download.
โ
All browser AI features are free and private . No data leaves your device.
OpenRouter API Key
Get your key at openrouter.ai/keys or click "Get Key" for instructions.
โ๏ธ Cloud Chat Model
Default (server chooses)
Qwen3.7 Plus โ Multimodal, vision+text โญ
Qwen3 235B โ Highest quality
GPT-4o โ Best OpenAI
GPT-4o Mini โ Fast & cheap
Claude Sonnet 4 โ Strong reasoning
Claude 3.5 Haiku โ Fast
Llama 4 Maverick โ Latest
Llama 3.3 70B
Mistral Large
Select "openrouter" as chat provider to use this model. Leave as "Default" to let OpenRouter choose.
๐๏ธ Cloud Transcription
๐ Microsoft MAI-Transcribe 1.5 โ Fastest, proper punctuation โญ
โ๏ธ OpenAI GPT-4o Transcribe โ Some punctuation ($0.014)
โ๏ธ OpenAI GPT-4o Mini โ Cheapest full transcript ($0.007)
โ๏ธ OpenAI Whisper 1 โ Original Whisper ($0.024)
โ ๏ธ OpenAI Whisper Large V3 โ May truncate long audio
๐ Microsoft MAI-Transcribe 1.5: 2690 chars / 4min. 3.6s. $0.024. Proper Arabic punctuation (ุ . ุ). Best quality.
No download needed. 4-min audio costs $0.007-$0.024. Upstream timeout: 60s per request.
๐ Subtitle Translation Model
DeepSeek V4 Flash โ $0.14/$0.28/M, ~4s latency โญ
DeepSeek V4 Pro โ $0.44/$0.87/M, 1.6T MoE
Gemini 2.5 Flash-Lite โ $0.10/$0.40/M, ultra-fast
Gemini 2.5 Flash โ $0.15/$0.60/M, strong multilingual
GPT-5 Nano โ $0.05/$0.40/M, cheapest OpenAI
Gemini 3.0 Flash โ $0.50/$3.00/M, frontier quality
GPT-4.1 Nano โ $0.10/$0.40/M, proven stable
Mistral Small 3.2 โ $0.06/$0.18/M, excellent European
Qwen3 30B A3B โ $0.10/$0.30/M, strong Asian
DeepSeek V3.2 โ $0.14/$0.28/M, great CJK
Flash models recommended for subtitles โ fast and cost-effective. Prices per 1M tokens via OpenRouter.
โก Audio Speed (for cloud transcription)
1.0x โ Full quality, no speed-up โญ Default
1.3x โ 23% smaller, 99.4% quality
1.4x โ 29% smaller, 99.7% quality
1.5x โ 33% smaller, ~99.6% quality
1.6x โ 37% smaller, 99.6% quality
1.7x โ 41% smaller, 99.6% quality
2.0x โ 51% smaller, 99.1% quality
1.0x (default): Full quality. 4-min audio = 7.3MB upload. No quality loss.
Speed-up reduces upload size and processing time. Higher speeds may slightly affect word accuracy for Arabic.
โ๏ธ Cloud models require an OpenRouter API key. Pay per use. Data is sent to OpenRouter servers for processing.