Synchronized Arabic ASR comparison

Cohere Transcribe vs Wit.ai through Tafrigh

Open source video
00:00 / 1:09:21
Cohere

ابدأ التشغيل أو حرّك المؤشر لعرض النص.

Wit.ai / Tafrigh

ابدأ التشغيل أو حرّك المؤشر لعرض النص.

Comparable display grouping: both sides use Tafrigh 1.7.8's exact compact rule: merge consecutive short cues until at least 30 whitespace-delimited words. Words are not corrected, normalized, or rewritten. Cue timing remains engine-specific.
0 results

Cohere Transcribe

Local 2B model

Wit.ai through Tafrigh

Cloud API

What is being compared

The identical 4,161-second lecture. Cohere uses the canonical production-gate transcript; Wit.ai uses a fresh stock Tafrigh 1.7.8 run with its default 15-second maximum region and eight independent Arabic apps.

Timing provenance

Cohere cue positions are the fast segment-timing approximation distributed over retained Silero speech spans. Wit positions come from Tafrigh's Auditok regions. These are navigation aids, not a timestamp ground truth.

Interpretation boundary

This page supports listening and qualitative comparison. The lecture has no human reference transcript, so it does not provide WER. The separate benchmark report contains MSA, dialect, and Classical Arabic WER.

Artifact provenance and checksums