Real-Time Speech Translation Latency Analysis
Empirical objective audio-visual latency evaluation measuring end-to-end performance from speech onset to target-language caption visibility (TTFC) and synthesized audio playback (TTFS). Conducted under IEEE 829 software performance testing guidelines.
Exbabel is up to 10.2× faster than Wordly during continuous speech.
While Wordly pauses and buffers audio for up to 7.22 seconds waiting for the speaker to stop talking, Exbabel streams live translated speech continuously in under 2 seconds without interruptions.
Streaming Architecture Prevents Continuous Speech Audio Lag
Wordly holds translated audio in a silence-detection buffer for up to 7.22 seconds during continuous speech. Exbabel streams live translated speech continuously in ~2 seconds, eliminating delays during keynotes, sermons, and lectures.
EXB-LAB-2026-001: Latency Benchmark Report
Complete 11-section research report including FFmpeg astats RMS profiles, frame extraction logs, and mathematical scaling formulas.
Machine-Readable Trial JSON Datasets
Access raw 30 FPS frame indices, RMS decibel logs, and runnable Python analyzer scripts for independent audit and verification.
Executive Head-to-Head Benchmark Findings
Wordly mean TTFC: 1.974s. Exbabel renders target Spanish captions in ~1.0s from speech onset (0.403s pure processing time).
Wordly mean TTFS: 5.680s. Exbabel begins audio playback in ~2.0s (1.417s pure processing time).
Wordly continuous speech lag: 7.220s. Exbabel maintains constant ~2.0s streaming delay without sentence buffering.
Continuous Speech Audio Latency Projections
Linear scaling formula: Wordly TTFS scales with speech duration (TTFS ≈ Speech_Duration + 0.4s) while Exbabel remains constant (TTFS ≈ 2.027s).
| Speech Duration | Wordly Audio Delay (TTFS) | Exbabel Audio Delay (TTFS) | Exbabel Advantage |
|---|---|---|---|
| 6.82 s (Measured Trial) | 7.220 s | 2.027 s | 3.6× Faster |
| 10.00 s Continuous Speech | ~10.400 s | ~2.000 s | ~5.2× Faster |
| 20.00 s Continuous Speech | ~20.400 s | ~2.000 s | ~10.2× Faster |
| 30.00 s Continuous Speech | ~30.400 s | ~2.000 s | ~15.2× Faster |
| 60.00 s Continuous Speech | ~60.400 s | ~2.000 s | ~30.2× Faster |
Testing Standards & Protocols Followed
Standard for Software & System Test Documentation
Governs master test planning, trial logging, apparatus calibration, anomaly reporting, and formal summary documentation under EXB-LAB-2026-001.
Software Quality Requirements and Evaluation (SQuaRE)
Establishes standard performance efficiency models, measuring system time-behavior, response latency, and resource scaling.
Relative Timing of Sound & Vision Signal Processing
Defines objective frame-accurate video demuxing (30.00 FPS) and acoustic energy RMS window segmentation for A/V alignment.
Algorithmic Energy & Silence Boundary Detection
Measures Root Mean Square (RMS) energy shifts from noise floor (-60 dB) to speech peak (-34 dB) at 21.3ms window resolution.
Download Official Reports & Machine-Readable Datasets
| Resource Title | Format | Size | Action |
|---|---|---|---|
EXB-LAB-2026-001 Executive Report (PDF)OFFICIAL REPORT | PDF Document | 39 KB | Download↓ |
EXB-LAB-2026-001 Markdown Source (MD)SOURCE CODE | Markdown Whitepaper | 18 KB | Download↓ |
Exbabel Trial Latency Dataset (JSON)RAW DATA | JSON Dataset | 12 KB | Download↓ |
Wordly Trial Latency Dataset (JSON)RAW DATA | JSON Dataset | 14 KB | Download↓ |
Download Official PDF Report & Whitepaper
Get full access to EXB-LAB-2026-001 including mathematical models, RMS decibel logs, PDF document, and raw datasets.