AI roundup
Gemini 3.5 Transcribe ships
Thu 27 August 2026
85-language STT with 5.5% WER and inline disfluency editing
| Streaming WER | 5.5% (Google FLEURS) / 4.0% (Artificial Analysis) |
| Real-time Factor | 79.6x (audio sec / proc sec) |
| Pricing | $5.00 per 1,000 min |
| Languages | 85+ |
| Speaker Diarization | Up to 3 speakers (pre-recorded) |
| Parameters | Not disclosed |
Key takeaways
- Filler removal: Strips 'ums' and self-corrections during transcription rather than post-processing
- Latency cut: 70% faster voice-to-final-text vs Chirp 3; 79.6x throughput on batch audio
- Custom vocab: Supports specialized jargon and alphanumeric entities (order IDs, postal codes)
Source: Ars Technica AI