Evaluation · Open source
Swiss German Voice Benchmark
This benchmark picks ETH SwissDial recordings across eight Swiss German dialects, runs voice-capable models on two tasks (verbatim dialect transcription and translation to High German), and lets you inspect every transcript with word-level differences. It supports different prompting strategies, ensembles, speaker-weighted scores based on how many people speak each dialect, and exports of every saved run.
- My role
- Creator and builder
- Started
- Updated
- Visibility
- Open Source
- Built with
- Python · Microsoft Foundry · OpenAI Realtime · SwissDial · WER / CER / chrF · React
- SwissDial clips, 25 per dialect
- 200
- Swiss German dialects compared
- 8
- best paired word match (Realtime 2.1, High German)
- 68.6%

Problem
Speech models are usually evaluated on standard languages. For Swiss German, teams have little evidence about which model and prompting strategy actually works, per dialect.
Solution
A reproducible benchmark with balanced dialect samples, multiple tasks and strategies, paired comparisons, and a UI to inspect individual mistakes instead of trusting one headline number.
Architecture
A local API and dashboard orchestrate model runs over a selected set of SwissDial clips, stream results, store every transcript and timing in a local database, and compute paired metrics so models are only compared on clips both handled successfully.
Lessons learned
Single numbers hide a lot. Dialect-level breakdowns, paired comparisons, and the ability to read individual errors changed my conclusions more than any aggregate score.
More screenshots (3)



speech-ai · evaluation · swiss-german · benchmarks