← All projects

Evaluation · Open source

Swiss German Voice Benchmark

This benchmark picks ETH SwissDial recordings across eight Swiss German dialects, runs voice-capable models on two tasks (verbatim dialect transcription and translation to High German), and lets you inspect every transcript with word-level differences. It supports different prompting strategies, ensembles, speaker-weighted scores based on how many people speak each dialect, and exports of every saved run.

My role
Creator and builder
Started
Updated
Visibility
Open Source
Built with
Python · Microsoft Foundry · OpenAI Realtime · SwissDial · WER / CER / chrF · React
SwissDial clips, 25 per dialect
200
Swiss German dialects compared
8
best paired word match (Realtime 2.1, High German)
68.6%
Swiss German speech lab benchmark setup
Benchmark setup: choose a task, prompting strategy, dialects, utterances, and the models to compare.

Problem

Speech models are usually evaluated on standard languages. For Swiss German, teams have little evidence about which model and prompting strategy actually works, per dialect.

Solution

A reproducible benchmark with balanced dialect samples, multiple tasks and strategies, paired comparisons, and a UI to inspect individual mistakes instead of trusting one headline number.

Architecture

A local API and dashboard orchestrate model runs over a selected set of SwissDial clips, stream results, store every transcript and timing in a local database, and compute paired metrics so models are only compared on clips both handled successfully.

Lessons learned

Single numbers hide a lot. Dialect-level breakdowns, paired comparisons, and the ability to read individual errors changed my conclusions more than any aggregate score.

More screenshots (3)
Saved 200-utterance model comparison
A saved 200-clip comparison between Realtime 2 and Realtime 2.1, with model and dialect bars.
Model results grouped by dialect
Results grouped by dialect and ordered by number of speakers.
Searchable result table with word-level differences
Every transcript can be inspected with reference and output differences, sorted by the lowest word match.

speech-ai · evaluation · swiss-german · benchmarks