Open suite for phone recognition: phonetic-feature error on IPA transcripts plus clinical, L2, and multilingual probes of speech models.
unassessed
| Category | multimodal |
|---|---|
| Subcategory | phone recognition and phonetic downstream probes |
| Page status | active |
| Metric | PFER (intrinsic); task F1 / Recall@1 (extrinsic) |
| Direction | lower_is_better |
| Publisher | Changeling Lab / Carnegie Mellon and collaborators |
PRiSM tests whether a speech model hears phones, not words. Intrinsic tasks ask for an IPA transcript of an utterance and score articulatory-feature edits against gold phones (PFER), rather than token-level phone error rate. Extrinsic tasks reuse those transcripts or hidden states on dysarthria intelligibility, atypical child speech, L1 classification, L2 assessment, language id, geolocation, and phone-inventory induction. The paper argues that transcription error alone hides failures on clinical and sociophonetic work.
Audio in, IPA transcript and/or frozen representation out. Hydra configs under github.com/changelinglab/prism. Kaldi-style test sets for PR; Hugging Face repos for downstream probes.
No model card in ModelSpec reports this benchmark yet.