PRiSM

Open suite for phone recognition: phonetic-feature error on IPA transcripts plus clinical, L2, and multilingual probes of speech models.

Also known as: Phone Realization in Speech Models

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorymultimodal
Subcategoryphone recognition and phonetic downstream probes
Page statusactive
MetricPFER (intrinsic); task F1 / Recall@1 (extrinsic)
Directionlower_is_better
PublisherChangeling Lab / Carnegie Mellon and collaborators

What it measures

PRiSM tests whether a speech model hears phones, not words. Intrinsic tasks ask for an IPA transcript of an utterance and score articulatory-feature edits against gold phones (PFER), rather than token-level phone error rate. Extrinsic tasks reuse those transcripts or hidden states on dysarthria intelligibility, atypical child speech, L1 classification, L2 assessment, language id, geolocation, and phone-inventory induction. The paper argues that transcription error alone hides failures on clinical and sociophonetic work.

Task format

Audio in, IPA transcript and/or frozen representation out. Hydra configs under github.com/changelinglab/prism. Kaldi-style test sets for PR; Hugging Face repos for downstream probes.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub