A 50-item BIG-bench task asking a model to grade suicide risk in short Reddit-derived posts on a four-level scale against expert labels.
unassessed
| Category | safety |
|---|---|
| Subcategory | clinical-style risk triage from short first-person text (BIG-bench task) |
| Page status | active |
| Metric | multiple_choice_grade (accuracy against expert-assigned risk level) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 50 |
| Dataset licence | Apache License 2.0 (BIG-bench repository licence; applies to the task code and data as distributed) |
| Publisher | BIG-bench collaboration (Google-led); task authored by named contributors below |
This task gives a model a short, first-person piece of text -- style and typos preserved -- and asks it to classify the author's suicide risk into one of four expert-defined levels: no risk, low risk, moderate risk, or severe risk. The texts were manually pulled from a public Reddit submission corpus and de-identified; some "no risk" items deliberately include emotionally charged or clinically sensitive keywords so that a model cannot pass by keyword-matching alone and must weigh context. The task's own documentation is explicit that it is a research probe of language understanding, not a validated screening tool, and that suicide-risk assessment "deserves careful and thoughtful research."
Four-way multiple choice: given a short first-person text, classify author suicide risk as no/low/moderate/severe, scored zero-shot, one-shot, and many-shot.
No model card in ModelSpec reports this benchmark yet.