G-MASS: Ghana Medical AI Safety Screen v1.1.0
Open cross-lingual safety evaluation for medical AI in Ghanaian languages.
| Medical query | Language | Model to evaluate | Failure category |
|---|
Upload CSV/JSONL probes. Files with language or language-specific prompt columns are evaluated per row/column; the dropdown is only a fallback for plain prompt files.
Scored results
This tab reads real combined outputs when available. It does not display placeholder benchmark claims.
Model profiles
gemini-2.5-flash | 55.56 | 0 | 0 | 100 | 50 | 55.56 | 55.56 | not_ready |
gemini-2.5-flash | 0 | 0 | 0 | 50 | 0 | 0 | 0 | not_ready |
gpt-4o | 55.56 | 0 | 0 | 100 | 50 | 55.56 | 55.56 | not_ready |
G-MASS: Ghana Medical AI Safety Screen
G-MASS evaluates whether medical AI assistants respond safely across English, Ghanaian English, and Twi. The app is a public interface over the same pipeline used by the repository CLI.
Scorer identities:
- LlamaGuard3: primary scorer for English and Ghanaian English.
- Gemma: secondary cross-validator for English and Ghanaian English.
- AfroLM: primary scorer for detected Twi responses.
- LlamaGuard3 also cross-validates detected Twi after Khaya back-translation.
gemini is an evaluated model key. SCORER_BACKEND=policy_api is a scorer
runtime option that may call Gemini API to execute policy prompts, but Gemini is
not counted as a scorer identity.
Outputs are preliminary evaluation evidence, not deployment certification for clinical care.