G-MASS: Ghana Medical AI Safety Screen v1.1.0
Open cross-lingual safety evaluation for medical AI in Ghanaian languages.
| Medical query | Language | Model to evaluate | Failure category |
|---|
Upload JSONL or CSV probe datasets. Files with bilingual columns (e.g. english_prompt, twi_prompt, ghanaian_en_prompt, source_standard_english, final_approved_twi) or single prompt columns are automatically parsed across all rows.
Scored results
This tab reads real combined outputs when available. It does not display placeholder benchmark claims.
Model profiles
gemini-2.5-flash | 55.56 | 0 | 0 | 100 | 50 | 55.56 | 55.56 | not_ready |
gemini-2.5-flash | 0 | 0 | 0 | 50 | 0 | 0 | 0 | not_ready |
gpt-4o | 55.56 | 0 | 0 | 100 | 50 | 55.56 | 55.56 | not_ready |
Personalisation & Local Execution Settings (Vision §7)
Auto detects RAM/GPU or forces a specific judge tier
G-MASS: Ghana Medical AI Safety Screen
G-MASS evaluates whether medical AI assistants respond safely across English, Ghanaian English, and Twi. The app is a public interface over the same pipeline used by the repository CLI.
Scorer identities:
- LlamaGuard3: primary scorer for English and Ghanaian English.
- Gemma: secondary cross-validator for English and Ghanaian English.
- AfroLM: primary scorer for detected Twi responses.
- LlamaGuard3 also cross-validates detected Twi after Khaya back-translation.
gemini is an evaluated model key. SCORER_BACKEND=policy_api is a scorer
runtime option that may call Gemini API to execute policy prompts, but Gemini is
not counted as a scorer identity.
Outputs are preliminary evaluation evidence, not deployment certification for clinical care.