G-MASS: Ghana Medical AI Safety Screen
Open cross-lingual safety evaluation for medical AI in Ghanaian languages.
Language
Model to evaluate
Failure category
Result will appear here.
Examples
| Medical query | Language | Model to evaluate | Failure category |
|---|
Upload a CSV with required columns probe_id and prompt.
Model to evaluate
Language
Scored results
This tab reads real combined outputs when available. It does not display placeholder benchmark claims.
Model profiles
No combined results found at data/eval_outputs/combined/all_models_scored.jsonl |
No combined results found at data/eval_outputs/combined/all_models_scored.jsonl |
G-MASS: Ghana Medical AI Safety Screen
G-MASS evaluates whether medical AI assistants respond safely across English, Ghanaian English, and Twi. The app is a public interface over the same pipeline used by the repository CLI.
Scorer identities:
- LlamaGuard3: primary scorer for English and Ghanaian English.
- Gemma: secondary cross-validator for English and Ghanaian English.
- AfroLM: primary scorer for detected Twi responses.
- LlamaGuard3 also cross-validates detected Twi after Khaya back-translation.
gemini is an evaluated model key. SCORER_BACKEND=policy_api is a scorer
runtime option that may call Gemini API to execute policy prompts, but Gemini is
not counted as a scorer identity.
Outputs are preliminary evaluation evidence, not deployment certification for clinical care.