I am a computational linguist and Hausa NLP/AI data specialist working on a Hausa LLM Healthcare & Safety Preference Evaluation Dataset.
The dataset is developed collaboratively by computational linguistics specialists and healthcare professionals, including doctors, nurses, and community healthcare workers. Healthcare experts create and review the raw healthcare content, while computational linguistics specialists handle Hausa linguistic review, annotation, adaptation, and LLM evaluation/fine-tuning.
I am interested in connecting with researchers and developers working on multilingual and low-resource healthcare AI, particularly those interested in Hausa-language evaluation and safety data.
I would be happy to share more information and discuss potential research, collaboration, or dataset licensing opportunities.
Great initiative — healthcare preference & safety datasets in low‑resource languages are extremely valuable, especially when the linguistic review and the medical content are both handled by domain specialists.
A few feedback points that may help future iterations:
• Schema: a clear separation between raw healthcare content, linguistic adaptation, and safety‑preference evaluation is important. If the schema keeps these layers distinct, it becomes easier to run cross‑language comparisons or plug the Hausa data into multilingual evaluation pipelines.
• Provenance: the fact that medical professionals generate and review the source material is a strong foundation. Explicit provenance fields (author role, medical domain, review stage) make the dataset more reliable for downstream safety benchmarks.
• Ergonomics: if the dataset includes structured preference labels (e.g., refusal, safe alternative, escalation, uncertainty), it becomes easier to integrate into evaluation frameworks for LLM alignment and healthcare safety.
If you plan future versions, adding a small “error taxonomy” for unsafe or ambiguous outputs could help researchers build more robust Hausa‑language safety evaluators.
Happy to discuss more if you expand the schema or release additional healthcare subsets.
Thank you very much for the detailed feedback. I really appreciate the suggestions, especially around schema separation, provenance, structured safety-preference labels, and error taxonomy.
These are very helpful for strengthening the dataset for multilingual healthcare and safety evaluation. I would be interested in discussing how I can incorporate these improvements and potentially collaborate on future versions.
If useful, I would also be happy to share more details about the current schema and sample data for further discussion.
Thanks for your kind reply and for being so open to the feedback.
I’m glad the points on schema separation, provenance, and the safety-preference / error taxonomy were useful. My goal was exactly to help make the dataset more robust for multilingual healthcare evaluation.
I’d be very happy to continue the discussion and collaborate on the next version. If you can share the current schema and a small sample of the data (20-30 examples), I can take a look and give you more concrete suggestions on how to structure it.