Frontline Community Triage Agent
Supporting community health workers (CHWs) with voice-driven Amharic-English triage documentation.
Clinician-reviewed voice documentation for multilingual care teams
Supporting community health workers (CHWs) with voice-driven Amharic-English triage documentation.
Automating clinical note-taking during doctor consultations with structured FHIR / EHR JSON output.
Click "Start Consultation Recording" to capture live audio, or "Load Consultation Demo" for a scripted example...
Record the live doctor-patient conversation; the chart auto-fills as fields are recognized.
Auto-extracted fields require clinician verification before committing to EHR.
Editable clinical draft. Streaming updates fill fields that have not been manually changed.
Simulating automated voice call check-ins post-discharge with automated risk escalation.
Post-care demo focus: Case 2 of 5 Gold Standard Cases.
Runs the clinical team's post-care protocol on real transcript text: recovery risk scoring plus, for maternal patients, the 5-level obstetric acuity scale (Annex 1-3).
Decision support from reported symptoms only -- a nurse/doctor must confirm with real vitals and exam.
Not a simulation: Intron speaks each question aloud, you answer by mic, Intron transcribes it, and the protocol engine decides the next question -- or escalates immediately if it hears an emergency. No LLM involved anywhere in this loop.
Press a simulation button on the left to initiate the voice call...
Recovery status is parsed from code-switched transcript in real time.
Monitoring outpatient recovery telemetry...
The matrix below is a fixture/demo view. It is not a real model ranking; use the measured validation summary and linked reports for performance claims.
The displayed matrix uses embedded hypotheses and illustrative case values. Its WER/CER values are fixture examples, not independent audio inference or proof of production superiority. The aggregate row is calculated live from the five gold standard demo cases below.
15-case Amharic-English benchmark: 34.91% normalized WER, 57.78% target recall, and 42.22% M-WER. Clinician review is mandatory; autonomous care is not supported.
Normalized transcript scores on the saved Amharic-English manifest outputs. Target recall is not a fairness metric.
| Baseline | Normalized WER | Target recall | M-WER |
|---|---|---|---|
| Google Gemini Flash | 11.46% | 93.33% | 6.67% |
| OpenAI Whisper Tiny | 99.56% | 23.67% | 76.33% |
| Meta Wav2Vec2 Base 960h | 108.23% | 2.22% | 97.78% |
Offline scoring uses stored hypotheses; it does not rerun model inference. Wav2Vec2 is an English-only baseline. See BENCHMARK_RESULTS.md for the metric method and limitations.
Factuality and safety are qualitative human-evaluation fields; no case-level ratings are supplied for this demo fixture.
| Speech AI Model | WER | CER | Factuality | Safety |
|---|
Clinician Review Notes:
Intron Sahara v2.5 maintains accurate phonetic grounding for local disease descriptions without hallucinatory translations.
Upload or record an audio sample in Amharic-English to evaluate live Intron Sahara v2.5 ASR performance.
| Provider | Status | Latency | WER | CER | Transcript |
|---|
Amharic-English v1.0 · Current prototype status and proposed priorities
Product direction
AfriHealth AI is a clinician-reviewed voice documentation prototype for community health teams working in Amharic-English. Intron Sahara v2.5 is the primary speech engine; clinicians remain responsible for reviewing transcripts and clinical outputs.
Captures reviewed audio, returns code-switched transcription, and presents a clinician-review triage summary.
Supports draft documentation workflows. SOAP processing accepts structured input or falls back to manual review; prose-to-SOAP/ICD-10 generation is not complete.
Supports follow-up simulations and highlights possible escalation signals for human assessment.
Amharic-English v1.0. Afaan Oromoo-English is Phase 2; Tigrinya-English is Phase 3.
Transcripts and generated clinical artifacts are drafts for qualified clinician verification.
Follow-up and triage support decisions; they do not diagnose, contact patients autonomously, or dispatch care.
This is a saved-transcript comparison, not fresh model inference or population-performance evidence. Gemini reports 11.46% normalized WER; Whisper Tiny 99.56%; English-only Wav2Vec2 108.23%. A separate 15-recording clinical review found 56.38% mean WER, 44.33% target-term recall, and critical-term misses in 6 cases. All results require clinician review.
Review benchmark methodology and limitationsNot ready for identifiable patient data: transcript retention/deletion controls, built-in clinician authentication and role-based access, and verified production gateway configuration remain open. Medication safety is not enforced end to end; medication details require independent clinician verification.
Provider credentials remain server-side, and private validation inputs stay outside Git. Complete clinician sign-off and backend safety controls before operational use.
Implement transcript retention/deletion and clinician authentication with role-based access; verify gateway authentication and exact production CORS origins before identifiable data use.
Validate transcript-to-SOAP and coding behavior, require sign-off, and implement a backend medication safety gate before presenting medication checks as enforced.
Complete missing annotator metadata, retain model/configuration details, and expand consented, de-identified Amharic-English evaluation while keeping saved transcripts distinct from fresh inference.
Confirm frontend/backend deployment and CORS together, refresh the demo against the current build, and complete the recording and team handoff checklist.