Illustrative pilot projection
~5 minutes
Average consultation time saved
Planning assumption per patient; not measured in a clinical study.
The authoritative manifest contains 100 cases (CS-01 to CS-100), but its source text and provisional labels are not audio-verified. Mock outputs are synthetic; performance claims require audio verification. Validated model metrics and CEAS are currently unavailable.
Illustrative workflow-value assumptions for discussion and pilot planning. These figures are not observed results, validated clinical outcomes, or findings from a clinical study.
Illustrative pilot projection
~5 minutes
Planning assumption per patient; not measured in a clinical study.
Illustrative pilot projection
~200 minutes
Illustrates 5 minutes across an assumed 40 consultations per workday; neither input is a measured outcome.
Illustrative pilot projection
~70%
Hypothesis for future evaluation; not yet validated or measured.
Illustrative pilot projection
~15%
Projected daily capacity contribution; definition and real-world impact require prospective measurement.
Prototype controls only; these badges are not production certification.
Not ready for identifiable patient data: authentication/RBAC, transcript retention and deletion controls, auditability, and production deployment configuration still require implementation and verification. Benchmark findings are limited to their documented dataset and methods.
Supporting community health workers (CHWs) with voice-driven Amharic-English triage documentation.
Automating clinical note-taking during doctor consultations with structured FHIR / EHR JSON output.
Click "Start Consultation Recording" to capture live audio, or "Load Consultation Demo" for a scripted example...
Record the live doctor-patient conversation; the chart auto-fills as fields are recognized.
Extracted fields are clinical reference suggestions and require clinician review before any record action.
Editable clinician-support draft. Streaming updates fill fields that have not been manually changed; all content requires clinician verification.
Review Required before FHIR Commit
Review Required before FHIR Commit
Prototype demonstration only. This state machine does not schedule or place a real call, contact a patient, or submit an alert to clinical staff.
Workflow Simulation: Patient Discharged / Intake Completed.
Clinician-support voice follow-up simulations with risk indicators for human review.
Post-care demo focus: illustrative fixture 2 of 5; separate from the clinical benchmark.
Applies the configured post-care protocol to supplied transcript text. Risk and acuity results are clinical reference suggestions and require clinician review.
Decision support from reported symptoms only -- a nurse/doctor must confirm with real vitals and exam.
Interactive voice session: Intron speaks each prompt, you answer by microphone, and the service transcribes the response. Returned questions and risk indicators are clinical reference suggestions for clinician review; this panel does not automatically contact clinical staff.
Press a simulation button on the left to initiate the voice call...
Clinical reference indicators are derived from the supplied transcript and require clinician review.
Monitoring outpatient recovery telemetry...
Requirement-to-evidence map for a fast review. Evidence claims below describe repository contents and prototype behavior; they are not a claim of clinical efficacy or production readiness.
| Requirement | Project Implementation | Evidence Location |
|---|---|---|
| Voice downstream clinical workflows | Browser audio capture and transcription, clinician-review triage and intake surfaces, FHIR-compatible export with review gate, and post-care API workflows. The added VoiceBot state machine is a separate local workflow simulation; it does not place or schedule calls. | Triage and intake modules; VoiceBot simulation; ARCHITECTURE.md; main.py routes `/api/v1/post-care/analyze`, `/api/v1/fhir/export`, `/api/v1/ehr/commit`. |
| Three-or-more speech model comparison | A six-adapter evaluation pipeline is wired to the 100-case master manifest. Current source transcripts are not independently audio-verified, and mock outputs are not accuracy evidence; no live comparison is represented as validated benchmark performance. | BENCHMARK_RESULTS.md; 100-case master manifest; inference_engine.py; evaluator.py. |
| Health category focus | Healthcare documentation and voice-support prototype focused on clinician-reviewed Amharic-English code-switching. SOAP/ICD output is not automatically generated from ordinary prose; clinical artifacts require human verification. | README.md; ARCHITECTURE.md; CLINICAL_RECORDING_PROTOCOL.md. |
| Responsible AI, safety, and human review | UI disclaimers and review badges; explicit client review confirmation before export; backend sign-off assertion required by FHIR/EHR endpoints. Reviewer identity/RBAC, auditability, retention/deletion, subgroup fairness, and clinical efficacy are not established by this prototype. | RESPONSIBLE_AI.md; tests/test_safety.py; index.html; main.py. |
| Working prototype and judge-ready demonstration | Single-page browser prototype with one-click impact dashboard, VoiceBot workflow simulation, challenge map, clinical modules, and benchmark view. Confirm the linked video matches this build before submission. | DEMO_SCRIPT.md; SUBMISSION_READINESS.md; README.md. |
Evidence note: the 100-case master manifest contains source text, not independently audio-verified gold transcripts. Mock outputs are synthetic and not model-performance evidence. Do not present the VoiceBot simulation or projected impact values as measured or live patient outcomes.
The matrix below uses five illustrative application fixtures, separate from the 100-case clinical benchmark. Its scores are synthetic examples and must not be used for model ranking or performance claims.
The displayed matrix uses embedded demonstration hypotheses and illustrative case values. Its WER/CER values are synthetic fixture examples, not independent audio inference or proof of production superiority. The five demo fixtures are separate from the authoritative CS-01 to CS-100 clinical manifest.
The active clinical corpus is benchmark/metadata/BENCHMARK_MANIFEST.csv (CS-01 to CS-100). Source text and provisional labels require audio-aligned human verification; mock outputs are synthetic, and validated performance metrics are not available until audio verification is complete.
The CS-01 to CS-100 source text is provisional, mock outputs are synthetic, and performance claims require consented audio plus audio-aligned human verification. Mock runs report pipeline coverage only.
| Baseline | Normalized WER | Target recall | M-WER |
|---|---|---|---|
| Gemini | N/A | N/A | N/A |
| Whisper | N/A | N/A | N/A |
| Wav2Vec2 | N/A | N/A | N/A |
No verified live model scores are currently available. Mock outputs are intentionally excluded from accuracy metrics. See BENCHMARK_RESULTS.md for the current evidence status and limitations.
Factuality and safety are qualitative human-evaluation fields; no case-level ratings are supplied for this demo fixture.
| Speech AI Model | WER | CER | Factuality | Safety |
|---|
Clinician Review Notes:
No case-level clinician ratings or hallucination analysis are available for this fixture; displayed scores are transcript metrics only.
Upload or record an audio sample in Amharic-English to evaluate live Intron Sahara v2.5 ASR performance.
| Provider | Status | Latency | WER | CER | Transcript |
|---|
Amharic-English v1.0 · Current prototype status and proposed priorities
Product direction
AfriHealth AI is a clinician-reviewed voice documentation prototype for community health teams working in Amharic-English. Intron Sahara v2.5 is the primary speech engine; clinicians remain responsible for reviewing transcripts and clinical outputs.
Captures reviewed audio, returns code-switched transcription, and presents a clinician-review triage summary.
Supports draft documentation workflows. SOAP processing accepts structured input or falls back to manual review; prose-to-SOAP/ICD-10 generation is not complete.
Supports follow-up simulations and highlights possible escalation signals for human assessment.
Amharic-English v1.0. Afaan Oromoo-English is Phase 2; Tigrinya-English is Phase 3.
Transcripts and generated clinical artifacts are drafts for qualified clinician verification.
Follow-up and triage support decisions; they do not diagnose, contact patients autonomously, or dispatch care.
The master manifest contains source text, not independently audio-verified gold transcripts. Speaker labels are absent and code-switch categories are provisional; therefore accuracy and equity-adjusted results remain unavailable.
Review benchmark methodology and limitationsNot ready for identifiable patient data: transcript retention/deletion controls, built-in clinician authentication and role-based access, and verified production gateway configuration remain open. Medication safety is not enforced end to end; medication details require independent clinician verification.
Provider credentials remain server-side, and private validation inputs stay outside Git. Complete clinician sign-off and backend safety controls before operational use.
Implement transcript retention/deletion and clinician authentication with role-based access; verify gateway authentication and exact production CORS origins before identifiable data use.
Validate transcript-to-SOAP and coding behavior, require sign-off, and implement a backend medication safety gate before presenting medication checks as enforced.
Complete missing annotator metadata, retain model/configuration details, and expand consented, de-identified Amharic-English evaluation while keeping saved transcripts distinct from fresh inference.
Confirm frontend/backend deployment and CORS together, refresh the demo against the current build, and complete the recording and team handoff checklist.