# AfriHealth AI Architecture

## 1. System purpose

AfriHealth AI is a browser-based clinical voice documentation and
decision-support prototype for community health workers and clinicians. The
v1.0 active scope is Amharic-English code-switching. Afaan Oromoo-English
is a Phase 2 roadmap item and Tigrinya-English is Phase 3; neither is enabled
in the active language selector.
The system produces draft transcripts and clinician-review artifacts; it is
not an autonomous diagnostic or prescribing service.

## 2. Runtime components

```text
Browser (index.html)
  |-- microphone PCM capture and local WAV review
  |-- WebSocket client for live STT
  |-- REST client for saved-recording STT upload
  |-- triage, EHR intake, follow-up, benchmark, and audit views
  |
  +--> FastAPI service (main.py :8000)
          |-- server-side Intron API-key boundary
          |-- streaming STT WebSocket proxy
          |-- follow-up TTS generation bridge
          |-- synchronous STT upload bridge
          |-- server-side FHIR export and optional EHR commit
          |-- synchronous TTS generation and status bridges
          |-- clinical artifact and medication safety gate
          |-- async SQLAlchemy persistence adapter (PostgreSQL in production; SQLite locally)
          |
          +--> Intron Sahara STT/TTS APIs
          +--> Configured FHIR/EHR endpoint (optional)

Optional local static server (server.js or npm start :3000)
  |-- serves the browser files
  +-- optional legacy Intron proxy route
```

The browser must not receive or store the Intron API key. Production
deployments should expose only the FastAPI service through an allowlisted
origin and configure `ALLOWED_ORIGINS`.

## 3. Main request flows

### Live transcription

1. The browser requests microphone access and captures audio.
2. Audio bytes and control messages are sent to `/ws/stream`.
3. FastAPI adds the server-side Authorization header and connects to Intron.
4. Partial and final transcript messages are returned to the browser.
5. Final text populates the transcript, entity, triage, and safety-review
   panels.

### Saved recording review and upload

1. The browser captures PCM and encodes a local WAV blob.
2. The user reviews playback and may download the local copy.
3. Upload occurs only after the explicit upload action.
4. FastAPI forwards the multipart file to Intron's synchronous STT endpoint.
5. The browser displays the returned transcript and clinician-review triage
   output.

Synchronous uploads are limited by the provider's documented 120-second
maximum. The controlled clinical uploader uses one reviewed recording per
case and keeps raw audio and provider response data outside Git.

### Clinical artifact generation

`/api/v1/process-clinical` (also available as `/api/v1/clinical/process-text`)
accepts transcript text, applies limited email
and phone redaction, and attempts to validate structured SOAP JSON. Ordinary
prose falls back to a summary marked for manual review; this endpoint does not
invoke a generation model and is not currently called by `index.html`. The UI's
editable clinical workflow and this API endpoint are separate paths. Do not
describe the current endpoint as an agentic SOAP/ICD-10 generator. Any medication
safety gate must be verified in the specific workflow before describing it as
enforced end to end.

### Clinician review and record export gates

The browser displays a persistent clinical safety disclaimer and marks
transcripts and derived artifacts as pending verification. The EHR workflow
requires a separate review-confirmation checkbox and sign-off action. Approval
does not automatically commit data; FHIR export, clinical summary export, and
EHR commit controls remain unavailable until the reviewer signs off. Editing
the current SOAP draft invalidates its approval.

Both `/api/v1/fhir/export` and `/api/v1/ehr/commit` require the
`clinician_signed_off` request field. This is a caller-supplied assertion, not
authenticated clinician identity or an auditable digital signature. The
prototype still requires deployment-level authentication, authorization, and
audit controls before identifiable patient data can be used.

The UI does not synthesize ASR confidence values. When confidence is
unavailable, the clinical SOAP response carries `null` rather than a synthetic
zero. Until verified backend inference metadata provides a confidence tier, the
UI shows a pending-verification placeholder. A score supplied inside
structured SOAP input is not independently verified and must not be treated as
inference metadata. Frontend heuristics must not be presented as model
confidence.

## 4. Evidence and evaluation layers

The repository contains separate evidence types:

| Evidence | Purpose | Interpretation |
| --- | --- | --- |
| `benchmark/metadata/BENCHMARK_MANIFEST.csv` | Sole active 100-case clinical reference source | Source text only; audio alignment and adjudication remain pending |
| `inference_engine.py` + `evaluator.py` | Resumable ASR attempts and manifest-based metrics | Mock mode is synthetic; validated live comparison is not currently available |
| Embedded application fixture | Reproducible UI/scoring demonstration | Not a population performance claim or benchmark result |
| AfriSwitch pilot scripts | Separate general-purpose code-switched ASR experiment | Not clinical validation and not part of this clinical benchmark |
| Review protocols and safety documentation | Human review, safety scenarios, SOAP scoring | Maintained outside the public demo flow |

Do not combine application-fixture metrics or AfriSwitch results with the
clinical benchmark.

`clinical_validation_evaluator.py` is a compatibility scorer that resolves
references and vocabulary from the same master manifest and medical dictionary.
It does not train a model or support autonomous clinical decisions.

### Stream session metadata

Each live stream receives a session ID and clinic ID, which are stored in
PostgreSQL using Railway's `DATABASE_URL`; local development defaults to an
async SQLite database. The current WebSocket path returns partial and committed
transcripts to the browser but does not call the transcript persistence method.
Raw audio is processed in memory and is not stored by this layer. Session
metadata has no expiry or deletion policy. The schema is created automatically
at service startup.

### Live benchmark comparison

Module 4 keeps the fixture matrix separate from live measurements. A reviewed
audio sample and verified Amharic-English reference transcript can be sent to
`/api/v1/benchmark/live`, which benchmarks Intron Sahara v2.5 by default and
can later compare OpenAI and Gemini when explicitly selected. It returns
provider status, latency, WER, and CER. Missing provider keys are reported as
unavailable; no scores are fabricated. If the UI selects the Intron output as
the reference, the response is marked `provisional_intron_reference`; that
mode is useful for model-to-model error comparison but is not independent
gold-standard accuracy.

The separate clinical validation audit form is intentionally not exposed in
the public application navigation. The product demo focuses on the three care
modules and benchmark matrix; review protocols remain available for controlled
evaluation and future clinical governance.

## 5. Security and privacy boundaries

- Keep `INTRON_API_KEY` in a server environment variable.
- Never place credentials in `index.html`, commits, screenshots, or demo
  recordings.
- Raw audio is not stored by the stream persistence tables. The WebSocket path
  currently persists session metadata only; it does not write transcript text
  to `clinic_sessions` or `transcript_events`. No expiry or deletion policy is
  implemented for session metadata.
- The clinical text endpoint redacts common email addresses and Ethiopian
  phone numbers only. This limited redaction is not applied to live-stream
  persistence and is not comprehensive de-identification.
- Do not use identifiable patient data until retention, deletion, and access
  controls have been approved and implemented.
- Use de-identified or simulated cases for public demonstrations.
- Configure exact production CORS origins rather than broad wildcards.
- Treat generated transcripts, entities, codes, triage labels, and medication
  candidates as clinician-review drafts.

The persistence `session_id` is a stream-session UUID. It is not linked to the
separate, caller-supplied FHIR `encounter_id`. The repository has no patient or
clinician tables and no built-in clinician accounts or role-based access
control. `REQUIRE_PROXY_AUTH` enables only a trusted identity-header check.

## 6. Known limitations

- The current frontend uses a lightweight browser capture path and requires a
  compatible microphone/browser for live recording.
- The current release targets English-Amharic code-switching. Oromo is not
  established as production-supported; provider and checkpoint limitations
  still apply.
- Clinical validation currently reports a baseline, not a safety or efficacy
  claim.
- No demo video is included in this repository. The written
  [DEMO_SCRIPT.md](./DEMO_SCRIPT.md) is provided for a future recording or
  live judging session.
