Case Study · Consumer Health
ClinCalc
Consumer health self-check & multimodal interpretation platform
Pain points
Many people get a lab report full of numbers and jargon and don't know which values are concerning and which can wait. The default behaviors are: search online, ask family, or ignore it. Two structural problems sit underneath:
- Information asymmetry: medical terminology is hard for non-experts
- Privacy: pasting a whole report into a chatbot sends the values and personal details along with it
System architecture
"Rules first, then LLM": interpretation is done by rules in the browser; the language model only writes up the results, and the lab data it receives contains no raw values.
- 1 · User input
Enter lab values (35 indicators available), upload a report photo, or pick symptoms on the body map. eGFR is the value the user copies from the report; the app does not compute it.
- 2 · In-browser rules engine
Looks up
referenceRanges.tsand labels each indicator normal / high / low / critically high / critically low; eGFR is also staged G1–G5 (G3 split into a/b) by KDIGO thresholds. Interpretation needs no server round-trip. - 3 · When the user asks for AI analysis
For lab data, only each indicator's interpretation and reference range is sent, e.g. "eGFR: low (reference ≥ 60)", never "eGFR = 52"; age, sex and the user's own symptom notes are included. This restriction was added on 2026-09-27; before that the values were included, as noted in the thesis errata.
- 4 · Google Gemini 2.5 Flash (proxied server-side)
Writes the interpretation up in plain language, and handles report-photo text recognition and zh/en medical translation. Requires login; the API key stays on the server.
- 5 · Supabase (PostgreSQL + Auth)
Signed-in users' health records and medication reminders are stored here under RLS; without login, records stay in the browser.
Design point 1: rules first, then LLM
After the user fills in values, the frontend looks up referenceRanges.ts per field and calls
checkAbnormal() for a five-level verdict, then builds the prompt from the verdicts.
Correctness is decided by the rules, so the model never has to recall reference ranges and cannot invent them.
Trade-off
- Limited rule coverage: holistic questions ("how does my report look overall?") are beyond the rules, and the model is limited to restating them, so the app feels like it does less
- Rules are updated by hand: guideline revisions mean editing the rules
- Single axis: without urine albumin, staging reflects GFR only, and one reading cannot diagnose CKD (it must persist over three months)
Design point 2: multimodal models only "translate"
Gemini's multimodal ability is used for two things only: reading text from report photos, and zh/en medical translation. It makes no clinical decisions.
Trade-off
Photos and translation text are sent by the user as-is; the app does no automatic de-identification. Recognition results are shown as text and may be wrong, so users need to check them against the original report.
Design point 3: symptom exploration by body map
Instead of typing keywords, users pick a region, tick symptoms, and give duration and severity. Gemini then returns an urgency level (go now, see a doctor soon, routine visit, self-monitor), the reasons, and home-care notes. The prompt forbids recommending any medication.
Derived feature: personal health tracking
When signed in, results are stored in Supabase under RLS as a personal time series. Users can view each indicator's history, set medication reminders, and create a single-use invitation link, valid for 7 days, that lets a chosen physician read their records until revoked.
Security & privacy
- Interpretation in the browser
All 35 indicators and the KDIGO eGFR staging are computed in the browser; no values need to go to a server for interpretation.
- Minimal model input
For lab data, AI analysis sends only verdicts and reference ranges, no values; age, sex and symptom notes are included. Photo recognition and translation send what the user chose to submit, as-is.
- RLS at the storage layer
Records and reminders live in Supabase with PostgreSQL Row Level Security on every table; users can only read and write their own rows, and a physician needs the patient's active consent plus two-factor authentication to read them.
- Keys stay server-side
The Gemini API key is server-side only; AI features require login and are rate-limited to 10 requests per user per minute with an atomic counter in the database. GitHub Secret Scanning and Push Protection are enabled.
Tech stack
Research questions surfaced
- Does "rules first, then LLM" extend to fuzzy tasks? It works where there are clear thresholds, like KDIGO staging, but rules can't cover holistic judgement. How should the work be split between rules and model?
- How do guideline revisions reach the system? Guidelines change every few years and hand-editing rules doesn't scale. Could guidelines become an updatable knowledge base?
- How much error is acceptable for lay users? For people without medical training, how should disclosure and fallbacks be designed?