Scribe, an ambient scribe for therapy sessions
A silent assistant that listens to a psychology session, drafts the notes, remembers what matters for next time, and lets the therapist decide what stays private. Designed with a clinical psychologist at a large Indian hospital group, and built on Echo.
The problem
Therapists re-read hours of old notes before a session, fill hospital forms by hand after it, and prepare redacted copies for other doctors manually.
What I did
Ran discovery, benchmarked 20+ scribes and 8 speech models, then specified an MVP that turns a live session into notes, memories and a form, with privacy enforced in code.
Expected result
About 3 hours of clinician time back each day. Two versions of every note in under 45 seconds. Up to four speakers told apart.
The problem
I started with a long discovery session with a clinical psychologist at the hospital. The hospital already runs three systems: an EHR, a clinician app that mostly tracks admitted patients, and a patient app. None of them helps with the session itself. She described her workflow as entirely manual.
Before a session she re-reads notes from earlier sessions to spot patterns and changes. That can take two to three hours.
After every appointment the hospital’s consultation form is filled in by hand. A doctor seeing about 40 patients a day does that 40 times.
One session needs three versions: what only she sees, what another doctor may see, and what goes into the hospital record. Redacted copies are prepared by hand.
What already existed
I researched more than 20 scribes in India and abroad, and read the independent benchmarks for Indian speech models. Three findings shaped the product.
- Indian scribes compete on price and language count. Most cost about ₹1,500 per doctor per month and claim 20 or more languages, but none publishes accuracy by language. None handles therapy sessions, prompts during the session, or memory across sessions.
- Transcription alone returns little time. A randomised trial of one of the best-known global scribes found it saved 1.7% of note-writing time, which was not significant.
- Real rooms are much harder than benchmarks. On real psychiatric interviews in India, the best speech model with timestamps got about half of the Kannada words wrong. On clean phone speech, the same class of models gets about 16% wrong.
We were selling time and memory, not transcription.
Scope: one specialty, one week to build
Her wish list covered the whole hospital: cross-specialty records, referrals, scheduling, a patient chatbot and psychological assessments. I scoped the MVP to the part no one else does well, end to end, for psychology only: before, during and after one session, plus memory across sessions.
Built
- A 30-second brief before the session
- Live transcript with speaker labels and cards
- Hiding by hand or by typing a request
- Two notes at the end, the hospital form on request
- Memories the therapist accepts, and an assistant that answers from them
Designed, not built yet
- Psychological assessments and scoring, her biggest gap. We named it in the demo instead of hiding it
- Referral letters and a patient version of the note
- The records of other specialties
- Scheduling and an after-hours patient chatbot
The plan fitted a seven-day build: one day reserved purely for tuning prompts on a seeded case, and a feature freeze at noon on the last day.
How it works
What changed since last time, and the patient’s memories grouped by topic. No old notes to open.
Every 20 seconds, new lines are read against all of the patient’s memories. Four kinds of card come back: fact, flag, link and change.
Two notes write themselves in under 45 seconds. The hospital form fills only when asked. New memories wait for the therapist to accept them.
Decisions that make it safe to use
A scribe in a therapy room can do harm in two ways: by saying something the patient never said, and by showing a private detail to the wrong reader. Most of my decisions target one of those two.
No quote, no card
Every card, memory, note section and form field carries the patient’s exact words.
Code checks the quote really appears in that line. If it doesn’t, the item is dropped.
The model suggests, code decides
Topics come from a fixed list of twelve. Patterns are counted by code. Code sets what is private.
The model writes sentences. It is never the last line of defence.
“Possibly related”, never “because”
Cards never state a cause or a diagnosis. A risk statement is shown with its quote, and nothing else happens.
This keeps Scribe on the safe side of UK and Australian guidance: transcribing and summarising are fine; diagnosing makes it a medical device.
Every memory, every call
All 20 to 60 of a patient’s memories go into each call, instead of searching for the relevant ones.
A search could miss the one memory that matters, like a father who died two weeks ago.
Make doubt visible
A doubtful line is shown faded with the reason. Fixing who spoke takes one click and updates every card built on that line.
Published research found that 80 to 100% of errors in what the patient said reach the final note, against 20 to 30% for the clinician’s words.
Buy now, own later
Hosted Indian speech models and voiceprints for the MVP; our own models later.
We switch when a hospital requires audio to stay on site, or once we have about 25 hours of consented audio to tune on.
Privacy is built in, not filtered out
The therapist can hide any words at two levels. Private means only she sees them. Not for the record means they stay in her own note but never reach the hospital record. Code then builds three views of the session, and each job reads only the view its reader may see.
| FullHer own note, memory, assistant | SharedThe version for another doctor | RecordThe hospital form |
|---|---|---|
| Sleep down to about 3 hours | Sleep down to about 3 hours | Sleep down to about 3 hours |
| Dispute with his brother over the family house (private) | Never seen | Never seen |
| Hasn’t told his wife how bad it is (not for the record) | Hasn’t told his wife how bad it is | Never seen |
The hospital form waits until the therapist asks for it. She marks what to keep out first, so the model filling the form never reads it and cannot repeat it in the wording of another field. Diagnosis and anything that can only be seen, like appearance, stay with the clinician.
Designing the screens
I designed the screens to look like a clinical case file, not a consumer health app, and tested them as a clickable prototype.
- The dashboard is home. Three queues in order of urgency: notes that need review, today’s sessions, and versions waiting to be sent.
- The session is a mode. Full screen, no navigation, one way out. Earlier sessions appear only as a quiet suggestion.
- Review only what’s doubtful. After the session she sees “N places to check”, not the whole transcript, then signs once.
- No percentages on screen. A doubt is a state plus a reason in words. Colour is never the only signal, and each privacy level has its own icon.
Results
of clinician time expected back each day, from note review and form filling
to draft two versions of the note, one private and one shareable
told apart in one session, including a family member
These are expected results; Scribe is still in production. We will measure minutes to a signed note against the therapist’s own baseline, count form fields filled without typing, and log the connections a manual review would have missed.
What I learned
- The value starts after the words arrive. Transcription is table stakes. Memory, privacy and the form are what give time back.
- Put guardrails in code, not only in prompts. Quote checks, fixed topics and views chosen by code make the rules hold even when the model slips.
- Say what you are not building. Naming assessments as the biggest gap made the demo more credible, and gave us the next thing to build.
Next: assessments and scoring, referral letters, a patient version of the note, and tuning speech models on consented session audio.
Want the details behind any of this?
The full spec, architecture and research are available on request, with client details removed.