How to Evaluate an AI Medical Scribe Before Using Patient Audio
By Sphygmos · Updated 2026-09-16 · For South African healthcare practitioners
A polished demonstration cannot tell you how much correcting a note will need in your own consulting room. Use the same fictional examples with every supplier, inspect what was actually saved, and measure the work from recording to reviewed note. This guide includes a free worksheet; no account or patient information is needed.
Plain-text worksheet with six fictional tests and space for observations.
In this guide
Choose the workflow you are evaluating
Dictation captures the doctor's spoken account. Consultation recording captures a conversation between the doctor and patient. Evaluate them separately: a good dictated note does not demonstrate reliable speaker attribution in a conversation.
Set up a fictional patient and use colleagues who agree to participate in sample recordings. Record the product version, device, browser, microphone, language and test conditions. Do not use identifiable patient data to try an unapproved service.
Six practical tests to take to a demonstration
These original examples test record fidelity and workflow behaviour, not the clinical correctness of a treatment. They are deliberately small enough to inspect line by line. Repeat them in comparable conditions instead of comparing different rehearsed demonstrations.
For each test, record whether the requirement was met, a correction was needed, it failed, or it was not tested. Save the incorrect or missing fact in your private evaluation notes. An untested function is not a pass.
Keep the speaker and family history straight
Fictional patient Lindiwe Test: "My sister has asthma. I have never been diagnosed with asthma." Doctor: "Family history: sister with asthma. The patient reports no previous asthma diagnosis."
What to check: The sister, not Lindiwe, has the reported asthma history. The note must not turn a family history into a patient diagnosis.
Keep an unknown medicine unknown
Lindiwe Test: "I take a tablet for blood pressure. I cannot remember its name or strength." Doctor: "Medication reconciliation is still needed. No new prescription has been issued in this encounter."
What to check: The medicine and strength remain unconfirmed. No invented drug, dose, frequency or prescription appears.
Preserve a correction and a negative
Lindiwe Test: "I said left wrist earlier, but I meant my right wrist. I have not had a fever." Doctor: "Right wrist symptoms; patient reports no fever."
What to check: The final note records the right wrist and the reported absence of fever. It does not silently change a reported negative into an examination finding.
Separate a plan from a completed action
Doctor: "We will request the previous report. It has not arrived. A follow-up is proposed for next week, but no appointment has been booked."
What to check: The report is outstanding and follow-up is proposed. Neither receipt of the report nor a confirmed booking is invented.
Open a second patient without carrying the first note
Leave Lindiwe Test with an unfinished draft. Open a separate fictional record named Peter Test and begin a new encounter. Do not copy any details into it.
What to check: Peter has his own identity and encounter. Lindiwe's draft is not inserted into Peter's record, and returning to Lindiwe offers the correct unfinished work.
Find out what survives an interruption
Use a disposable fictional test encounter. Interrupt the connection or reload during a short sample recording, then reopen the same patient and encounter.
What to check: Record what was recovered and what was lost, whether the app explains the state, and whether a retry duplicates notes or usage. Do not assume recovery from a success message alone.
Measure a reviewed note, not a fast first draft
Start timing when you begin the task and finish when the corrected note is saved in the intended patient record. Separately record capture, generation, review and correction, and saving time. Include retries, assistance and work moved to another system.
Complete comparable fictional tasks with your current method. Alternate the order where practical, record the number of attempts, and compare the median and range rather than the best run. Separate unfamiliarity with the interface from repeated factual errors.
Do not combine every finding into an impressive-looking percentage. Wrong-patient content, an invented medicine or lost work deserves its own investigation even when the rest of the note reads well. A short trial cannot establish population-level accuracy or clinical safety.
Check what happens after the note
Ask to see the final record again from the patient timeline. Correct a detail and inspect how the change is represented. Prepare a referral or other document, check the correct patient and letterhead, and confirm that an outstanding action can be found again.
Test the phone and computer you will actually use. Compare what can be completed on each. Establish an alternative working process for an outage rather than assuming that a browser-based service works offline.
Separate product testing from permission to use patient data
HPCSA Booklet 20 places clinical judgement and accountability with the practitioner, addresses disclosure and data protection, and calls for assessment of an AI tool's suitability and reliability. It does not constitute approval of a named product. Read the relevant sections alongside your institution's requirements.
WHO identifies inaccurate or incomplete outputs and automation bias among the risks of generative AI in health. Fluent text alone is therefore not evidence that a note is correct.
Before an actual patient pilot, ask the supplier for its current processing terms, subprocessors, data locations, training-use policy, recording retention, export and deletion arrangements. Have the responsible people in your practice or institution assess them. Do not infer these answers from a padlock, a logo or a general compliance claim.
References: HPCSA Booklet 20: Ethical Guidelines on the Use of Artificial Intelligence, September 2025; sections 3-5, 7, 9 and 10 · WHO: Generative AI in health, potential benefits and risks, 18 January 2024
Include the cost of completing the work
Compare the allowance you would really consume: dictation minutes, consultation minutes, AI requests, images, storage and any paid seats. Ask what happens at the limit and what a top-up costs. Include retained billing or laboratory services and the time spent correcting notes.
For Sphygmos, the pricing page distinguishes Dictation from Consult recorder. The free plan includes limited dictation and three patient records; do not assume that it includes consultation recording. The fictional demo illustrates the interface, not the quality of a live model response.
Frequently asked questions
- How should I compare AI medical scribes?
- Use the same fictional material and comparable devices, record omissions and incorrect facts, and time review and correction as well as generation. Check the saved patient record and unfinished work. Use supplier claims as questions to test, not as your result.
- Does a successful trial prove that an AI scribe is clinically safe?
- No. These examples are a practical product evaluation, not an independent clinical study, legal assessment or regulatory certification. A patient pilot needs an appropriate scope, governance and clinician review.
- Can I use the worksheet with another supplier?
- Yes. The fictional examples and observation fields can be used to evaluate any supplier. Sphygmos publishes the worksheet and has a commercial interest in clinical software; it is not an independent product ranking.
Sources
- HPCSA Booklet 20: Ethical Guidelines on the Use of Artificial Intelligence, September 2025; sections 3-5, 7, 9 and 10
- WHO: Generative AI in health, potential benefits and risks, 18 January 2024
An original evaluation resource published by Sphygmos, not independent clinical validation or legal advice. The references inform the risk discussion, not an endorsement of the worksheet or product. No named clinician review is claimed.

