Module overview
Section 4 of 6 · Open sections

Required section · Section 4 of 6

Working the in-service note: catching the fabricated citation

Return to the student's draft. The AI tool cited "CLSI H21-A6, Collection, Transport, and Processing of Blood Specimens, 2019." The current CLSI guideline for collection of venous blood specimens is PRE02, 8th edition, which replaced the earlier GP41/H3 document in February 2025; H21 concerns coagulation specimens. The citation is fabricated in the exact pattern described by two 2023 studies of chatbot-generated references: a JAMA Network Open study found GPT-3.5 fabricated 98.1% of the journal citations it generated for medical topics and GPT-4 fabricated 20.6%, and a JAMA Ophthalmology study found chatbot-generated reference lists were unverifiable roughly 29-31% of the time, including in an updated model version. Those rates describe specific 2023 model versions; they are not a current, fixed fabrication rate for any product.

The chart below shows those three reported figures side by side. The pattern to take from it is not a specific percentage to memorize, it is that a generated citation can be wrong at a rate high enough to matter, that the rate is not the same across model versions, and that a citation a model produces is not evidence the source exists or supports the claim. Only checking the citation against the original document, here the CLSI catalog itself, settles the question.

The student also checked the numeric threshold the tool stated, 20% or 1.0 mmol/L, against the laboratory's own delta-check standard operating procedure (SOP), and the value did not match. This is the second half of verification: a calculation or a stated criterion needs the same check as a citation, against the laboratory's own controlled document, because the model has no access to that document unless the tool is explicitly built to retrieve it.

The student corrected the draft: replaced the fabricated citation with the laboratory's actual controlled-document number, replaced the threshold with the value verified from the SOP, and kept the AI-drafted explanatory language on hemolysis and pseudohyperkalemia after confirming it did not conflict with local policy. The student logged the exact prompt, model name and version, date, and name of the person who verified the output. The draft was routed to the bench supervisor for sign-off before posting. At no point was the note used to release, reinterpret, or override an actual patient result, which kept the whole task in the lower-risk band from the risk ladder in the previous section.

Separate what the model did, produce fluent prose and a citation-shaped string, from what it demonstrated, nothing about whether that citation or number was real; the laboratory's own controlled document, not the model's confidence, is the source of truth.

Illustrative drawing — this picture was drawn rather than captured.

Bar chart with three bars. GPT-3.5 in coral at 98.1%. GPT-4 in navy at 20.6%, with a teal dashed reference line at that height. Ophthalmology chatbot reference set as a light blue band spanning 29 to 31 percent. Caption notes the rates are model- and date-specific.
Figure 1Reported 2023 chatbot citation fabrication rates: GPT-3.5 98.1%, GPT-4 20.6%, and an ophthalmology chatbot reference set unverifiable 29-31% of the time.

Illustrative drawing — this picture was drawn rather than captured.

Five connected steps: prompt entered with no PHI or lot numbers, AI draft returned with a citation and threshold, verify step in coral where the citation and threshold are checked against CLSI catalog and local SOP, correct and log step, and supervisor sign-off before the note is posted, with a note that the laboratory's controlled document is the source of truth.
Figure 2The five-step verification workflow the guided case follows, from prompt to a supervisor-reviewed, logged note.
Reported chatbot citation fabrication rates from two 2023 studies, by model or tool.
Model or toolStudyReported fabrication or unverifiable rate
GPT-3.5JAMA Network Open, 202398.1% of generated citations fabricated
GPT-4JAMA Network Open, 202320.6% of generated citations fabricated
Ophthalmology-focused chatbot setJAMA Ophthalmology, 202329-31% of references unverifiable

Knowledge checks

Reading and checks are open. Sign in only to save.

Knowledge check 1

The in-service note draft cites "CLSI H21-A6" for blood specimen collection guidance. What does a citation generated by an LLM tell you about whether that source exists and supports the claim?

Choose one option.

Knowledge check 2

The AI-drafted note also states a delta-check threshold of '20% or 1.0 mmol/L.' The student does not recognize that number from the local SOP. What is the correct next step?

Choose one option.

Knowledge check 3

Which fields belong in a defensible prompt/output audit record, per NIST's Generative AI Profile logging recommendation and the guided case?

Choose at least 4 options.

Section status

Finish this section

Reading and checks are open. Sign in only to save.

The module finishes after every required section is marked done and every check in those sections is correct.