Laboratory Operations

Quality Control

Quality Control, Method Evaluation, and Quality Management

Laboratories assess quality across the entire testing path. During the analytic phase, control results, calibration records, and method-evaluation data show whether a measurement procedure is performing as intended. Proficiency testing compares the laboratory’s results with an external target or peer group. A quality management system links this work to the laboratory’s other processes.

Statistical control of a measurement procedure

Repeated measurements of a stable control material produce a distribution of results around a central value. If that distribution is approximately Gaussian, about 68.3% of results fall within 1 standard deviation (SD) of the mean, 95.4% within 2 SD, and 99.7% within 3 SD. These proportions are the basis for common statistical control rules, but the laboratory must first confirm that its control data are suitable for those rules.1,2

  • The mean is the arithmetic average.
  • The standard deviation describes dispersion around the mean in the analyte’s units.
  • The variance is the SD squared.
  • The coefficient of variation compares imprecision across concentrations or analytes: CV (%) = SD ÷ mean × 100.

A potassium control produces nine results: 4.1, 4.2, 4.2, 4.3, 4.3, 4.3, 4.4, 4.4, and 4.5 mmol/L.

StatisticCalculationResult
Mean38.7 ÷ 94.30 mmol/L
MedianMiddle ordered value4.3 mmol/L
ModeMost frequent value4.3 mmol/L
VarianceΣ(xi − mean)² ÷ (n − 1)0.0150 (mmol/L)²
SD√0.01500.122 mmol/L
CV0.122 ÷ 4.30 × 1002.8%

The resulting control lines are:

IntervalLower limitUpper limit
±1 SD4.18 mmol/L4.42 mmol/L
±2 SD4.06 mmol/L4.54 mmol/L
±3 SD3.93 mmol/L4.67 mmol/L

These nine results demonstrate the calculations but cannot establish that the underlying distribution is Gaussian. For a reported ±2 SD interval of 80 to 110 mg/dL, the midpoint is 95 mg/dL. The full interval spans 4 SD, so 1 SD is (110 − 80) ÷ 4 = 7.5 mg/dL.

Control materials

Control materials contain one or more analytes in a matrix that is the same as, or similar to, the patient specimen when an appropriate material is available. Laboratories commonly monitor at least two concentrations at clinically relevant points. Materials may be liquid, frozen, or lyophilized. Lyophilized material must be reconstituted with the manufacturer-specified volume and diluent because reconstitution error changes the observed concentration.

For most quantitative nonwaived procedures, the CLIA default is two control materials at different concentrations once each day patient specimens are assayed. Qualitative procedures use a negative and a positive control. Graded or titered, extraction, and molecular procedures have additional method-appropriate control requirements. More stringent manufacturer instructions apply, and an approved Individualized Quality Control Plan can provide an eligible test system with a different risk-based plan. CLIA also requires control procedures after a complete reagent change and after major preventive maintenance or replacement of a critical part that can affect performance. The laboratory’s written procedure may call for controls after additional events, such as troubleshooting or calibration.3

Local means and SDs should capture routine day-to-day variation. A common starting design uses at least 20 measurements over at least 10 working days, preferably about 20 working days and, when practical, more than one calibration cycle. This is practice guidance rather than a CLIA requirement because CLIA sets no single universal design for every control material. Manufacturer ranges can guide initial use. After sufficient local data are available, locally evaluated limits usually detect changes in the laboratory’s own procedure more effectively.2

Commutability describes whether a material and authentic patient specimens show equivalent relationships among different measurement procedures. A noncommutable control can still monitor within-method stability when the laboratory establishes suitable local limits. It is unsuitable for demonstrating calibration traceability across methods. Because commercial processing, stabilizers, and added analyte can affect commutability, the material’s intended use and supporting evidence matter more than whether it came from the instrument manufacturer or a third party.2

Levey-Jennings charts

A Levey-Jennings chart plots successive control results on the vertical axis against time or run number on the horizontal axis. Horizontal lines mark the mean and the ±1, ±2, and ±3 SD limits. For a hemoglobin control with a mean of 15.0 g/dL and an SD of 1.5 g/dL, the lines are:

LineValue
+3 SD19.5 g/dL
+2 SD18.0 g/dL
+1 SD16.5 g/dL
Mean15.0 g/dL
−1 SD13.5 g/dL
−2 SD12.0 g/dL
−3 SD10.5 g/dL

Results from one defined control lot are plotted against that lot’s mean and SD. A new lot needs its own evaluated target and limits. Overlap testing helps separate a true method change from a difference between control lots.2

Random error creates unpredictable scatter. Examples include an air bubble, an inconsistent pipette delivery, or transient electrical noise. Systematic error moves results in one direction. Reagent deterioration, a calibration problem, or a temperature shift can cause systematic error. An abrupt sustained displacement appears as a shift; gradual movement in one direction appears as a trend. The number of points used to define a shift or trend belongs in the laboratory’s written control policy because conventions differ among procedures.2

Multirule interpretation

Westgard multirules combine control observations to improve detection of random and systematic error. Laboratories commonly include them in their control policies. CLIA requires effective control procedures but does not mandate this rule set.4

RuleTriggerUsual interpretation
12sOne result exceeds ±2 SDWarning that prompts review of the other rules
13sOne result exceeds ±3 SDReject; large random error or systematic error is possible
22sTwo consecutive results exceed the same 2 SD limit in the same directionReject; systematic error is likely
R4sTwo results in one run differ by more than 4 SD, with one above +2 SD and one below −2 SDReject; random error is likely
41sFour consecutive results exceed the same 1 SD limit in the same directionReject; a small systematic change is likely
10x, 9x, or 8xThe specified consecutive results fall on one side of the meanReject under the selected policy; systematic change is likely
(2 of 3)2sTwo of three control levels exceed the same 2 SD limit in the same directionReject; systematic error is likely

A rule pattern points to a class of problem; it does not prove the cause:

  • A glucose control has a mean of 120 mg/dL and SD of 4 mg/dL. Ten results are 123, 125, 124, 126, 122, 127, 124, 125, 126, and 123 mg/dL. All lie within ±2 SD, but all are above the mean. The sequence satisfies the 10x rule and indicates a shift.
  • CK controls have targets of 60 U/L with SD 3 and 180 U/L with SD 9. Results of 66.3 and 199.8 U/L in one run are +2.1 SD and +2.2 SD, satisfying 22s. Results of 53.7 and 199.8 U/L in one run are −2.1 SD and +2.2 SD, a 4.3 SD spread that satisfies R4s.
  • A sodium control has a mean of 145 mmol/L and SD of 1.5 mmol/L. Four results of 147.0, 147.5, 147.0, and 147.5 mmol/L are all above +1 SD and satisfy 41s.

When a rejection rule is met, the laboratory stops reporting affected patient results, investigates the control failure, corrects the cause, and documents acceptable control performance before resuming testing. It evaluates all patient results from the unacceptable run and all results reported since the last acceptable run to determine whether any were adversely affected. Corrections follow the laboratory’s procedure and the clinical significance of the error.3

Individualized quality control plans and calibration

An Individualized Quality Control Plan (IQCP) is a CLIA quality-control option for eligible nonwaived test systems. Its three parts work together:5

PartRequired work
Risk assessmentEvaluate hazards associated with the specimen, test system, reagent, environment, and testing personnel
Quality control planState the control practices that reduce the identified risks
Quality assessmentMonitor the plan, investigate failures, and revise the assessment when relevant conditions change

Before implementation, the laboratory director approves, signs, and dates the plan. The plan must meet manufacturer requirements and applicable regulatory requirements. Because some testing is excluded from IQCP, the laboratory checks eligibility before developing a plan.

Calibration establishes the relationship between an instrument signal and analyte concentration. Calibrators must be specified or validated for the measurement system. Calibration mainly addresses systematic bias; routine control results remain necessary to detect subsequent changes and imprecision. When a higher-order reference system exists, traceability links the patient result through an unbroken calibration chain to that reference.3

For nonwaived testing, CLIA requires calibration verification according to the manufacturer’s instructions and at least once every six months. Verification generally spans the reportable range with materials near the low, middle, and high values. It is also required after specified changes or events that could affect performance. These include a complete reagent change unless the laboratory demonstrates that lot changes leave the reportable range and control values unaffected, major maintenance or critical-part replacement, and control evidence of an unusual shift, trend, or unacceptable value after other assessment and correction fail to identify or correct the problem.3

Method verification and performance establishment

Before reporting patient results, the laboratory documents that a procedure performs adequately in its own setting. Its responsibility depends on the status of the procedure.3

ProcedureLaboratory responsibility before patient testing
Unmodified FDA-cleared or FDA-approved systemVerify accuracy, precision, and reportable range comparable with the manufacturer’s claims; verify that the reference interval is appropriate for the laboratory’s patient population
Modified FDA-cleared or FDA-approved system, laboratory-developed procedure, or procedure without FDA clearance or approvalEstablish accuracy, precision, analytical sensitivity, analytical specificity including interfering substances, reportable range, reference intervals, and other performance characteristics needed for test performance

When the laboratory performs the same test with different methods, instruments, or testing sites, it evaluates the relationship between results at least twice each year. The laboratory should define acceptance criteria before reviewing the data.3

Each performance characteristic answers a different question:

  • Accuracy or trueness: How close are results to an accepted target?
  • Precision: How closely do repeated results agree with one another?
  • Analytical sensitivity: What is the method’s detection capability, including the limit of blank, limit of detection, or limit of quantitation when applicable?
  • Analytical specificity or selectivity: How well does the method measure the intended analyte in the presence of interferents?
  • Linearity: How closely does the response follow the stated mathematical relationship across a tested interval?
  • Reportable range: What interval can the laboratory report after applying its validated specimen-processing procedures?
  • Reference interval: What range describes the designated reference population under defined conditions?

Comparing a candidate method

A paired t-test asks whether paired results have a statistically detectable mean difference. For each specimen, calculate d = candidate − comparison, then use:

t = d̄ ÷ (SDd ÷ √n)

Eight paired WBC specimens produce differences of +0.2, −0.1, +0.3, −0.2, +0.2, −0.1, +0.3, and −0.1 × 109/L. Their mean difference is 0.0625 × 109/L and their SD is 0.207 × 109/L. Therefore t = 0.0625 ÷ (0.207 ÷ √8) = 0.86, with 7 degrees of freedom. The two-tailed critical value at P = .05 is 2.365, so this set does not show a statistically significant mean difference. A nonsignificant result leaves equivalence unresolved.

A complete method comparison evaluates bias across the measuring interval, examines differences between paired results, and uses regression suited to the data and error structure. High correlation can coexist with unacceptable bias.6

For ten calcium specimens from 7.6 to 11.2 mg/dL, paired candidate results from 7.7 to 11.4 mg/dL give a least-squares line of y = 1.024x − 0.108, r = 0.9997, and a standard error of the estimate of about 0.032 mg/dL. The slope describes proportional bias, and the intercept describes constant bias. Correlation measures association between the methods, while the standard error of the estimate describes residual scatter around the fitted line. Acceptance still depends on the predefined allowable bias across the interval.

Evaluating precision

The F statistic compares two variances:

F = larger variance ÷ smaller variance

If two coagulation analyzers each measure one prothrombin-time control 18 times and have SDs of 1.2 and 1.6 seconds, their variances are 1.44 and 2.56 seconds². The calculated F is 2.56 ÷ 1.44 = 1.78, with degrees of freedom (17, 17). A one-sided upper-tail critical value at α = .05 is about 2.27, so these data do not demonstrate a precision difference. This single-level calculation illustrates variance comparison. A precision verification study uses a defined design with suitable concentrations, days, runs, and replicates.7,8

Recovery, interference, linearity, and reference intervals

A recovery experiment adds a known amount of analyte and measures the observed increase:

Recovery (%) = [(spiked result − unspiked result) ÷ amount added to the final aliquot] × 100

For five sodium specimens spiked to add 12 mmol/L, observed increases of 12.2, 11.8, 12.1, 12.3, and 11.9 mmol/L give recoveries of 101.7%, 98.3%, 100.8%, 102.5%, and 99.2%. Mean recovery is 100.5%. The result is judged against a predefined allowable recovery or bias criterion.3

An interference study compares an interferent-spiked aliquot with a paired aliquot containing an equal volume of diluent:

Bias = mean(interferent-spiked aliquot) − mean(diluent-control aliquot)

Five potassium pairs show biases of +0.4, +0.5, +0.6, +0.4, and +0.5 mmol/L after adding hemolysate. The mean bias is +0.48 mmol/L. This finding applies to the tested hemolysate level and the evaluated procedure. Other interferent levels require their own evidence.9

Linearity and reportable range are related but distinct. A linearity study evaluates the response pattern across selected concentrations. The reportable range includes the values the laboratory can release after validated specimen dilution or concentration procedures. Study design follows the manufacturer’s instructions, the laboratory’s protocol, and current guidance; CLIA specifies no one universal number of concentrations or replicates. A result outside the established range is diluted and reanalyzed only under a validated procedure. Otherwise, the laboratory uses its approved above-range or below-range reporting convention.3,10

Reference-interval verification asks whether a proposed interval fits the procedure and the laboratory’s reference population. A common CLSI approach tests 20 qualified reference individuals. The interval is verified directly when no more than two results fall outside it. If three or four fall outside, a second set of 20 is tested. The interval fails verification when five or more in the first set fall outside, or when three or more in the second set fall outside. The laboratory then evaluates another interval or establishes a new one.11

Selected two-tailed critical t values

dfP = .10P = .05P = .01
52.0152.5714.032
91.8332.2623.250
101.8122.2283.169
151.7532.1312.947
201.7252.0862.845
251.7082.0602.787
291.6992.0452.756
301.6972.0422.750
401.6842.0212.704
601.6712.0002.660
1201.6581.9802.617
1.6451.9602.576

As degrees of freedom increase, the t critical values approach the corresponding standard-normal values. The calculation uses the standard error of the estimated mean difference, not the SD of individual observations.12

Proficiency testing and external quality assessment

External quality assessment includes several kinds of external comparison. Proficiency testing is a regulated form in which an outside program sends blinded samples for testing. The laboratory compares its results with an assigned target or with a peer group using similar instruments, methodologies, and reagent systems. Material commutability determines how confidently results can be compared across measurement procedures. Accuracy-based programs use commutable material and reference-method targets when available.13

For analytes covered by CLIA proficiency-testing requirements, the laboratory enrolls in an HHS-approved program. For other testing, it verifies accuracy at least twice yearly when CLIA requires an alternative assessment. During a PT event, the personnel who ordinarily perform the test process samples through the routine patient-testing workflow. The laboratory uses its routine method and number of repeats and keeps PT records for at least two years. Participating laboratories must not communicate with one another before the submission deadline.3,14

If the patient-testing procedure normally calls for reflex, distributive, or confirmatory testing at another laboratory, PT testing stops at the point where the patient specimen would be referred. A PT specimen must not be sent to another laboratory for analysis the sending laboratory is certified to perform. A first referral limited to otherwise routine reflex, distributive, or confirmatory testing is still improper, but CLIA treats it as subject to alternative sanctions instead of intentional referral. A laboratory that receives a PT specimen must notify CMS.14

An unacceptable result requires a documented investigation. The laboratory reviews transcription, sample handling, calculations, reagents, calibration, control performance, instrument function, and personnel technique. Corrective action addresses the identified cause, evaluates possible effects on patient testing, and includes evidence that the action worked. The laboratory director reviews the investigation and conclusion, including cases in which the evidence supports random error as the likely explanation. An initial PT failure can lead to corrective or regulatory action according to the analyte, event pattern, investigation, and applicable CLIA provisions.3

Representative current CLIA performance limits

These selected limits are a study aid. They summarize the current federal proficiency-testing criteria but do not replace the full regulation or a program’s grading instructions.15

DisciplineAnalyteAcceptable limit
ChemistryGlucose±8% or ±6 mg/dL, whichever is greater
ChemistrySodium±4 mmol/L
ChemistryPotassium±0.3 mmol/L
ChemistryTotal calcium±1.0 mg/dL
ChemistryUrea nitrogen±9% or ±2 mg/dL, whichever is greater
ChemistryTotal bilirubin±20% or ±0.4 mg/dL, whichever is greater
ChemistryTotal cholesterol±10%
ChemistryALT or AST±15% or ±6 U/L, whichever is greater
ChemistryAlkaline phosphatase or amylase±20%
Blood gaspH±0.04
Blood gaspCO2±8% or ±5 mm Hg, whichever is greater
EndocrinologyTSH±20% or ±0.2 mIU/L, whichever is greater
EndocrinologyFree T4±15% or ±0.3 ng/dL, whichever is greater
EndocrinologyhCG, excluding waived visual urine pregnancy tests±18% or ±3 mIU/mL, whichever is greater, or positive or negative
ToxicologyDigoxin±15% or ±0.2 ng/mL, whichever is greater
ToxicologyBlood lead±10% or ±2 µg/dL, whichever is greater
HematologyHemoglobin or nonspun hematocrit±4%
HematologyPlatelet count±25%
HematologyProthrombin time or activated partial thromboplastin time±15%
ImmunohematologyABO grouping, D typing, unexpected antibody detection, and compatibility testing100% accuracy
ImmunohematologyAntibody identificationAt least 80% accuracy

Quality management system

A quality management system coordinates the processes that support reliable laboratory service.16

TermScope
Quality controlEvidence that a particular measurement procedure or item of equipment performs as expected
Quality assuranceEvaluation of a complete process, such as specimen transport or critical-result notification
Quality managementOrganization-wide planning, control, assessment, and improvement of laboratory quality

CLSI organizes the laboratory quality management system into twelve quality system essentials:16

Quality system essentialExample activity
OrganizationQuality planning and management review
Customer focusComplaint investigation and service feedback
Facilities and safetyHazard assessment and emergency preparation
PersonnelTraining and competency assessment
Purchasing and inventoryIncoming inspection and lot tracking
EquipmentInstallation, calibration, and preventive maintenance
Process managementProcedure design and method evaluation
Documents and recordsVersion control and record retention
Information managementAccess control, data integrity, and downtime procedures
Nonconforming event managementDetection, documentation, and investigation of failures
AssessmentsInternal audits and proficiency testing
Continual improvementCorrective action and measured process improvement

A quality indicator is a defined measurement of a laboratory process. Examples include specimen-rejection rate by cause, critical-result notification time, turnaround time for a defined test group, amended-report rate, and PT performance. For each indicator, the laboratory states the numerator, denominator, population, data source, review frequency, threshold, and responsible owner. Review frequency depends on the risk and the laboratory’s quality plan; no universal monthly committee rule applies to every indicator.16

After a process failure, the laboratory documents the event and contains the immediate risk. It investigates contributing causes, evaluates affected patient results, implements corrective action, and checks effectiveness. Root cause analysis is one tool for serious or recurring events. CLIA requires investigation and correction of identified problems, while organization-specific accreditation or safety rules determine when a formal root cause analysis is mandatory.3

The cost-of-quality model groups resources and losses into four categories:17

CategoryExamples
PreventionTraining, method evaluation, preventive maintenance
AppraisalControl materials, calibration verification, PT, internal audit
Internal failureRecollection, repeat testing, discarded reagents, downtime
External failureCorrected reports, patient follow-up, complaints, recalls, legal costs

These categories let a laboratory compare the cost of preventing or detecting a problem with the cost created after failure. Their relative size depends on the process and the event.

References

  1. National Institute of Standards and Technology. The normal distribution. NIST/SEMATECH e-Handbook of Statistical Methods. Accessed August 29, 2026.
  2. Clinical and Laboratory Standards Institute. Statistical Quality Control for Quantitative Measurement Procedures: Principles and Definitions. 4th ed. CLSI guideline C24. Published 2016.
  3. Electronic Code of Federal Regulations. 42 CFR part 493: Laboratory Requirements. Updated through August 27, 2026. Accessed August 29, 2026.
  4. Westgard JO, Barry PL, Hunt MR, Groth T. A multi-rule Shewhart chart for quality control in clinical chemistry. Clin Chem. 1981;27(3):493-501.
  5. Centers for Medicare & Medicaid Services. Individualized Quality Control Plans. Accessed August 29, 2026.
  6. Clinical and Laboratory Standards Institute. Measurement Procedure Comparison and Bias Estimation Using Patient Samples. 3rd ed. CLSI guideline EP09c. Published 2018.
  7. Clinical and Laboratory Standards Institute. User Verification of Precision and Estimation of Bias. 3rd ed. CLSI guideline EP15. Published 2014. Reaffirmed 2019.
  8. National Institute of Standards and Technology. Upper critical values of the F distribution. NIST/SEMATECH e-Handbook of Statistical Methods. Accessed August 29, 2026.
  9. Clinical and Laboratory Standards Institute. Interference Testing in Clinical Chemistry. 3rd ed. CLSI guideline EP07. Published 2018. Reaffirmed 2022.
  10. Clinical and Laboratory Standards Institute. Evaluation of Linearity of Quantitative Measurement Procedures. 2nd ed. CLSI guideline EP06. Published 2020.
  11. Clinical and Laboratory Standards Institute. Defining, Establishing, and Verifying Reference Intervals in the Clinical Laboratory. 3rd ed. CLSI guideline EP28-A3c. Published 2010. Reaffirmed 2020.
  12. National Institute of Standards and Technology. Critical values of the Student's t distribution. NIST/SEMATECH e-Handbook of Statistical Methods. Accessed August 29, 2026.
  13. College of American Pathologists. Proficiency Testing/External Quality Assessment Frequently Asked Questions. Accessed August 29, 2026.
  14. Electronic Code of Federal Regulations. 42 CFR §493.801: Condition: Enrollment and testing of samples. Updated through August 27, 2026. Accessed August 29, 2026.
  15. Electronic Code of Federal Regulations. 42 CFR part 493, subpart I: Proficiency Testing Programs for Nonwaived Testing. Updated through August 27, 2026. Accessed August 29, 2026.
  16. Clinical and Laboratory Standards Institute. A Quality Management System Model for Laboratory Services. 5th ed. CLSI guideline QMS01. Published 2019.
  17. American Society for Quality. Cost of quality. Accessed August 29, 2026.