Proficiency testing (PT) is the process by which a calibration laboratory compares its measurement results against those of other labs to confirm it’s performing within accepted limits. ISO/IEC 17025:2017, clause 7.7.2, makes this mandatory: accredited labs must participate in PT schemes or, where none exist, document a justified alternative. It’s one of the clearest ways a lab can demonstrate that its results are technically valid, not just administratively compliant. Understanding how ISO 17025 accreditation impacts calibration starts with recognizing that PT is a core pillar of that system, not an optional add-on.

Key Takeaways

  • ISO/IEC 17025:2017 clause 7.7.2 requires accredited labs to participate in proficiency testing as part of ongoing performance monitoring.
  • Proficiency testing (PT) is a specific type of interlaboratory comparison (ILC), scored against pre-established criteria using z-scores or En numbers.
  • A z-score of |z| ≤ 2 is satisfactory; |z| ≥ 3 triggers mandatory corrective action under ISO 13528:2022.
  • ISO/IEC 17043:2023, published May 2023, is the current standard governing PT providers — updated to align with ISO/IEC 17025:2017.
  • Over 810 PT providers are accredited by ILAC MRA signatories worldwide, serving more than 114,600 accredited laboratories.
  • An unsatisfactory PT result doesn’t automatically mean lost accreditation, but it requires documented root-cause analysis and verified corrective action.
Calibration lab technician performing proficiency testing measurement on round robin artifact

What Is Proficiency Testing and Why Does It Matter?

A calibration lab can pass every internal audit, maintain pristine documentation, and still produce measurements that don’t agree with anyone else’s. Internal checks catch procedural errors, but they can’t catch systematic bias in your measurement system. That’s the gap PT fills.

At its core, proficiency testing is an external check. A PT provider distributes the same artifact, reference material, or data set to multiple labs, collects everyone’s results, and evaluates performance against a pre-established reference value and acceptance criteria. If your result falls within the acceptable band, you’ve demonstrated external validity. If it doesn’t, you have a problem worth finding.

The distinction between PT and interlaboratory comparison (ILC) is worth clarifying, because the terms are often used interchangeably when they shouldn’t be. An ILC is the broader concept: the organization, performance, and evaluation of measurements by two or more laboratories on the same or similar items. PT is a specific, formal type of ILC managed by an independent third-party provider, evaluated against pre-established criteria, and typically scored using statistical tools like z-scores or En numbers. Not every ILC is a PT scheme, but every PT scheme is an ILC.

Why does this matter for labs? Because ISO/IEC 17025 uses both terms deliberately. When PT is available, labs must use it. When it isn’t, an ILC or other documented comparison may be acceptable, but only with justification. The standard expects accreditation bodies and auditors to distinguish between them, and so should your quality team.

The scale of the global PT infrastructure is larger than most lab managers realize. As of 2024, over 810 proficiency testing providers are accredited by ILAC MRA signatories, serving a global community of more than 114,600 accredited laboratories across 122 economies. For most measurement disciplines and most industries, finding an appropriate PT scheme isn’t the challenge. Selecting the right one for your scope is.

What ISO 17025:2017 Requires (Clause 7.7)

Clause 7.7 of ISO/IEC 17025:2017 covers the monitoring of the validity of results. Clause 7.7.1 deals with internal monitoring, things like control charts, check standards, and repeat measurements. Clause 7.7.2 is specifically about external comparison, and this is where proficiency testing lives.

The language in clause 7.7.2 is direct: laboratories shall monitor their performance by comparison with results of other laboratories. This monitoring shall be planned and reviewed, and it shall include participation in proficiency testing. The standard doesn’t specify frequency, it leaves that to the accreditation body, but it is unambiguous that participation is required, not optional.

What happens when PT isn’t available? The standard does allow for alternatives, intralaboratory comparisons, certified reference materials, or other documented mechanisms, but these alternatives must be justified and documented. An auditor reviewing your quality records will expect to see that you actively sought a PT scheme, confirmed none existed or was appropriate for your scope, and selected the documented alternative on that basis.

ILAC P9:01/2024, published January 2024, is the current ILAC policy governing how accreditation bodies apply PT requirements to conformity assessment bodies, including calibration labs. It’s enforced by all 121 signatories to the ILAC Mutual Recognition Arrangement (MRA) and references ISO/IEC 17025:2017, ISO 15189:2022, and ISO/IEC 17043:2023 as the underlying normative standards. If your accreditation body is an ILAC MRA signatory, this policy directly shapes your PT obligations. Source: ILAC P9:01/2024.

One practical implication many labs overlook: the planning requirement. Clause 7.7.2 says PT participation shall be planned and reviewed. That means your quality management system should include a documented PT schedule, with specific schemes, frequencies, and responsible parties named. An ad hoc approach, enrolling in rounds as you hear about them, doesn’t satisfy the intent of the standard. Your accreditation body will look for evidence of a systematic, forward-looking program.

For labs operating under scope extensions or seeking to add new measurement disciplines, PT participation in that discipline is typically a prerequisite. Demonstrating competence in a new area requires external evidence, and a satisfactory PT result in the relevant scheme is the most direct form of that evidence.

Types of Proficiency Testing Schemes

Not all PT schemes work the same way. The format depends on the measurement type, the analyte or artifact being measured, and the logistics the PT provider can manage. ISO/IEC 17043:2023, the governing standard for PT providers, describes several scheme types. The most common in calibration contexts are sequential (round-robin) comparisons, simultaneous distribution schemes, and reference material-based schemes.

Sequential Round-Robin (Interlaboratory Comparisons)

In a sequential ILC, a single artifact, a gauge block, a reference weight, a calibrated oscilloscope, travels from lab to lab in a defined order. Each lab measures the item, records results, and ships it to the next participant. The pilot lab or PT provider holds the reference value, typically established at a national metrology institute (NMI) or reference lab, and compares each participant’s result against it at the end.

Sequential schemes are common for physical measurement artifacts where the item can be transported without significant change. The trade-off is time: a round-robin involving 15 labs takes months, during which the artifact’s condition may drift. Good PT providers build in recirculation checks and reference re-measurements to detect this.

Simultaneous Distribution Schemes

In simultaneous schemes, the PT provider distributes identical or near-identical items to all participants at the same time. This is practical for consumables, reference materials, or items that can be manufactured to tight tolerances in quantity. All labs measure within the same time window, eliminating the drift risk of sequential circulation.

Reference Material-Based Schemes

For chemical, physical-chemical, or dimensional measurements, certified reference materials (CRMs) can serve as the basis for PT. Labs measure a CRM with a certified value and uncertainty, then compare their result against the certified value using En numbers or z-scores. This approach works well for labs where the measurement type doesn’t lend itself to artifact circulation, such as analytical chemistry or certain electrical standards.

Pilot Lab and NMI Key Comparisons

At the highest metrological level, national metrology institutes participate in key comparisons organized by the BIPM (Bureau International des Poids et Mesures) and regional metrology organizations. These comparisons establish the degree of equivalence between national measurement standards and underpin the entire traceability chain that flows down to accredited calibration labs. The results are published in the BIPM key comparison database (KCDB). Understanding how working and reference standards relate to this hierarchy helps explain why NMI comparisons matter even for commercial labs that never participate in them directly.

ISO/IEC 17043:2023 was revised in May 2023, replacing the 2010 edition. Key changes include structural harmonization with ISO/IEC 17025:2017, updated terminology aligned with ISO 13528:2022, and an expanded scope covering conformity assessment activities beyond just testing and calibration. If your PT provider’s accreditation certificate still references the 2010 edition, that’s a conversation worth having. Source: ANAB, ISO/IEC 17043:2023 Main Changes.

How Labs Select the Right PT Program

Selecting a PT scheme isn’t just about finding one that covers your measurement discipline. It’s about matching the scheme to your accreditation scope, your customer base, and the performance claims you make. A torque lab that primarily calibrates tools used in aerospace assembly has different PT needs than one serving general industrial maintenance.

Start with your accreditation body’s PT requirements. Most ILAC MRA signatories publish a list of approved or recognized PT providers for different measurement categories. In the US, A2LA and NVLAP both maintain guidance on PT participation requirements and acceptable schemes. In the UK, UKAS publishes similar guidance. These lists aren’t exhaustive, and labs can often propose alternative schemes, but starting with your AB’s recognized list is the path of least resistance during assessment.

One factor labs underestimate is the reference value’s uncertainty in the PT scheme. A scheme where the reference value has uncertainty comparable to, or larger than, your own measurement uncertainty will produce uninformative results. Before enrolling, request the PT provider’s uncertainty budget for the scheme. If their reference uncertainty is more than a third of your claimed uncertainty, the scheme may not be sensitive enough to detect realistic performance problems in your lab. This is directly relevant to measurement uncertainty in calibration, where the relationship between reference uncertainty and test uncertainty determines whether a comparison is even statistically meaningful.

Frequency is another consideration the standard leaves to the accreditation body, but a common baseline for most disciplines is at least one PT round per accreditation cycle (typically four years), with many ABs requiring annual or biannual participation for high-risk or technically demanding scopes. Some labs choose to participate more frequently than required, using PT as an ongoing internal check rather than a compliance checkbox.

Geographic and logistical factors matter too. For international labs or labs with multiple sites, schemes offered through ILAC MRA-accredited providers carry mutual recognition across member economies, meaning satisfactory results in one jurisdiction support credibility in others. Over 121 economies participate in the ILAC MRA, making provider selection a more global decision than it once was. Source: ILAC Facts and Figures, 2024.

Finally, consider the scheme’s historical performance data. Reputable PT providers publish summary statistics from previous rounds, showing the distribution of participant results. If 40% of labs scored unsatisfactory in the last round, that’s either a sign of a poorly designed scheme or a discipline with widespread competency issues. Either way, knowing this before you enroll helps you set realistic expectations and prepare your team appropriately.

lab technician working

What Happens When PT Results Are Unsatisfactory

An unsatisfactory PT result, a z-score of |z| ≥ 3 or an En number of |En| > 1.0, is classified as an action signal. This doesn’t immediately suspend your accreditation, but it does trigger a formal nonconformance process that your accreditation body will review. How you respond matters as much as the result itself.

The required response follows the same corrective action framework that governs any ISO 17025 nonconformance. You must document the nonconformance, conduct a root-cause investigation, formulate corrective actions, implement them, and then provide objective evidence to your accreditation body that the actions were effective. Repeating the PT round and achieving a satisfactory result is the most direct form of that evidence, though not always the only option.

Root-cause analysis for PT failures typically falls into a handful of categories: equipment out of tolerance (the instrument used for the PT measurement was itself uncalibrated or drifting), procedural error (the method wasn’t followed as written), personnel competence gaps (the technician performing the measurement hadn’t been properly qualified for that task), or reference standard issues (the traceability chain used for the PT measurement had a break or an unrecognized uncertainty contribution). The relationship between calibration vs. verification is relevant here: a verification check that passed doesn’t rule out calibration drift as the root cause.

What are the consequences if the corrective action isn’t adequate? Accreditation bodies have several tools available. They can issue a corrective action request (CAR) with a defined response deadline. They can reduce the lab’s scope of accreditation to exclude the affected measurement discipline. In persistent cases, they can suspend or withdraw accreditation entirely. Under the ILAC MRA framework, full reassessment of a lab’s accreditation occurs before the end of each four-year cycle, and unresolved PT deficiencies will surface at that point if not before.

A warning signal (|z| between 2 and 3) doesn’t require formal corrective action, but it shouldn’t be ignored. Labs that consistently score in the warning zone are tracking toward an action signal, and a pattern of warning scores may itself attract scrutiny from an assessor. Documenting your awareness of a warning result and any investigative steps you took is good practice even when no formal response is required. For ISO 17025-accredited calibration with full traceability documentation, contact Micro Precision.

How to Read and Interpret a PT Certificate

When your lab completes a PT round, you receive a certificate or report from the PT provider summarizing your performance. Reading this document correctly is a skill in itself. The components vary by provider, but there are standard elements that should appear in any well-run scheme under ISO/IEC 17043:2023.

The assigned value (sometimes called the reference value or consensus value) is the target your result is compared against. This may be a value certified by an NMI, a mean derived from a reference lab’s measurements, or a robust statistical consensus of all participant results. The method used to establish the assigned value directly affects how informative your z-score is. A consensus mean is less metrologically rigorous than an NMI-certified value, but it’s often what’s available for specialized or novel measurement types.

The standard uncertainty of the assigned value (u_ref) is a critical number that many labs skim past. It contributes to the denominator of the En number calculation: En = (x_lab – x_ref) / sqrt(U_lab² + U_ref²). If u_ref is large relative to U_lab, your En number will be compressed toward zero regardless of how well your result actually compares. That sounds favorable, but it means the scheme has low sensitivity and may not be catching real problems. Reviewing calibration certificates and issuing authority context helps when evaluating whether the PT provider’s own reference values are adequately traceable.

Under ISO 13528:2022, the statistical companion standard to ISO/IEC 17043, z-scores are calculated as z = (x – X) / σ, where x is the participant’s result, X is the assigned value, and σ is the standard deviation for proficiency assessment (set by the PT provider, not derived from participants’ data alone). A |z| ≤ 2 is satisfactory, |z| between 2 and 3 is a warning signal, and |z| ≥ 3 is an unsatisfactory action signal. For calibration labs, En numbers (which incorporate both the lab’s and the provider’s measurement uncertainties) are often used instead of or alongside z-scores. Source: ISO 13528:2022 and ISO OBP/Shapypro z-score methodology documentation.

Beyond your individual result, most PT reports include a graphical summary of all participants’ z-scores or En numbers, typically as a ranked bar chart or scatter plot. You won’t see other labs identified by name (anonymization is standard practice), but you can see where your result falls in the distribution. Consistently landing in the low-positive or low-negative range (z between 0 and 1.5) across multiple rounds suggests your measurement system is well-controlled. Consistently landing near the boundary of the satisfactory zone should prompt you to investigate before you cross it.

PT certificates should be retained as quality records. Your accreditation body will request them during assessments, and they form part of the objective evidence that your monitoring of result validity is active and effective. Some ABs require labs to submit PT results directly; others review them only at assessment. Know which approach your AB uses and document your PT records accordingly. Understanding the full scope of your instrument calibration services obligations under ISO 17025 means treating PT records with the same rigor as calibration certificates.

z score performance zones under ISO 13528 2022

Calibration certificates your auditor can accept. ISO 17025 accreditation at 50+ labs.

Micro Precision’s ISO/IEC 17025:2017-accredited instrument calibration services are delivered across a global network of 50+ labs, covering mechanical, electrical, temperature, pressure, dimensional, and RF/microwave disciplines. Request a quote and we’ll confirm scope and turnaround.

Frequently Asked Questions

Proficiency testing is an external performance check where an independent PT provider distributes the same measurement artifact or reference item to multiple labs, collects results, and scores each lab against a reference value using z-scores or En numbers. ISO/IEC 17025:2017 clause 7.7.2 requires accredited calibration labs to participate in PT schemes as part of ongoing result validity monitoring. As of 2024, over 810 accredited PT providers support this global system.

Clause 7.7.2 requires that accredited labs monitor performance through comparison with other labs, including participation in proficiency testing. This monitoring must be planned and reviewed, not ad hoc. When no suitable PT scheme exists, the lab must document a justified alternative such as an intralaboratory comparison or reference material measurement. ILAC P9:01/2024 governs how accreditation bodies apply this requirement across all 121 ILAC MRA signatory economies.

An interlaboratory comparison (ILC) is the broad category: any organized measurement by two or more labs on the same or similar items. Proficiency testing is a specific type of ILC managed by an independent third-party provider, evaluated against pre-established acceptance criteria, and scored statistically. All PT is ILC, but not all ILC qualifies as formal PT. ISO/IEC 17025 uses both terms, and auditors recognize the distinction.

A z-score compares a lab’s result to the assigned reference value, normalized by the standard deviation for proficiency assessment set by the PT provider. Under ISO 13528:2022, |z| ≤ 2 is satisfactory, |z| between 2 and 3 is a warning signal requiring investigation, and |z| ≥ 3 is an unsatisfactory action signal requiring formal corrective action. For calibration labs, En numbers (which incorporate measurement uncertainty) are also commonly used, with |En| ≤ 1.0 considered satisfactory.

The lab must document the nonconformance, conduct a root-cause investigation, implement corrective actions, and provide objective evidence of effectiveness to the accreditation body, typically by repeating the PT round. Failure to submit adequate corrective action can result in scope reduction, accreditation suspension, or withdrawal. Full reassessment occurs before the end of each four-year accreditation cycle, and unresolved PT deficiencies will surface at that assessment if not before.

ISO/IEC 17025:2017 doesn’t set a fixed frequency; that’s determined by the accreditation body. Most ILAC MRA signatories require at least one PT round per four-year accreditation cycle for each measurement discipline, with many requiring annual or biannual participation for higher-risk scopes. Labs seeking scope extensions must typically demonstrate satisfactory PT results in the new discipline before accreditation is granted for that area.

As of 2024, over 810 PT providers are accredited by ILAC MRA signatory bodies worldwide. In the US, NIST operates comparison programs, and providers accredited by A2LA and NVLAP are commonly used. Internationally, EUROLAB, APMP, and other regional metrology organizations coordinate PT schemes. ILAC maintains a database of accredited PT providers at ilac.org. Selecting an ILAC MRA-accredited provider ensures results carry international recognition across all 121 member economies.

The En number (normalized error) is a statistical measure used specifically in calibration PT when measurement uncertainty is a key factor. It’s calculated as En = (x_lab – x_ref) / sqrt(U_lab² + U_ref²), where U_lab and U_ref are the expanded uncertainties of the lab and the reference value, respectively. |En| ≤ 1.0 is satisfactory. The En number is preferred over z-scores in calibration contexts because it accounts for the uncertainty claims of both the participant and the PT provider.