A participant in one of our recent seminars asked a question that comes up surprisingly often: "We only have three units available for testing. How do we justify this to a Notified Body or the FDA?"
It is an excellent question, and one that cuts straight to the heart of a real tension in medical device development: statistical rigor vs. economic reality. If a single prototype costs €500,000 to manufacture, a sample size of n = 29 or n = 59 is simply not realistic.
The good news is that n = 3 can absolutely hold up under regulatory scrutiny,as long as the rationale is sound and clearly documented. The key is to shift the mindset from "more units equals stronger validation” to “maximum information per unit."
The following six strategies help you do exactly that. Along the way, this post covers:
-
Why the number of physical units is only one part of the equation
-
How to build statistical power through repeated testing on the same unit
-
Why a worst-case unit beats ten average units every time
-
How stress testing lets you argue for fewer samples
-
Why switching from attribute to variable data is a game changer
-
When analysis-based evidence is not just acceptable, but preferred
-
How to anchor everything in ISO 14971 for a bulletproof audit argument
Strategy 1: Redefine "Sample Size": Units vs. Repetitions
In medical device verification testing, sample size is not only about the number of physical units. It is also about how much relevant information each unit can generate.
When within-unit variation is the dominant source of variation for the performance endpoint, repeated testing on the same unit can contribute meaningful statistical information [9,10]. For example, testing 3 units multiple times may provide a stronger evidence base than testing 3 units once, especially when the endpoint is repeatable and the measurement system is well controlled.
However, repeated measurements are not the same as independent units. Thirty measurements from 3 units x 10 repetitions are not statistically equivalent to 30 independent units. If between-unit variation matters, treating all repetitions as independent observations will overstate precision and weaken the credibility of the rationale.
This is where a GR&R or repeatability study becomes useful [1]. It can help determine whether the main source of variation is the unit itself, the test method or normal repeatability within the same unit. If repeatability dominates and the analysis accounts for the repeated-measures structure, repeated testing can support a small-n argument. If part-to-part variation dominates, more physical units may be needed, no matter how many repetitions are performed.
Audit point: Repeated testing can strengthen a sample size design verification medical device rationale, but only when the variance structure is understood and documented.
Strategy 2: Use Worst-Case Testing for Medical Devices Deliberately
A well-selected worst-case unit can be more informative than several average units.
The purpose of worst-case testing is to challenge the design at the limits of its intended performance range [1,2,7,8]. Instead of testing only units that represent typical production conditions, you identify the parameters most likely to influence performance, such as dimensional tolerances, material properties, surface geometry or assembly conditions, and then select or build units that represent the most demanding combinations.
If these boundary units pass, the result provides stronger evidence that the design can perform under challenging conditions. This is especially relevant when the number of available units is limited, because the value of each tested unit matters more.
The argument only works if the worst case is technically justified. A unit is not “worst case” because it is available, expensive or difficult to manufacture. It must be linked to the design features, use conditions and failure modes that matter most.
Also, n = 1 per worst-case condition should not be treated as a general rule. A single worst-case unit may be useful evidence, but the rationale must explain why that unit represents the relevant boundary of the design space [1].
Audit point: Document which parameters define the worst case, why they are critical and how the selected units represent the relevant design limits.
Strategy 3: Make the Test Harder Than Real Life
Accelerated stress testing for a medical device can increase the value of each tested unit by exposing it to conditions more demanding than normal use [7,8].
This may include cycling beyond the expected service life, applying mechanical loads above the specification limit, testing at temperature or humidity extremes or combining multiple stress factors. The goal is not simply to make the test harsher. The goal is to create a justified link between the test conditions and the real-world stresses the device may encounter [7,8].
This distinction matters. A more severe test does not automatically justify a smaller sample size. The manufacturer still needs to explain why the stress level is relevant, how it relates to intended or foreseeable use and which failure mode it is meant to challenge.
A well-designed stress test can make n = 3 more meaningful than a larger test that only examines easy or poorly justified conditions. But the argument has to be built. “We tested harder, therefore we can test fewer units” is not enough.
Audit point: Stress testing strengthens a small sample size rationale when the elevated conditions are technically justified and connected to the risk management file.
Strategy 4: Switch From Pass/Fail to Variable Data
The choice between attribute and variable data can dramatically affect sample size requirements.
Attribute data, or pass/fail data, is easy to interpret, but it is statistically inefficient. A pass/fail result tells you whether the unit met the acceptance criterion, but it does not tell you how close the result was to the limit, how much margin exists or how the results are distributed.
Variable data, such as force, pressure, deflection, leakage rate or time to failure, provides much more information per unit tested. This is especially important when physical samples are limited. With variable data, you can estimate variation, calculate margins and apply statistical methods such as tolerance intervals [1,4].
A tolerance interval according to ISO 16269-6 can support population-level statements from a smaller number of measured results, provided the underlying assumptions are appropriate and the measurement system is validated [4]. This makes variable data one of the strongest tools available when sample size is constrained.
The tradeoff is that the measurement method must be suitable for the claim being made. Poor measurement data does not become strong evidence simply because it is numerical.
Audit point: Use variable data whenever possible, but make sure the measurement method is validated and appropriate for the acceptance criterion.
Strategy 5: Let Analysis Support the Physical Testing
When physical testing is constrained by cost or unit availability, engineering analysis can strengthen the evidence package.
This may include Finite Element Analysis (FEA), tolerance stack-up analysis, computational fluid dynamics (CFD) or comparison with equivalent predicate designs [5]. These methods are especially useful when they help explore worst-case conditions, sensitivity to design parameters or performance across a broader design space than physical testing alone can cover.
The FDA computational modeling guidance provides a risk-informed framework for assessing the credibility of computational modeling and simulation in medical device submissions [5]. ASME V&V 40 can also help structure the credibility argument by linking verification and validation activities to the level of reliance placed on the model and the consequence of an incorrect decision [11].
But analysis is not a shortcut around testing. A simulation or calculation must be credible. That means defining the context of use, documenting assumptions, justifying boundary conditions, using appropriate material properties and verifying and validating the model where needed.
Even a small number of physical tests can be powerful when used to confirm or calibrate the analysis. The weak argument is: “We ran a simulation and it looked fine.” The stronger argument is: “We used analysis to evaluate the design space, justified the model and confirmed the relevant assumptions with targeted physical testing.”
Audit point: FEA medical device verification and other analysis methods can reduce the physical test burden, but only when they are transparent, credible and linked to physical evidence.
Strategy 6: Use Standards When They Apply
Recognized standards can be one of the strongest ways to justify sample size.
Many medical device standards include defined sample sizes, sampling plans or test expectations for specific use cases [6]. When your device, test method and acceptance criteria fall within the scope of such a standard, the sample size is easier to defend because it is based on established expert consensus.
Examples may include standards or regulatory requirements for gloves, needle-based injection systems, small-bore connectors or stent securement testing. In these cases, the argument is not simply “we chose n = 3 because that is all we had.” It becomes: “We followed the sampling plan defined in the applicable standard.”
That said, using a standard does not remove the need for judgment. You still need to show that the standard applies to your specific device, test method and intended claim. If you deviate from the standard, or if the standard only partially covers your use case, the rationale must be documented.
Audit point: State which standard was used, why it applies and how any deviations were justified.
Examples include:
-
21 CFR 800.20 – Patient examination gloves and surgeons' gloves – sample plans and test method for leakage defects
-
ISO 11608-1:2022 – Needle-based injection systems for medical use
-
ISO 80369-7:2021 – Connectors for intravascular or hypodermic applications
-
ASTM F2394 – Balloon- Expandable Stent Securement
The Thread That Ties It All Together: ISO 14971
No sample size justification is complete unless it connects back to risk.
Regulators and auditors are not evaluating sample size in isolation. They are asking whether the evidence is sufficient for the risk associated with the design requirement. A low-risk cosmetic feature may not require the same level of evidence as a load-bearing structural element in a high-risk implantable device.
This is where ISO 14971 becomes the anchor [3]. The risk analysis should inform the verification strategy, the sample size rationale, the selection of worst-case conditions, the use of analysis and the level of statistical confidence required [3].
For an ISO 14971 sample size argument, the key is traceability: from hazard and foreseeable sequence of events, to design requirement, to verification method, to sample size rationale, to acceptance criteria. When that connection is clear, a small sample size becomes much easier to defend.
The central question is not simply: “Is n = 3 enough?”
The better question is: “Does this evidence package adequately address the clinical risk of failure?”
If the answer is clearly documented and supported by one or more of the strategies above, n = 3 can be a defensible starting point.
Putting It All Together
The question is never just “how many units do we have?” It is always: “How much do we know about this design, and is that enough for the clinical application at hand?”
For expensive medical device prototypes, a small sample size may be unavoidable. But “we only had three units available” is not a sample size justification. It is a constraint. The justification comes from how you design the evidence package around that constraint.
A strong rationale may combine several elements:
-
Repeated measurements when the variance structure supports them
-
Worst-case units selected from the relevant design limits
-
Stress testing connected to foreseeable use conditions
-
Variable data instead of pass/fail results
-
Tolerance intervals where appropriate
-
Engineering analysis supported by credible model assumptions
-
Recognized standards when they apply
-
A clear risk-based link to ISO 14971
In other words, n = 3 is not automatically acceptable and it is not automatically unacceptable. It depends on the device, the endpoint, the risk, the test method and the quality of the rationale.
A small sample size can satisfy an auditor when it is part of a documented, risk-based and technically justified verification strategy.
Why Your Visual Inspection Fails the Audit
The 6 steps that turn visual inspection into evidence that holds up – learn what auditors expect to see, the gaps they keep finding, and the framework that closes them.
Supplier Documentation
Stop Relying on Certificates – Learn the 5 Non-Negotiables Every Auditor Actually Probes
In this webinar, you'll learn the 5 Non-Negotiables that decide every ISO 13485 audit – and how to satisfy them without dragging your team onto the audit-prep treadmill.
Close the gaps. Pass the audit. Stay qualified.
Risk-Based Samples Sizes
Learn how to justify sample sizes using a clear, risk-based and statistically sound approach that reduces validation effort and cost.
Gain a practical framework you can confidently defend in audits across TMV, design verification, packaging, and process validation.
The 7 Deadly Sins of TMV
Stop reacting to audit findings – start leading.
Learn where MedTech companies repeatedly fail, what regulators truly expect, and how to set the right priorities – before inspections force your hand.
Join our free live webinar and walk away with:
- - Clear insights of the 7 most common mistakes in TMV
- - Practical principles you can apply immediately
- - Confidence to lead TMV decisions instead of firefighting them
Get audit-ready. Gain clarity. Take the lead. Secure your seat now!
Frequently Asked Questions
How many samples are needed for medical device verification?
There is no universal minimum sample size for medical device verification. The required number of samples depends on the device, the risk associated with the design requirement, the test method, the type of data collected and whether applicable standards define a sampling plan.
For high-risk claims or endpoints with high between-unit variation, more physical units may be needed. For lower-risk claims, highly repeatable endpoints or tests supported by variable data, worst-case units, standards or credible analysis, a smaller sample size may be defensible.
Can n = 3 be acceptable for a medical device sample size?
Yes, n = 3 can be acceptable in some medical device verification contexts, but not by default. The rationale must explain why three units are sufficient for the specific claim being made.
A defensible n = 3 sample size medical device argument usually needs more than simple pass/fail testing. It may rely on repeated measurements, worst-case units, stress testing, variable data, analysis-based evidence, applicable standards and a clear link to ISO 14971 risk management.
How do you justify sample size to FDA reviewers or a Notified Body?
To justify sample size to FDA reviewers or a Notified Body, document the logic behind the evidence package. This includes the risk associated with the requirement, the reason for selecting the units, the test method, the data type, the statistical approach and any standards or analysis used to support the rationale.
The strongest argument is not “this is the number of units we had available.” The strongest argument is “this is why the selected evidence is sufficient for the risk and intended claim.”
What is the minimum sample size for design verification?
There is no single minimum sample size for design verification that applies across all medical devices and all test types. A minimum may be specified by an applicable standard, but otherwise the number must be justified based on risk, variability, confidence requirements and the purpose of the test.
For expensive prototypes, the practical minimum may be driven by availability. The regulatory justification, however, must be driven by evidence quality and risk.
When should variable data be used instead of pass/fail data?
Variable data should be used whenever the measurement is feasible, reliable and relevant to the acceptance criterion. It provides more information than pass/fail data because it shows the actual result, the margin to the limit and the distribution of results.
This is especially valuable when the number of physical units is limited, because each tested unit needs to generate as much useful information as possible.
About the Author
Simon Föger is the founder and CEO of SIFo Medical. With more than a decade in medical device engineering, he has led validation, supplier qualification and compliance projects worldwide – from setting up MedTech manufacturing sites in Asia to training quality professionals at the TÜV SÜD Academy.
He shares his hands-on experience beyond consulting – in blog posts, in our newsletter, and as a guest on the Medical Device made Easy Podcast, where he talked about validation and supplier management.
Related SIFo Medical Resources
If you found this helpful, explore the related blog posts:
-
Statistical Tolerance Intervals: Explained simply and practically based on ISO 16269-6
-
Attribute Agreement Analysis (Go/No Go Gage Pass/Fail Test Systems)
Are you facing a design verification challenge and not sure how to structure your statistical justification? Contact SIFo Medical to build a strategy that satisfies both your budget and your auditors.
References
[1] Taylor, Wayne (2017). Statistical Procedures for the Medical Device Industry. Taylor Enterprises, Inc., www.variation.com
[2] FDA (1997). Design Control Guidance for Medical Device Manufacturers. U.S. Food and Drug Administration.
[3] ISO 14971:2019. Medical devices, Application of risk management to medical devices. International Organization for Standardization.
[4] ISO 16269-6:2014. Statistical interpretation of data, Part 6: Determination of statistical tolerance intervals. International Organization for Standardization.
[5] FDA (2023). Assessing the Credibility of Computational Modeling and Simulation in Medical Device Submissions, Guidance for Industry and Food and Drug Administration Staff. U.S. Food and Drug Administration, CDRH. Available at: https://www.fda.gov/media/154985/download
[6] ASTM F3334-19. Standard Practice for Finite Element Analysis (FEA) of Metallic Orthopaedic Total Knee Tibial Components. ASTM International. FDA-recognized standard (FR Recognition Number 11-368). Scope: "can be used for worst-case assessment within a series of different implants of the same implant design to reduce the physical test burden."
[7] IMDRF N47:2024. Essential Principles of Safety and Performance of Medical Devices and IVD Medical Devices. International Medical Device Regulators Forum. [Sections 5.1.2–5.1.8 require documented risk management, assessment of known and foreseeable hazards, and demonstration that performance is maintained throughout the expected service life under normal and foreseeable use conditions, including transport and storage.]
[8] GHTF SG1/N68:2012. Safety and Performance of Medical Devices. Global Harmonization Task Force. [Establishes that essential requirements for safety and performance must be demonstrated across the product lifecycle, including under external influences, and requires risk-based documentation of design and manufacturing evidence.]
[9] FDA (1998). E9 Statistical Principles for Clinical Trials, Guidance for Industry. U.S. Food and Drug Administration. [Recognizes repeated-measures designs as valid study designs; requires correct modeling of the variance structure and pre-specification of statistical assumptions.]
[10] Bloch, D.A. & Lai, T.L. (2004). "Sample size requirements in trials using repeated measurements." Statistics in Medicine. [Demonstrates that within-individual variability and number of measurement time points both influence required sample size, but only when the design correctly accounts for the variance structure.]