Pharma Engineering Insights

Pharmaceutical Water Sampling Plans: How to Select and Justify Sampling Points

A sampling plan inherited from qualification does not survive an inspection. How to build the rationale that ties each point to a precise question, and how to tell a sample representing the system from one representing the water the process actually uses.

G GuideGxP 16 min read
✓ Official sources and references ✓ Practical approach ✓ For pharmaceutical professionals
GUIDEGXP · PRACTICAL GMP INSIGHTS
Schema dei punti di campionamento lungo un loop di distribuzione di acqua farmaceutica, dalla generazione alle utenze

Inspectors rarely ask why a sample went out of limit. They ask why that point is in the plan and the other one is not. That question dismantles most sampling plans, because the honest answer in too many sites is that the plan was inherited: born during qualification as a list of the points that happened to be physically available, frozen into an SOP, extended by analogy to every new branch, and never revisited. No document explains which question each point is there to answer.

The problem is not a paperwork one: a plan built that way produces two opposite ways of being wrong. In the first, every sample passes because the sampling technique systematically removes the contribution of the point of use, while the water reaching the process travels through a hose nobody samples. In the second, data drift upward because of poor sampling technique, and the site investigates a system that is under control until someone proposes widening the limits.

A defensible sampling plan comes from an explicit choice, not from a list: which questions the plan must answer, how each point is justified, how sampling technique determines what the data mean, and on what basis frequency is set and reviewed.

The two questions a plan must keep apart

A pharmaceutical water sample can answer two different questions, and confusing them is the most common structural error.

Question 1 — is the system under control? The sample represents the water held in the generation, storage and distribution system. It feeds the trending that alert levels rest on, detects progressive deterioration, and verifies sanitisation effectiveness. The technique is designed to eliminate the contribution of the sampling point itself: dedicated and sanitised valve, defined flushing, standardised aseptic technique. A high result here points to the system.

Question 2 — what water is the process actually using? The sample represents the water reaching the product through the same physical chain production really uses: same connection, same hose, same opening procedure, the same flushing prescribed in production — no more, no less. The technique is designed to include everything sitting between the loop and the product. A high result here may point to the system, but equally to the hose, the connection or the operator's practice.

These are two different plans, with different points, techniques, frequencies and investigation criteria. A “system” sample says nothing about the water the process receives; a “process” sample is a poor indicator of loop condition, because noise from the transfer device can dominate the signal. Neither question is optional: water quality is an element of the contamination control strategy under Annex 1 clause 2.5(v), and a CCS that demonstrates loop control without demonstrating control of transfer to the user point has a hole in the middle. A drafting rule follows: the plan must state, for each point, which of the two questions it answers and with which technique. A list of points with no question attached is not justifiable, because there is no criterion for deciding whether the set is sufficient.

Sample point and point of use are not synonyms

Point of use denotes a process function: where water leaves the system to enter an operation. Sample point denotes a control function: a point equipped to draw a sample representative of something. They often coincide, but out of physical convenience, not logical equivalence.

Three cases make the distinction operational. Many user points are served by a dedicated sampling valve upstream of the process connection: that is a sample point answering question 1, not a point of use. Some user points have a point-of-use filter: sampling downstream answers question 2 but masks loop condition, and using that data for system trending builds a blind trend. Other points are not user points and never will be — generation outlet, loop return — and serve question 1 only. The number of sample points is therefore not derivable from the number of points of use in either direction: it has to be built from a rationale.

The applicable regulatory framework

No text in force publishes a list of mandatory points or a table of universal frequencies: the texts define objectives and leave the rationale to the site.

Annex 1 of EudraLex Volume 4 (C(2022) 5938 final, in operation since 25 August 2023) requires at clause 6.13 regular and ongoing chemical and microbiological monitoring, with alert levels based on initial qualification data, periodically reassessed on the basis of requalification, routine monitoring and investigations. Clause 6.14 requires alert excursions to be documented, reviewed and investigated, distinguishing an isolated event from an adverse trend or system deterioration. Clause 6.8 requires water systems to be qualified and validated for physical, chemical and microbiological control taking seasonal variation into account. Clause 6.15 requires WFI systems to include continuous monitoring, such as TOC and conductivity.

Read together they describe the architecture of a sampling plan without dictating its content: levels derive from system data, not from an external table; the plan must produce data sufficient to tell an isolated event from a trend; the collection window must cover seasonal variability; and for WFI part of the monitoring is not sampling at all, but continuous in-line measurement.

The EMA Q&A EMA/INS/GMP/443117/2017 on the production of WFI by non-distillation methods (in force since 1 August 2017) is the only text in this set that indicates a frequency: for those systems it expects extended testing, daily testing of all critical points in the initial phase and data over approximately one year to capture seasonal variation, and it sets as a minimum an annual evaluation of monitoring effectiveness. Two clarifications prevent over-extrapolation: the daily testing expectation is framed for non-distillation WFI systems and concerns the initial phase, not routine operation; applying the same model to other systems is a site choice, reasonable but not a requirement, and must be documented as such. The same document states explicitly that increasing limits is not good practice and may mask a failing system.

WHO TRS 1033, Annex 3 (2021) closes the loop on method: §12.8 establishes that alert and action levels are determined from reported historical data, and §4.3.8 that alert and action limits for bulk purified water derive from system knowledge and data trending. The informational chapter USP <1231> (official since 1 December 2021) explicitly covers both online and offline sampling and recalls that users establish in-house specifications or fitness-for-use microbial levels; the action levels it reports are not binding. USP <643> also gives a methodological cue by subtraction: it states that it intentionally says nothing about how often the system suitability test should be run, leaving the frequency to the user's risk assessment. That is exactly the logic governing a sampling plan.

The risk rationale is built with the intended tool, ICH Q9(R1) (Step 4, 18 January 2023), applied through the method of quality risk management on water systems; the dedicated industry reference is the ISPE Good Practice Guide Sampling for Pharmaceutical Water, Steam, and Process Gases, an industry guide and not a regulatory text.

Mapping the points: what each location answers

Every location answers a specific question and leaves others open. Which locations are actually included depends on the system architecture.

LocationWhat it answersWhat it does not say
Feed water and pretreatment stagesPerformance of upstream stages, seasonal variation of the incoming loadNothing about the quality released for use: these are process controls
Generation outletGeneration performance in isolation: separates a production problem from a distribution problemNothing about tank and loop, where most microbiological problems arise
Storage tankCondition of stored water, effect of residence times, protection of the volumeNothing about downstream branches and user points
Loop supplyQuality entering distribution: the reference for reading every downstream pointNothing about accumulation along the route
Loop returnThe most sensitive point on overall distribution condition: it integrates the whole routeIt does not localise: a degraded return does not say which branch caused it
Branches and sub-loopsContribution of branches: low flow, intermittent use, terminal legsNothing about the water the operator actually transfers to the process
Critical user point, system sampleLoop condition at the furthest or highest-risk point; feeds system trendingNothing about hose, connection and sampling practice in production
Critical user point, process sampleQuality of the water actually used, with the operator's physical chain and techniqueIt does not isolate the cause among loop, hose, connection and technique

Loop return and user points are not alternatives: the former gives sensitivity, the latter give localisation and product relevance. And rarely used outlets deserve attention out of proportion to their frequency of use, because stagnation between uses is precisely the condition monitoring has to catch. The geometry that makes a branch critical — length, flow, terminal leg — belongs to distribution loop design, and the plan must read it rather than ignore it.

Building the risk rationale

Justifying a point is not a sentence in an SOP: it is a traceable assignment linking that point to identified risk factors — product impact of the operation served, hydraulic position relative to the flow regime established at qualification, actual coverage by the sanitisation cycle, pattern of use, historical data for the point, and alternative observability, that is whether the point is already covered by continuous in-line measurement or the sample is the only available window. The matrix is for the site to complete. The weight column is deliberately empty: weights depend on context, product type and architecture, and must be minuted and approved before scores are assigned.

FactorEvidence supporting the scoreWeightScoreWeighted
Product impact of the operation servedWater quality required, type of contact, process stage
Hydraulic position and flow regimeP&ID, qualification flow data, branch length
Coverage by the sanitisation cycleCycle mapping, parameters achieved at the point
Pattern and frequency of useUsage records, stagnation periods, transfer devices
History of the pointTrends, excursions, investigations, requalification outcomes
Redundancy of controlIn-line measurement or equivalent points upstream and downstream
Detectability of the anomalyTime from sampling to result, capability of the method

The output is not a number but a classification — points under extended, routine, periodic or event-driven monitoring — from which frequency, analytical panel and escalation criteria follow. The document must show the chain, not just the outcome.

Sampling technique: the variable that decides what the data mean

The most defensible sample, drawn with an undefined technique, produces meaningless data: microbiologically, the variability introduced by technique can exceed that of the system being measured. The variables must be defined, validated and trained; none has a universal value to copy.

Flushing. This is the variable that changes the question the sample answers. Prolonged flushing discards the content of the terminal zone and returns loop water: question 1. No flushing, or the flushing prescribed in production, returns the water the process receives: question 2. There is no absolutely correct duration or volume; there is an obligation to state the choice, justify it against the question, make it reproducible across operators and shifts, and keep it stable — because a silent change in practice produces a level shift in the trend that will be read as a system event.

Point preparation and asepsis. Sanitisation or disinfection of the valve before sampling, management of agent residue, the operator's aseptic technique and protection of the point between samples are factors the method must fix. Where the system is chemically sanitised or ozonated, the microbiological sample requires an appropriate neutraliser, verified not to interfere with the analytical method; without neutralisation what is counted is the residual effect of the agent, not the system population.

Containers, volumes, hold time and transport. Container material, sterility, presence of the neutraliser and absence of interference with the measured parameter — TOC is particularly sensitive to contamination from the container and from handling — are attributes to specify and verify, not to assume. Volume follows from the method and from the level that must be detectable: for WFI the non-binding action level reported in USP <1231> is expressed per 100 mL, with direct consequences for volume and filtration technique.

Hold time between sampling and the start of analysis has no regulatory value: it must be set by the site and supported by data showing that the result does not change significantly within the stated interval, under the storage and transport conditions actually used. The variables that matter are water type, expected population, container material, temperature and its stability along the route; transport itself must be defined, monitored and documented. The chain from tap to incubator is part of the analytical method, not support logistics: every uncontrolled link becomes an alternative explanation available during an investigation, and therefore a weakening of the data even when the data comply.

Qualification of the sampler. The operator is a variable of the method: documented training, practical verification of execution and periodic observation in the field are controls on the data, not formalities. An excellent plan executed by unqualified personnel produces noise that no statistical analysis recovers.

Online and offline are not interchangeable

USP <1231> covers both. The decisive difference is temporal: in-line chemical measurement — TOC and conductivity, required as continuous monitoring for WFI systems by Annex 1 clause 6.15 — gives a real-time signal, whereas the microbiological result arrives days after sampling, once the water has already been used. The plan must make that asymmetry explicit and build the strategy on it: microbiological control is predominantly preventive — design, thermal regime, sanitisation, flow — and sampling verifies afterwards that it is holding. Expecting microbiological sampling to protect an individual batch asks the method for something it cannot give. The in-line architecture is covered in the article on TOC and conductivity monitoring, the microbiological side in the article on microbiological and endotoxin control.

Frequency: what is verifiable and what must be justified

“How often do we sample” has no compendial answer: the texts in force supply a method for building one and two verified time references. The method is that of Annex 1 6.13 and WHO TRS 1033 §12.8 and §4.3.8: levels and monitoring structure derive from initial qualification and from trending, and are periodically reviewed. The verified time references are those of the EMA Q&A 443117/2017: in the initial phase, for non-distillation WFI systems, daily testing of all critical points and data over approximately one year to cover seasonal variation — consistent with Annex 1 6.8 — and, as a minimum, an annual evaluation of monitoring effectiveness.

Everything else must be justified per site. No requirement imposes a weekly, monthly or rotating frequency: these are widespread industry practice, not prescriptions, and whoever adopts them must be able to show why that cadence is sufficient to detect the event they fear, given the method's response time and the expected dynamics of the system.

  1. Initial phase: high frequency at all points to characterise the system and generate the data set from which alert levels derive, over a window covering seasonal variability.
  2. Transition: reduction only once data support it, with criteria defined a priori and a point-by-point rationale. It is a change to the control system and goes through change control.
  3. Routine: frequency and analytical panel differentiated by risk class, with escalation criteria that return a point to high frequency when defined triggers occur.
  4. Review: periodic verification that the frequency still answers the question, returning to the initial phase when the system changes substantially.

A sanity criterion: if the chosen frequency would not allow an isolated event to be distinguished from an adverse trend — the distinction required by Annex 1 6.14 — then it is too low for the stated purpose, however widespread it may be in the industry.

Worked example: Site Delta

Site Delta is a fictitious site, used here only as an example. It has a hot WFI loop serving a sterile area and a cold, chemically sanitised PW loop serving non-sterile production and equipment washing. The current plan lists twenty-two points, all sampled with the same method and the same frequency, and the SOP prescribes flushing before every sample.

The review starts from the question, not from the points: all twenty-two samples answer question 1, and none represents the water the process receives, even though four user points are served through hoses connected and disconnected at every use. Three outcomes follow. A second category of samples is introduced, drawn at user points with the production physical chain and technique, with its own investigation criteria: an out-of-limit result in this category starts from the transfer device, not from the loop. The risk classification redistributes frequencies, raising them on intermittently used branches and terminal legs and lowering them where loop return and in-line measurement already provide coverage, with every reduction documented and handled through change control. And the flushing instruction is rewritten to separate the two cases, because applied indiscriminately it made answering question 2 structurally impossible. The final number of points changes little; what changes is that each one now has a line stating which question it answers and why it is there.

Reviewing the plan

Annex 1 6.13 makes the periodic review of alert levels explicit, on the basis of requalification, routine monitoring and investigations; the same applies to the structure of the plan. The triggers to be codified in procedure recur: physical modifications to the system (new branch, new user point, removal of a leg, replacements that change the hydraulic profile); change of use of a point, which shifts its risk class even if no pipework was touched; change of the sanitisation or thermal regime, which changes cycle coverage; adverse trends, repeated excursions and investigation outcomes pointing to blind spots; changes in source water relative to the qualification characterisation; periodic requalification and evaluation of monitoring effectiveness, with the annual minimum indicated by the EMA Q&A for systems in scope.

The review must also be able to conclude in a restrictive direction: a plan that shrinks at every cycle and never grows is optimising analytical workload, not control. When data drift, the correct response is to investigate the system and revise the plan, not to move the limits. Reading anomalous outcomes is covered in the article on biofilm, rouging and investigation; the full set of topics is collected in the Pharmaceutical Water & WFI Systems hub.

Common mistakes and red flags

  • A plan inherited from qualification and never reviewed, with points whose original justification can no longer be retrieved; no distinction between system and process samples, with the site convinced it has two layers of coverage.
  • Flushing prescribed uniformly, which makes it structurally impossible to assess the water actually used by the process.
  • Points added by analogy with every new user point, without reassessing the risk classification of the whole; sampling downstream of a point-of-use filter used for system trending.
  • Hold time and transport conditions undefined or unsupported by data: they offer an alternative explanation for every anomalous result.
  • Frequencies presented as a regulatory requirement when they are industry practice, or reduced without change control and without a data-based rationale.
  • Rarely used outlets excluded from the plan precisely because they are rarely used; a discussion about limits opened before the discussion about the plan after a run of unfavourable results.

Three questions to ask in an audit: for any given point, which question it answers and on what evidence it entered the plan; who took the last sample there and with what technique; when the plan was last reviewed and what changed. If the answers do not come quickly and consistently, the problem is not the plan: it is that there isn't one.

If you work on decisions of this kind, The Pragmatic GMP collects technical and regulatory analysis on GMP systems with the same approach.

Key takeaways

  • A sample can represent the system or the water used by the process: two different questions, different techniques, different plans. The plan must state for each point which of the two it answers.
  • Sample point and point of use are not synonyms: the number of points does not follow from the number of user points but from a risk rationale built with ICH Q9(R1) and reviewed at every system change.
  • Annex 1 requires alert levels derived from initial qualification and periodically reassessed (6.13), investigation of excursions distinguishing isolated event from trend (6.14), qualification taking seasonal variation into account (6.8) and continuous monitoring such as TOC and conductivity for WFI systems (6.15).
  • The only verified frequencies are those of the EMA Q&A 443117/2017 for non-distillation WFI systems: daily testing of critical points in the initial phase, data over approximately one year, and as a minimum an annual evaluation of monitoring effectiveness. Any other cadence must be justified per site.
  • Flushing, point preparation, container, volume, hold time and transport have no universal values: they must be defined, validated and kept stable, because a silent change in technique reads in the trend as a system event.
  • The microbiological result is retrospective: control is preventive and sampling verifies that it is holding. When data drift, system and plan are revised, not the limits — the EMA Q&A calls raising limits a practice that may mask a failing system.

Regulatory and technical references

  • EudraLex Volume 4, Annex 1 (C(2022) 5938 final), in operation since 25 August 2023 — clauses 2.5(v), 6.8, 6.13, 6.14, 6.15. health.ec.europa.eu
  • EMA/INS/GMP/443117/2017 Q&A Production of WFI by non-distillation methods – reverse osmosis, biofilms and control strategies, in force since 1 August 2017; EMA/CHMP/CVMP/QWP/496873/2018 Guideline on the quality of water for pharmaceutical use, in force since 1 February 2021. ema.europa.eu
  • WHO TRS 1033, Annex 3 (2021) Good manufacturing practices: water for pharmaceutical use — §4.3.8 and §12.8. who.int
  • USP <1231> Water for Pharmaceutical Purposes (official since 1 December 2021), informational chapter; USP <643> Total Organic Carbon; USP <645> Water Conductivity.
  • ICH Q9(R1) Quality Risk Management (Step 4, 18 January 2023); ICH Q10 Pharmaceutical Quality System. ich.org
  • PIC/S PI 009-4 Inspection of Utilities, rev. 4, in force since 1 January 2021. picscheme.org
  • ISPE Good Practice Guide Sampling for Pharmaceutical Water, Steam, and Process Gases; ISPE Baseline Guide Vol. 4 Water and Steam Systems, 3rd ed. (2019) — industry guides, not regulatory texts.

THE PRAGMATIC GMP · EVERY MONDAY

The GMP topics that matter, in 7 minutes.

One GMP topic, one real-world example and one practical action, based on official sources and inspection trends.
Discover The Pragmatic GMP →