The infrastructure of an Environmental Monitoring System is the part of the project nobody looks at until it stops. Network, power, uninterruptible supplies, servers, time synchronisation, backup: elements that do not appear in functional specifications, that often fall to functions other than those writing the URS, and that determine whether the system will be available when it is needed and whether the data it produces will be defensible.
One rule governs every decision in this phase: every external dependency of the system must have an owner, a defined behaviour on failure, and documented verification that the behaviour is the one expected. A system that relies on a site network with no formal agreement on availability and maintenance, on a power supply with no defined response to an interruption, or on a backup never tested by restore, has not managed those risks: it has inherited them.
This is also the phase in which a distinction that causes recurring confusion must be kept sharp: commissioning and qualification are not the same activity. Commissioning is good engineering practice and serves to bring the installation to a state where it works as intended; qualification is the documented activity by which, under approved protocols, the system is demonstrated to be fit for its intended GMP use. Good commissioning reduces the burden of qualification but does not replace it.
Why this phase determines system availability
An EMS that is unavailable during a critical operation confronts the company with an immediate and unwelcome question: stop, continue with an alternative measurement, or continue and document. If the answer has not been defined and tested beforehand, it will be improvised under pressure — which is precisely the situation that produces the deviations hardest to close.
Availability is not a property of the instrument: it is a property of the chain. Probe, instrument, power, network, server, storage, support services: overall availability is governed by the weakest link and by how quickly it can be restored. Designing the infrastructure means deciding deliberately where to accept a single point of failure and where not to — and writing that down.
The frame: what is required and what is a design choice
| Level | What it establishes regarding infrastructure |
|---|---|
| Regulatory requirement (Annex 11, January 2011 revision) | For computerised systems, sets expectations on validation, access security, data integrity and protection, backup, incident management, business continuity and formal agreements with service providers. It does not prescribe specific technical solutions. |
| Regulatory requirement (EudraLex Volume 4, Annex 1) | Requires monitoring appropriate to criticality and, where applicable, continuous during critical operations: a requirement that translates into availability requirements for the supporting infrastructure. |
| Regulatory requirement (EudraLex Volume 4, Annex 15) | Sets out the qualification and validation framework, including the relationship between design-phase verification activities and qualification activities. |
| Good engineering practice | Commissioning, route separation, redundancy of critical components, orderly and identified installations, handover documentation. Reduces the qualification burden but does not replace it. |
| Company IT policies | Network segregation, patch management, cybersecurity, remote access management. To be reconciled with GMP requirements, not applied automatically. |
| GuideGxP operational recommendation | Put in writing, before installation, an agreement between quality, engineering and IT establishing owner, service level and expected behaviour for each infrastructure dependency. |
Technical guidance
Network and segregation
The network carrying monitoring data is part of the system. What to define:
- Boundary of the computerised system: where the EMS ends and site infrastructure begins. From this definition follow the validation perimeter, responsibilities and change management.
- Segregation: whether and how the EMS network is separated from the general corporate network, with what traffic rules and what controls. Separation reduces exposure but introduces dedicated management needs.
- Remote access: if the supplier needs access for support, the arrangements must be defined, authorised, logged and time-limited, with responsibilities set contractually.
- Physical routing: cables crossing classified areas follow the same coordination and sealing rules described for sampling lines in the article on probe installation.
- Behaviour when the network is unavailable: what happens to data acquired in the field while the connection is down, and how it is reconciled on restoration. To be defined, tested and documented.
Power supply and continuity
- Classification of loads: which components must stay powered during an interruption and for how long, according to what the system must be able to demonstrate.
- Uninterruptible power supplies: sizing and autonomy must be defined from the chosen scenario, not from a generic convention; the autonomy required depends on how long it takes to complete or safely interrupt the operation in progress.
- Shutdown and restart behaviour: how the system behaves when power is lost and when it returns; whether partial data is retained; whether restart is automatic or requires intervention; whether alarm states are correctly restored. These are checks to be performed, not assumed.
- Maintenance of uninterruptible supplies: batteries and components have a service life; the maintenance plan must be defined at handover, not at the first outage.
Time synchronisation
This is one of the infrastructure elements with the greatest impact on data integrity, and one of the most frequently neglected. If instruments, servers and connected systems do not share a reliable time reference, correlation between events becomes unreliable and investigations lose their basis. It must be defined what the reference source is, how the various components align to it, at what frequency, how daylight-saving changes are handled, who can change the system clock and how such a change is traced. In fleets of standalone instruments, where each unit has its own clock, this becomes an operating procedure in its own right.
Servers, storage and restore
- Hosting and responsibility: physical server, virtual machine or managed service imply different responsibility models, which must be made explicit in writing.
- Backup: scope, frequency, retention and location of copies must be defined consistently with applicable data retention requirements.
- Restore: an untested backup is not a backup. The restore test must be performed, documented and repeated periodically.
- Long-term archiving: how data remains readable and reconstructable for the entire required retention period, including after the system is decommissioned.
- Patch management: operating system, database and application have their own update cycles; the process for impact assessment and any re-verification must be agreed with IT before go-live.
Cybersecurity
Information security measures must be reconciled with GMP requirements: an automatic update policy applied without impact assessment can modify a qualified system; an over-aggressive session lock rule can interrupt surveillance. The answer is not to exempt the EMS from security policy, but to define how the two needs are combined, with explicit responsibilities and decision process. This intersects with access and audit trail management, covered in the article on EMS software between Annex 11, Part 11 and data integrity.
Commissioning and qualification: where the line falls
| Aspect | Commissioning | Qualification |
|---|---|---|
| Nature | Good engineering practice | Documented GMP activity |
| Objective | Bring the installation to work as intended | Demonstrate fitness for intended use |
| Typical ownership | Engineering and supplier | Quality, with engineering and the user |
| Documentation | Records and technical reports | Approved protocols and reports |
| Criteria | Technical specifications | Approved acceptance criteria |
| Relationship | Reduces the qualification burden if planned and documented | Cannot be replaced by commissioning |
For commissioning to lighten qualification, it must be planned consistently with the validation strategy, executed by competent personnel and documented verifiably. Undocumented commissioning produces no benefit at qualification stage: it simply has to be done again.
Working tool: infrastructure dependency matrix
To be completed before installation and filed with the project documentation. The columns are deliberately empty: fill them with your own site's data.
| Dependency | Owner | Expected behaviour on failure | Planned verification |
|---|---|---|---|
| Data transport network | |||
| Normal electrical supply | |||
| Uninterruptible power supply | |||
| Server / virtual machine | |||
| Storage and backup | |||
| Time synchronisation source | |||
| Supplier remote access | |||
| Support and on-call services | |||
| Integrations with other systems |
A practical scenario
At a site we will call Site Delta — realistic but fictional — the EMS is installed and qualified without findings. Some months after go-live, during processing, the system stops recording for a period of time. The investigation reconstructs the sequence: planned maintenance on the network infrastructure, communicated to the IT department but not to the function operating the system; no formal agreement identifying the EMS as a system with particular availability requirements; no fallback procedure defined for that scenario.
None of the three is a technical problem. They are three governance gaps, all foreseeable at design stage and all resolvable with a dependency matrix completed and approved before installation.
A second element emerges at the same site during review: the data restore test had never been performed after go-live. The backup was running; restore in the real configuration had never been verified. It is a gap typically discovered at the worst possible moment.
Common mistakes and red flags
- Treating network and power as “site services” outside the project. If the system depends on them, they are part of the project and must be governed as such.
- Not defining system behaviour on interruption. The question “what do we do if it stops during processing” must be answered at the desk, not on the shop floor.
- Sizing continuity by convention. The autonomy required follows from the operating scenario, not from inherited practice.
- Neglecting time synchronisation. It is the infrastructure gap that most often undermines the reconstructability of events.
- Considering backup sufficient without testing restore. Only a verified restore demonstrates that data is recoverable.
- Applying IT policies without impact assessment. Updates and configuration changes on a qualified system go through assessment, not automation.
- Confusing commissioning with qualification. Presenting commissioning activity as qualification is a classic finding; failing to document commissioning forfeits the benefit it could have produced.
- Not formalising service agreements. Without written responsibilities, every outage becomes a discussion instead of a procedure.
- Involving IT late. Brought in late, IT receives constraints instead of contributing to decisions, and the result is almost always worse for both sides.
How to document
- System boundary definition: what is part of the EMS and what is site infrastructure, with implications for the validation perimeter.
- Dependency matrix: owner, expected behaviour on failure, planned verification for each dependency.
- Service agreements: agreed levels, response times, responsibilities for maintenance, updates and restore.
- Commissioning plan: activities, criteria, responsibilities and relationship with the qualification strategy.
- Fallback procedures: defined and tested behaviour for system unavailability during operations.
- Evidence of continuity testing: behaviour on power loss, restart, data restore.
- Time synchronisation policy: source, alignment method and frequency, change management.
- Handover documentation: configuration, administration credentials, maintenance plan for infrastructure components.
Key takeaways
- Every external dependency must have an owner, a defined behaviour on failure and a documented verification.
- Availability is a property of the chain, not of the instrument.
- Time synchronisation is a data integrity requirement, not a configuration detail.
- A backup is worth exactly as much as its verified restore.
- Commissioning and qualification are different activities: the first can lighten the second only if planned and documented.
- IT should be involved as a co-designer, not as an executor downstream of decisions.
Frequently asked questions
Does an EMS need a dedicated network?
There is no obligation to that effect. Segregation reduces exposure and simplifies change control, but introduces dedicated management needs. The choice must be justified on the basis of risk, system criticality and the organisation's ability to sustain it over time.
How much autonomy should the uninterruptible supply provide?
Enough to complete or safely interrupt the operation in progress, and to keep available the functions that must remain active under the defined strategy. There is no standard figure: it derives from the site's operating scenario.
Can commissioning replace qualification?
No. It can reduce the burden if planned consistently with the validation strategy, executed by competent personnel and documented verifiably, but it does not replace the documented activity demonstrating fitness for intended use. The distinction is developed in the article on FAT, SAT, IQ, OQ and PQ of the system.
Who is responsible for the network: IT or quality?
Technical responsibility typically sits with IT; responsibility for adequacy against GMP requirements remains with quality. For the model to work, a written agreement is needed setting service levels, how maintenance is communicated, and the process for assessing changes.
How often should data restore be tested?
At a frequency defined by the company on the basis of system criticality and its own data management policy, and in any case after significant infrastructure changes. The test must be documented: it is the only evidence that restore works in the real configuration.
How are operating system and application updates managed?
Through change control, with impact assessment on the qualified system and definition of the verifications required before release to operation. The process must be agreed with IT before go-live, because updates will arrive regardless.
Regulatory and technical references
- EudraLex Volume 4 — EU Guidelines for Good Manufacturing Practice (European Commission): Annex 1 (applicable from 25 August 2024), Annex 11 (January 2011 revision), Annex 15 (in operation from 1 October 2015).
- ICH Quality Guidelines — ICH Q9(R1) Quality Risk Management.
- PIC/S — Guides and Guidance Documents.
Continue the project journey
This article is part of the GuideGxP Environmental Monitoring Systems pathway, which follows the life cycle of an EMS project from requirements definition through to operational management.
- Upstream: the choice of architecture and the installation of probes and lines.
- Downstream: software, Annex 11 and data integrity and FAT, SAT, IQ, OQ and PQ.
- In operation: calibration, maintenance and periodic review.
- GuideGxP regulatory foundations: data integrity, ALCOA+ and audit readiness.
Want analysis like this straight to your inbox? Subscribe to The Pragmatic GMP, the GuideGxP newsletter for professionals working daily with GMP, qualification and data integrity.