Skip to main content
ANCVEIRS
Professional workAnalysisCybersecurity

Security Research in Critical Systems: Open Reporting Does Not Mean Open Testing

How to preserve the value of independent security research without turning high-consequence production systems into uncontrolled testing environments.
Zigmārs AncveirsTechnology Leader in FinTech & RegTech · Cybersecurity & Ethical HackingPublished: 26 September 2026Reviewed: 26 September 202614 min

Introduction

A researcher notices a suspicious endpoint on a publicly reachable hospital, bank, energy or transport system.

What should happen next?

On an ordinary web application, one additional controlled request may create little operational risk.

In another environment, technically similar experimentation could:

  • alter clinical data;
  • interfere with a payment flow;
  • block authentication;
  • trigger control automation;
  • overload equipment;
  • affect third parties;
  • cause cascading disruption beyond the service being tested.

“Allow security research” is therefore not a binary policy choice.

The real question is:

Which research actions can this particular system tolerate safely, what evidence is actually necessary, and how much control should surround the test?

A second distinction is equally important:

An open vulnerability-reporting channel does not require an organisation to make every production system openly testable.

CVD can remain public.

Active-testing authority can be differentiated.

That distinction matters most where the cost of a testing mistake is high.

“Critical system” is not one legal status

European law contains several adjacent but different classifications.

NIS2 distinguishes essential and important entities, while Annex I lists sectors of high criticality, including energy, transport, banking, financial-market infrastructure, health, drinking water, waste water, digital infrastructure, managed ICT services, space and public administration.1

The Critical Entities Resilience Directive (CER) requires Member States to identify critical entities that provide essential services necessary for vital societal functions, economic activities, public health and safety or the environment.2

DORA creates a separate digital-operational-resilience framework for financial entities and an advanced threat-led penetration-testing regime for selected entities.3

These categories overlap.

They are not interchangeable.

A technically critical system can also exist inside an entity that is not labelled a “critical entity” for the particular legal question being discussed.

A small identity, scheduling or control component may be operationally critical because many other services depend on it.

In this article, critical system therefore describes potential consequence and operational sensitivity, not a substitute legal classification.

Criticality changes the risk of the test itself

Security teams usually think about vulnerability risk.

Critical systems require a second analysis:

testing risk.

Those are different questions.

A potentially severe race condition may be unsafe to demonstrate in production.

A read-only access-control flaw may sometimes be provable with one tightly constrained request.

The appropriate research mode therefore cannot be selected only from CVSS or assumed vulnerability severity.

Relevant testing-risk dimensions include:

  • availability sensitivity;
  • integrity sensitivity;
  • physical or public-safety impact;
  • confidentiality of the data;
  • reversibility;
  • blast radius;
  • third-party exposure;
  • cascading dependencies;
  • recoverability;
  • observability;
  • sectoral or supervisory requirements.
vulnerability risk:
what could happen if a real adversary exploits the weakness?

testing risk:
what could happen merely because we attempt to prove it?

Public CVD can coexist with narrow public test scope

This model already exists in practice.

Latvijas Banka's public vulnerability-disclosure policy permits research against specifically named public websites — bank.lv, e-monetas.lv and naudasskola.lv — while explicitly stating that testing of its other websites, resources and systems is not permitted under that initiative.7

The policy also prohibits:

  • social engineering;
  • accessing more information than the strict minimum needed to prove the vulnerability;
  • deleting or modifying information;
  • DoS/DDoS;
  • automated password guessing.7

Testing must stop and the organisation must be contacted if the researcher encounters personally identifying, financial, trade-secret or ownership information.7

That architecture is worth noticing:

There is no contradiction.

It is risk differentiation.

public vulnerability intake
+
limited public test scope
+
sensitive/core systems outside the open-testing scope

CERT.LV reinforces minimum-impact testing

CERT.LV's August 2026 guidance states that ethical vulnerability testing does not include DoS/DDoS unless a specific programme has explicitly authorised it and that researchers should use the minimum number of requests and actions needed to demonstrate the vulnerability.6

The platform terms require testing methods and intensity to be proportionate to the capacity of the tested resource and require researchers to avoid actions that harm the resource or its owner.5

That principle applies broadly.

In critical systems it becomes foundational:

The purpose of a PoC is to establish sufficient evidence, not to demonstrate the maximum damage theoretically available.

Open research is not the only research model

A public VDP or CVD policy is valuable.

But public CVD is not equivalent to:

anyone may actively test everything.

Higher-consequence systems can use several access profiles.

I would not treat these as maturity levels.

They are different risk modes.

1. Report-only / passive discovery

Anybody can report:

  • accidentally observed weaknesses;
  • exposed configuration;
  • leaked credentials;
  • public data exposure;
  • externally visible security defects.

Active exploitation is not authorised.

This can be appropriate where even limited active validation creates unacceptable operational risk.

2. Public low-impact testing

Defined public assets and methods are authorised, usually with constraints such as:

  • web/API scope;
  • minimum sufficient PoC;
  • no DoS;
  • no data modification;
  • no persistence;
  • no social engineering.

This is a conventional VDP/CVD profile.

3. Registered research

Before broader testing, the researcher:

  • registers;
  • accepts additional terms;
  • provides a reliable contact;
  • receives a test identifier;
  • uses assigned test accounts or data;
  • accepts rate limits or a test window.

The organisation gains operational control without turning the engagement into a conventional penetration test.

4. Vetted / invitation-only research

More sensitive scope is available only to researchers whose:

  • identity has been verified;
  • experience or competence is known;
  • prior quality is established;
  • evidence-handling obligations are explicit.

This can take the form of private bounty or another controlled research programme.

5. Controlled professional testing

Higher-risk methods use:

  • an approved test plan;
  • a defined test window;
  • test accounts or synthetic data;
  • active monitoring;
  • a control function;
  • emergency stop;
  • explicit escalation;
  • rollback/recovery readiness;
  • cleanup and retest.

DORA TLPT is a particularly formalised financial-sector example of this logic, not a universal definition of controlled testing.34

One organisation can use several profiles at the same time

Do not classify an entire enterprise with one testing mode.

A practical model can look like:

This is more precise than either extreme:

critical infrastructure is too sensitive for bug bounty;

or:

if a service is internet-facing, researchers should be free to test all of it.

Both are overbroad.

public information website
→ public low-impact CVD

customer self-service portal
→ registered/private bounty

administrative back office
→ invitation-only testing

core transaction / control system
→ controlled professional test

unexpected vulnerability anywhere
→ reporting route remains available

Test windows are a safety control

In a high-consequence production environment, when testing happens can matter almost as much as how it happens.

Testing during:

  • a low-load period;
  • settlement or payroll;
  • peak bookings;
  • a major infrastructure migration;
  • an active incident

are operationally different propositions.

A controlled test window lets the organisation:

  • staff monitoring appropriately;
  • observe service health;
  • distinguish exercise telemetry from attack telemetry;
  • stop quickly;
  • restore state.

This does not mean the entire blue team must know every red-team scenario.

Secrecy may be necessary to test detection and response.

But there should still be a control function that knows the exercise exists and has authority to stop it.

DORA TLPT shows how seriously test-induced risk can be treated

Commission Delegated Regulation (EU) 2025/1190 requires the TLPT control team to assess risks associated with testing live production systems, including potential effects on the financial entity, third parties and the broader financial sector.4

The regulation explicitly addresses risks relating to:

  • sensitive information provided to testers;
  • crisis and incident escalation;
  • interruption of critical activities;
  • corruption of data;
  • impact on third parties;
  • incomplete restoration after the test.4

Under exceptional risk conditions, active testing can be suspended and, in defined circumstances, continued only through a limited purple-team exercise.4

The generalisable lesson is:

High-realism security testing needs safety engineering of its own.

Emergency stop must be operational, not merely contractual

Rules of engagement may say:

Stop immediately if instructed.

Necessary, but incomplete.

A controlled programme should define:

  • who may issue the stop command;
  • which communication path is authoritative;
  • how quickly the tester is expected to see it;
  • which activities terminate immediately;
  • whether sessions or tokens must be revoked;
  • how persistence is removed;
  • how modified test data is restored;
  • who verifies cleanup.

For a high-risk test, emergency stop is an operational protocol.

Not a sentence buried in legal terms.

Identity verification is not proof of quality

Vetted programmes may use:

  • identity verification;
  • previous experience;
  • platform reputation;
  • references;
  • NDA;
  • specialist skill requirements.

Those controls can reduce certain risks.

They do not imply:

Vetting helps decide whom to trust with more sensitive testing authority.

Evidence still determines whether a vulnerability is real.

Behaviour still determines whether the researcher operates safely.

verified identity
= competent researcher

high reputation
= safe conduct in every environment

known researcher
= technically correct finding

Sandboxes and digital twins reduce risk but introduce fidelity gaps

High-risk functions can often be tested in:

  • staging;
  • dedicated test environments;
  • synthetic datasets;
  • hardware-in-the-loop labs;
  • digital twins;
  • isolated tenants.

That is often the right first step.

But non-production environments can fail to reproduce:

  • routing;
  • IAM;
  • third-party integrations;
  • production load;
  • caching;
  • edge controls;
  • timing;
  • firmware combinations;
  • operational workarounds.

So:

A mature programme does not frame the choice as “staging or production”.

It asks which hypothesis can be safely tested in which environment.

safer environment
→ lower test risk

but

lower fidelity
→ production-only weaknesses may be missed

Use synthetic data whenever real people are not needed to prove the point

If a weakness can be demonstrated with:

  • two test accounts;
  • a synthetic patient;
  • a fake payment;
  • a test booking;
  • a dummy identifier;

there is rarely a professional reason to use real-person data simply because it is available.

Evidence minimisation reduces:

  • privacy risk;
  • incident risk;
  • regulatory risk;
  • the consequence of a testing mistake.

When sensitive production data appears unexpectedly, a clear stop rule matters.

Latvijas Banka's public VDP provides a concrete example by requiring testing to stop when personally identifying, financial or trade-secret information is encountered.7

Third-party boundaries are especially important in critical services

Critical services are supply chains.

A healthcare service may depend on:

  • medical-device vendors;
  • laboratories;
  • identity services;
  • government registers;
  • cloud;
  • telecommunications.

Energy may depend on:

  • OT vendors;
  • telecom;
  • remote maintenance;
  • SCADA integrators;
  • cloud analytics.

Finance may depend on:

  • cloud;
  • payment schemes;
  • market infrastructure;
  • fintech integrations;
  • outsourcing providers.

One organisation's permission does not automatically authorise testing across every technical dependency.

Scope therefore needs more than:

It also needs an:

Who actually has the right to authorise testing of each technical layer?

IP / domain list
authority map

Cascading effects change what counts as a proportionate PoC

The CER Directive emphasises cross-sector interdependencies and the possibility that disruption of one essential service can create wider cascading effects.2

That matters for vulnerability research.

A PoC that appears locally to do nothing more than:

restart one service

can generate:

  • failover;
  • queue backlog;
  • downstream timeout;
  • partner retry storms;
  • synchronisation errors;
  • manual-recovery work.

In a critical environment, testing proportionality cannot be assessed from the immediate reaction of one component alone.

The system context matters.

Disclosure may need differentiated control as well

Responsible disclosure does not mean permanent secrecy.

Critical-system publication may, however, require coordination with:

  • vendors;
  • downstream operators;
  • CSIRTs;
  • sector regulators;
  • authorities in several countries;
  • patch deployment windows;
  • compensating controls.

NIS2's CVD architecture explicitly supports CSIRT coordination and cross-border cooperation for multi-party cases.1

Therefore:

permission to test publicly and timing of public disclosure

are separate policy variables.

This is not an argument for security through obscurity

Critical operators may have sound reasons to:

  • avoid publishing full architecture;
  • limit testing methods;
  • require researcher identity;
  • coordinate disclosure;
  • keep operational details confidential.

That is not automatically security through obscurity.

Obscurity becomes a problem when the organisation assumes:

If outsiders cannot inspect it, the vulnerability effectively does not exist.

Controlled research takes the opposite position.

It accepts the value of independent adversarial scrutiny while constraining the testing authority in proportion to possible harm.

Out of scope for testing should not mean out of scope for reporting

One principle deserves explicit treatment:

A system that is out of scope for active testing should not automatically be out of scope for reporting.

A researcher may:

  • notice a weakness accidentally;
  • find public data exposure;
  • discover a leaked credential;
  • receive evidence without exploiting anything.

The organisation still needs a safe route to receive that information.

This is why CVD intake should be separated from active-testing permission.

A risk-based research profile

Before granting active-testing authority, I would assess at least these dimensions:

The more the right-hand column describes the system, the weaker the case for uncontrolled public active testing.

A risk-based research profile
QuestionLower consequenceHigher consequence
Availabilitylocal, easily recoverable failuredisruption affects an essential function
Data integritydisposable/syntheticclinical, financial or control data
Confidentialitypublic/test datahealth, financial, state or other sensitive data
Physical impactnonepossible influence on physical processes
Blast radiusone test tenantmultiple operators, customers or sectors
Reversibilityimmediatedifficult or uncertain
Third partiesnoneshared/cloud/supply-chain dependencies
Observabilitycomplete monitoringtesting hard to distinguish from a real incident
Recoverytestedrecovery not confidently verified

A minimum controlled-research record

For a sensitive programme, I would preserve something like:

This is not a regulatory standard.

It is the author's practical minimum for controlled external research.

researcher:
  identity_verified:
  competence_basis:
  contact:

authorization:
  assets:
  methods_allowed:
  methods_prohibited:
  third_party_boundaries:

safety:
  test_window:
  rate_limits:
  test_accounts:
  synthetic_data:
  monitoring_owner:
  emergency_stop_contact:
  rollback_or_cleanup:

evidence:
  storage:
  encryption:
  retention:
  personal_data_stop_rule:

coordination:
  incident_escalation:
  disclosure_route:
  regulator_or_csirt_contact_if_needed:

closure:
  cleanup_verified:
  finding_owner:
  retest:

Controls should not kill spontaneous research

Risk controls can become excessive.

If every low-impact web request requires:

  • identity verification;
  • NDA;
  • prior approval;
  • a narrow testing window;
  • weeks of registration;

many capable external researchers simply will not participate.

The objective is not to turn all research into contracted penetration testing.

It is:

keep friction low for lower-consequence research, and add stronger controls only where they materially reduce plausible harm.

That is proportionality.

Conclusion

Critical infrastructure is not an ordinary website.

That does not mean independent security research has no place in it.

A better conclusion is:

As potential consequence rises, permission, methods, environment, researcher access and emergency controls need to become more precise.

A public vulnerability-reporting route can exist even in a highly sensitive organisation.

Only a small part of its infrastructure may be publicly testable.

Deeper access can be reserved for registered or vetted researchers.

The highest-risk activities may require professionally controlled testing with monitoring, test windows, synthetic data, stop protocols and recovery preparation.

The real choice is not:

It is:

In that model, the researcher is neither an uncontrolled threat nor an outsider who must be kept away from every sensitive system.

The researcher becomes one assurance layer in a system where testing freedom and testing safety are designed together.

open research
OR
no research
open reporting
+
risk-proportionate testing authority
+
stronger controls as potential harm increases

Frequently asked questions

Should critical infrastructure have a public CVD channel?

Yes, a public reporting route can be valuable even for high-consequence organisations. It does not mean every technical system is automatically open to active testing.

Does public CVD authorise scanning of the whole organisation?

No. Researchers must follow the programme's asset and method scope. Latvijas Banka's public policy, for example, names three public websites that may be tested and explicitly excludes other bank systems from that initiative.7

Do vetted researchers solve testing risk?

No. Vetting can reduce some identity and trust risks, but does not replace scope, monitoring, minimum-impact rules, evidence controls or emergency-stop procedures.

Should critical systems be tested only in staging?

Not always. Staging lowers operational risk but may miss production-only configuration, timing or integration weaknesses. Higher-fidelity testing may require controlled production activity with stronger safeguards.

Is DORA TLPT the model for every critical sector?

No. TLPT is a specific advanced-testing regime under DORA for selected financial entities. Its risk-management concepts are informative, but are not automatically binding models for healthcare, energy or transport.

What if a vulnerability is accidentally noticed in a system that is out of scope for testing?

Do not increase impact merely to create a stronger PoC. Preserve the minimum evidence and use the organisation's CVD/CSIRT route. “Not authorised for active testing” is not the same thing as “do not report”.

Source status

EU and Latvian source status checked on 25 September 2026. The five external-research profiles, risk matrix and controlled-research record in this article are the author's analytical model, not an official taxonomy of NIS2, CER, DORA, CERT.LV or Latvijas Banka.

This article analyses external security-research and testing-governance models. The legal status of a particular system and the acts authorised against it must be assessed under the rules applicable to that system and programme.

Sources

  1. Directive (EU) 2022/2555 (NIS2), particularly Articles 2–3 and 12 and Annexes I–II · EUR-Lex
  2. Directive (EU) 2022/2557 on the resilience of critical entities (CER), particularly Articles 1–2 and the Annex · EUR-Lex
  3. Regulation (EU) 2022/2554 (DORA), particularly Articles 24–27 on digital operational resilience testing and TLPT · EUR-Lex
  4. Commission Delegated Regulation (EU) 2025/1190 on TLPT regulatory technical standards, particularly Article 5 on risk management and provisions for suspension of active testing · EUR-Lex
  5. CERT.LV, Vulnerability Reporting Platform Terms of Use, effective 1 August 2026 · CERT.LV
  6. CERT.LV, “Atgādinām: ētiska ievainojamību testēšana neietver DoS un DDoS uzbrukumus”, 28 August 2026 · cert.lv
  7. Latvijas Banka, Vulnerability disclosure policy, checked 25 September 2026 · bank.lv