Skip to main content
ANCVEIRS
Professional workGuideCybersecurity

Security Assurance Under Pressure: Pentest, Red Team, CVD, Bug Bounty and TLPT

What each security-testing mode actually proves, what it does not prove, and how to combine planned, continuous and threat-led evidence without false confidence.
Zigmārs AncveirsTechnology Leader in FinTech & RegTech · Cybersecurity & Ethical HackingPublished: 26 September 2026Reviewed: 26 September 202615 min

Introduction

An organisation can pass a penetration test and suffer a serious incident the following month.

It can run a mature bug-bounty programme for years and still not know whether its SOC would detect a targeted intrusion.

A red team can reach its objective even when the product is not full of trivial vulnerabilities.

An empty vulnerability-disclosure inbox does not demonstrate the absence of vulnerabilities.

These are not contradictions.

They become confusing only when different security-testing modes are treated as interchangeable answers to the same question.

They are not stronger and weaker versions of one instrument.

They produce different kinds of evidence about different security properties.

The mature question is therefore not:

Which security test is best?

It is:

Which security claim are we trying to support, and which testing method can produce evidence relevant to that claim?

That is the core of security assurance.

Assurance is not the name of a test

NIST SP 800-115 has long treated technical security assessment as a collection of techniques with different benefits and limitations rather than one universal method.1

Modern organisations have an even broader set of evidence-producing activities:

  • vulnerability scanning;
  • source-code review;
  • penetration testing;
  • red-team exercises;
  • purple-team collaboration;
  • CVD;
  • bug bounty;
  • incident simulations;
  • resilience testing;
  • and, in the financial sector, DORA TLPT.

None of these is assurance by itself.

Assurance emerges when an organisation can connect:

Without that chain, testing can become evidence that an activity occurred rather than evidence that a security control is effective.

security claim
→ evidence-producing activity
→ finding / observation
→ owner
→ remediation
→ verification
→ residual uncertainty

Start with the question, not the tool

Different testing modes answer different primary questions.

This is not a maturity ladder.

A red team is not “a better pentest”.

Bug bounty is not “outsourced penetration testing”.

TLPT is not simply “an extremely deep red team”.

Start with the question, not the tool
ModePrimary question
Vulnerability scanDo we have known, technically detectable exposures?
Penetration testCan selected security boundaries be practically broken within a defined scope and period?
Red teamCan a realistic adversary achieve a defined objective, and would we detect or stop it?
Purple teamWhat can offensive and defensive teams learn and improve together?
CVDWhat can independent outsiders discover outside our planned testing schedule?
Bug bountyHow can we deliberately attract and incentivise external researcher attention?
TLPTFor selected critical functions, what does a threat-intelligence-led, controlled live-production exercise demonstrate about attack paths, detection, response and remediation?

Vulnerability scanning gives breadth, not completeness

Automated scanning can be excellent at identifying:

  • known CVEs;
  • exposed versions;
  • open services;
  • configuration weaknesses;
  • TLS and header issues;
  • repeatable technical patterns.

Its strength is scale.

A scanner can repeatedly check thousands of assets.

Its weakness is the class of question that depends on context:

  • business logic;
  • complex authorisation;
  • cross-system attack chains;
  • true exploitability in a particular configuration;
  • human and process controls.

So:

Scanning is a high-scale exposure signal, not a complete adversarial proof.

scanner clean
≠
system secure

Penetration testing is bounded by scope and time

A penetration test is normally a pre-authorised, time-bounded engagement with:

  • defined scope;
  • rules of engagement;
  • allowed methods;
  • start and end dates;
  • deliverables;
  • remediation and often retest.

NIST SP 800-115 treats penetration testing as one technical security-testing method used to identify vulnerabilities and assess the practical consequences of weaknesses.1

A strong pentest can examine:

  • application and API authorisation;
  • privilege escalation;
  • network segmentation;
  • credential abuse;
  • business logic;
  • chains of technical findings.

Its scope and time-box are simultaneously its limitation.

If the engagement is:

the report provides evidence about that tested boundary.

It does not prove security of:

  • the entire organisation;
  • code released the next day;
  • excluded cloud assets;
  • an unknown supplier dependency;
  • a patient six-month intrusion campaign.
5 days
3 applications
no social engineering
no production DoS

Red teaming asks whether the adversary can achieve an objective

The unit of a red-team engagement is usually not the number of vulnerabilities.

It is an objective.

For example:

obtain controlled access to a defined critical information set;

or:

reach a privileged position without crossing prohibited safety boundaries.

A red team may combine:

  • technical exploitation;
  • identity;
  • cloud;
  • endpoints;
  • physical access;
  • authorised social engineering;
  • simulated persistence;
  • detection avoidance.

The most useful result is often not:

We found twelve vulnerabilities.

It is:

  • which attack paths worked;
  • where defenders detected activity;
  • where they did not;
  • what prevented progression;
  • how response teams behaved;
  • which telemetry or process gaps mattered.

A red team can therefore produce high assurance value even when its report contains relatively few conventional “bugs”.

Red team is not exhaustive vulnerability assessment

This distinction matters.

Once a red team finds a viable path to its objective, it may have little reason to enumerate another twenty unrelated vulnerabilities elsewhere.

Its optimisation target is the attack objective.

A penetration test more often optimises for identifying weaknesses inside a defined scope.

So:

A red team can fail because of the selected scenario, time box, threat intelligence, rules of engagement or genuinely effective controls.

All of those outcomes can be useful evidence if the claim boundary is clear.

red team success
≠ pentest failure

red team failure
≠ organisation secure

Purple teaming is a learning mode, not just a colour

A purple team is not necessarily a permanent team.

In practice, purple teaming often means collaborative work between offensive/red and defensive/blue functions.

The goal may be to:

  • replay an attack technique;
  • inspect telemetry;
  • build a detection;
  • tune alerts;
  • validate an incident-response playbook;
  • understand why an activity was missed.

Secrecy is no longer the primary value.

The value is the feedback loop.

DORA's TLPT regime now provides a precise regulated example. Commission Delegated Regulation (EU) 2025/1190 defines purple teaming as collaborative testing involving testers and the blue team, and requires replay and a purple-teaming exercise during the TLPT closure phase.4

That does not make DORA the general definition of purple teaming for every sector.

It shows how one regulated assurance regime has formalised the learning step.

CVD produces unplanned external evidence

The assurance value of CVD is almost the inverse of penetration testing.

In a pentest, the organisation chooses:

  • when;
  • where;
  • who;
  • for how long;
  • using which methods.

A CVD report can arrive:

  • overnight;
  • from an unknown researcher;
  • against an asset internal teams considered uninteresting;
  • in a vulnerability class nobody planned to test.

That unpredictability is part of its value.

FIRST's PSIRT Services Framework treats external finders, vulnerability reporting and internal product-security assessments as complementary sources of vulnerability discovery.2

But CVD does not provide a coverage guarantee.

An empty inbox does not mean:

no vulnerability exists.

CVD is an opportunistic external discovery channel.

Bug bounty adds an attention market

Bug bounty adds an economic control to external research.

A programme can redirect researcher attention through:

  • higher rewards for a target;
  • new-scope bonuses;
  • vulnerability-class campaigns;
  • private invitations;
  • time-limited incentives.

That makes bounty valuable as an assurance input.

It still does not provide systematic test coverage.

Researchers choose where to invest time.

Coverage emerges from incentives, target attractiveness, expertise and competition.

For programme design, see Bug Bounty & Crowdsourced Security: What a Good Programme Actually Optimises.

TLPT is a regulated assurance regime in finance

DORA is a useful example of why mature assurance cannot collapse into one test.

Article 24 requires covered financial entities, subject to its proportionality and scope rules, to establish a sound and comprehensive digital operational resilience testing programme as part of ICT risk management.3

Article 25 lists a broad set of possible methods, including:

  • vulnerability assessments and scans;
  • network security assessments;
  • source-code review where feasible;
  • scenario-based tests;
  • end-to-end testing;
  • penetration testing.3

Article 24 also requires findings from testing to be prioritised, classified and remediated, together with internal validation that identified weaknesses, deficiencies and gaps have been fully addressed.3

TLPT then adds a distinct advanced-testing regime for designated financial entities.

Under Article 26, those entities perform TLPT at least every three years unless the competent authority adjusts frequency based on risk profile and operational circumstances. The test covers several or all critical or important functions and is conducted on live production systems supporting them.3

This is significantly more specific than ordinary penetration testing.

DORA TLPT is not simply “compliance red teaming”

Commission Delegated Regulation (EU) 2025/1190 specifies:

  • identification criteria;
  • scope specification;
  • threat-intelligence phase;
  • red-team test planning;
  • active testing;
  • risk management;
  • closure;
  • purple teaming;
  • remediation planning;
  • supervisory cooperation.4

Because live-production realism creates operational risk, the RTS also provides explicit controls for suspending testing when there is exceptional risk to data, assets, critical functions or the wider financial sector, with limited purple teaming available in defined circumstances.4

This is a fundamental assurance principle:

realism is not an absolute good.

The objective is to obtain realistic evidence without turning the assurance activity itself into an uncontrolled incident.

TLPT closure shows what “testing with consequences” looks like

The active red-team phase is not the end of the process.

The RTS requires a structured closure including:

  • red-team reporting;
  • blue-team reporting;
  • replay of offensive and defensive actions;
  • purple teaming;
  • process feedback;
  • a findings summary for the TLPT authority;
  • a remediation plan.4

The remediation plan must include, for each finding, elements such as:

  • the identified shortcoming;
  • remediation measures and priority;
  • expected completion;
  • root-cause analysis;
  • responsible staff/functions;
  • risks of not implementing the measures.4

That is a useful assurance principle far beyond finance:

A test is more valuable when the organisation can reliably convert findings into remediation, validation and learning.

TIBER-EU and DORA sit at different legal levels

The ECB updated TIBER-EU in February 2025 to align it with DORA TLPT regulatory technical standards.5

The revised framework aligns:

  • process steps;
  • DORA deliverables;
  • terminology;
  • closure and purple-team activities;
  • operational guidance for controlled TLPT execution.5

But the hierarchy matters.

DORA and Commission Delegated Regulation (EU) 2025/1190 are the binding regulatory basis.

TIBER-EU is an operational framework and guidance that can support DORA TLPT implementation to the extent that it is compatible with the regulation and RTS.

They should not be described as instruments of the same legal status.

Latvia shows TLPT as one layer, not the whole testing programme

Latvijas Banka's current DORA information links TLPT to RTS 2025/1190 and explains that selection considers ICT-governance/process maturity and financial-system criticality, with primary focus on systemically important market participants.6

That does not mean other financial entities have no testing duties.

DORA's architecture is layered:

TLPT is an advanced layer within a wider resilience-testing programme.

It is not a substitute for that programme.

broad resilience testing programme
+
annual appropriate testing of critical/important-function systems
+
advanced TLPT for selected entities

Coverage, depth and realism cannot all be maximised at once

Security-testing methods trade between at least three properties.

Coverage

How much attack surface can be examined?

Automated scanning can achieve very broad coverage.

Depth

How deeply can one boundary, business process or exploit chain be investigated?

Focused penetration testing and specialist research can go much deeper.

Realism

How closely does the exercise resemble an actual adversary operating in the real environment?

Red team and TLPT can provide greater realism.

Greater realism usually also increases:

  • operational risk;
  • coordination complexity;
  • cost;
  • skill requirements;
  • safeguards around the test itself.

No single testing method optimises all three simultaneously.

Planned and unplanned evidence complement each other

Penetration testing gives planned assurance:

We will test this scope on this date.

CVD gives unplanned assurance:

If somebody finds something between tests, we can receive and act on it.

Bug bounty tries to increase the frequency and quality of that external discovery.

Red team tests adversarial paths.

Purple team turns those paths into detection and response learning.

TLPT, where applicable, combines threat intelligence, production realism, offensive and defensive testing, formal closure and remediation.

These are complementary evidence sources.

Time is one of the most important assurance variables

A penetration-test report decays.

Its evidence can be invalidated by:

  • a release;
  • configuration changes;
  • a new dependency;
  • cloud migration;
  • IAM redesign;
  • a new API;
  • supplier changes;
  • new attack techniques.

So:

We had a pentest last year.

is a dated statement.

It is not a current security property.

Continuous assurance does not mean “run a pentest every day”.

It means maintaining evidence at different time scales:

per commit / per release
→ automated security checks

continuous
→ telemetry + vulnerability monitoring + CVD

periodic
→ focused penetration testing

scenario / campaign
→ red or purple team

risk-based / regulated
→ TLPT where applicable

Independence and collaboration are different quality mechanisms

Some assurance activities benefit from independence.

A test designed only to confirm a team's existing assumptions can become self-reassuring.

DORA Article 24, for example, requires relevant testing to be undertaken by independent parties, internal or external, with sufficient resources and conflict-of-interest safeguards for internal testers.3

Red-team value may increase when the blue team does not know the exact attack path.

Afterward, however, collaboration may create more learning than continued secrecy.

Purple teaming deliberately trades independence for:

  • detection tuning;
  • telemetry improvement;
  • playbook validation;
  • shared attack-path understanding.

So:

independence is not always better, and collaboration is not always weaker.

They produce different evidence.

Six common forms of false assurance

“The pentest was clean, therefore we are secure”

No.

It means the selected scope, time and methods did not produce certain validated findings.

“No Critical bug-bounty reports means no Critical vulnerabilities”

No.

Researcher attention is not complete coverage.

“The red team failed, therefore a real attacker would fail”

No.

The exercise was one scenario under defined constraints.

“The red team succeeded, therefore our defence is useless”

Not necessarily.

The useful questions include what was detected, how quickly, where progress was constrained and which controls behaved as expected.

“We have a certificate, therefore controls work”

Certification can provide valuable assurance for a defined scheme and scope. It is not automatically equivalent to current operating effectiveness. See Cybersecurity Certification Should Prove Effectiveness, Not Just Activity.

“We perform TLPT, therefore other testing is unnecessary”

DORA itself says otherwise.

TLPT sits inside a broader digital operational resilience testing programme.3

A layered assurance map

A simplified architecture might look like this.

Design / build

  • threat modelling;
  • architecture review;
  • secure coding;
  • source-code analysis;
  • dependency analysis;
  • automated tests.

Pre-release / change

  • targeted security review;
  • vulnerability assessment;
  • risk-based penetration testing;
  • regression/security tests.

Continuous production

  • vulnerability monitoring;
  • attack-surface monitoring;
  • logging and detection;
  • CVD;
  • bug bounty where suitable.

Periodic adversarial

  • penetration testing;
  • red team;
  • purple team;
  • sector-specific scenario testing.

Advanced / regulated

  • TLPT where applicable;
  • other controlled high-risk testing for critical systems.

Closure

  • root-cause analysis;
  • remediation;
  • deployment evidence;
  • retest;
  • residual-risk decision;
  • regression control.

This is the author's assurance-architecture model, not a regulatory taxonomy.

Every test needs a claim boundary

A security report should say not only:

What did we find?

It should also say:

What does this test allow us to conclude — and what does it not?

For example:

That may prevent more false confidence than another coloured risk matrix.

assurance_claim:
  "No practical cross-tenant access path was identified
   in the tested API scope during the engagement."

does_not_claim:
  - all APIs were tested
  - no unknown vulnerability exists
  - future releases preserve the result
  - identity infrastructure was assessed
  - social engineering was tested

Different discovery sources should converge on one finding lifecycle

A mature organisation may receive findings from:

  • scanners;
  • SAST;
  • penetration tests;
  • red teams;
  • CVD;
  • bug bounty;
  • TLPT;
  • incident response;
  • supplier advisories.

Discovery can remain specialised.

After validation, the evidence should converge on a common decision chain:

This prevents each assurance team from operating its own disconnected remediation universe.

finding
→ validate
→ owner
→ priority
→ remediation
→ deploy
→ retest
→ closure
→ root-cause feedback

What management should ask

Not:

Did we have a pentest this year?

But:

  • Which critical security claims are we testing?
  • Which method tests each one?
  • How fresh is the evidence?
  • What remains untested?
  • Which root causes repeat?
  • How many findings are verifiably remediated?
  • What did red-team activity teach the blue team?
  • Where did external research supplement planned testing?
  • Where is testing coverage inconsistent with business criticality?
  • Where do we have activity evidence rather than effectiveness evidence?

That turns security testing into governance.

Conclusion

Penetration testing is valuable.

Red teaming is valuable.

Purple teaming is valuable.

CVD and bug bounty are valuable.

In finance, TLPT can provide a particularly powerful and regulated form of adversarial resilience evidence.

The problem begins when any one of them becomes a universal claim:

We performed this test, therefore we are secure.

Security assurance is not one test.

It is a system in which different evidence sources intentionally test different security hypotheses and feed a common remediation and learning cycle.

In one sentence:

A pentest is a point-in-time test. Assurance is an evidence system over time.

A strong assurance model does not try to prove that vulnerabilities do not exist.

It repeatedly tests the most important claims about why the organisation expects its systems and controls to remain secure — and what happens when the evidence says otherwise.

Frequently asked questions

Is a red team better than a penetration test?

Not universally. Penetration testing and red teaming optimise different objectives. Pentests usually examine a defined scope for exploitable weaknesses; red teams more often pursue an adversarial objective while also testing detection and response.

Can bug bounty replace an annual penetration test?

Not automatically. Bug bounty provides ongoing external discovery but does not guarantee systematic coverage of a defined scope during a defined period.

Is purple team a separate team?

Not necessarily. It can be a collaborative mode between offensive and defensive specialists. In the specific DORA TLPT context, the RTS defines purple teaming as a collaborative testing activity involving testers and the blue team.4

Does DORA require TLPT for every financial entity?

No. TLPT applies to selected financial entities under DORA's criteria and supervisory process. Latvijas Banka describes a proportionate, risk-based selection approach with primary focus on systemically important entities.6

Is TLPT performed in production?

DORA Article 26 requires TLPT on live production systems supporting the critical or important functions in scope.3 That production realism is why the RTS contains detailed risk-management and suspension controls.4

Does a “clean” penetration-test report prove that no vulnerabilities exist?

No. It records the result of a defined scope, period and methodology. It cannot rule out unknown, excluded or subsequently introduced vulnerabilities.

Source status

Sources and DORA/TLPT legal status checked on 25 September 2026. The assurance map, claim-boundary method and layered model proposed in this article are the author's operational framework, not an official taxonomy of NIST, DORA, the ECB or Latvijas Banka. TIBER-EU is deliberately distinguished from the binding DORA and Commission Delegated Regulation (EU) 2025/1190.

This article analyses security-assurance and testing architecture. The DORA/TLPT discussion is not individual legal or supervisory advice.

Sources

  1. NIST SP 800-115, Technical Guide to Information Security Testing and Assessment, September 2008 · NIST
  2. FIRST, PSIRT Services Framework v1.1, particularly Product Security Assessment, vulnerability discovery, remediation and disclosure services · FIRST
  3. Regulation (EU) 2022/2554 (DORA), particularly Articles 24–27 · EUR-Lex
  4. Commission Delegated Regulation (EU) 2025/1190 on TLPT regulatory technical standards, including testing, closure, purple teaming and remediation planning · EUR-Lex
  5. European Central Bank, TIBER-EU Framework updated to align with DORA, 11 February 2025, and TIBER-EU Framework 2025 · ECB
  6. Latvijas Banka, DORA implementation and subjects, published 23 October 2025, updated 17 March 2026; TLPT section and RTS 2025/1190 · bank.lv