Skip to main content
ANCVEIRS
Professional workGuideResponsible AI

Proof Beats Prose: A Minimum Standard for AI-Assisted Vulnerability Reports

A good vulnerability report is not persuasive prose. It is a reproducible evidence package with a clear boundary between observation, inference and speculation.
Zigmārs AncveirsTechnology Leader in FinTech & RegTech · Cybersecurity & Ethical HackingPublished: 26 September 2026Reviewed: 26 September 202615 min

Introduction

A language model can produce a polished vulnerability report in seconds.

It can name the weakness class, propose a payload, describe an attack chain, generate a CVSS vector, recommend remediation and wrap the entire thing in professional security language.

None of that proves the vulnerability exists.

The atomic unit of a useful security report is not prose.

It is a verifiable relationship between a defined system state, an action and an observable security-relevant result.

That distinction has become operationally important in 2026. HackerOne's current Code of Conduct permits and encourages responsible AI assistance but keeps the human researcher responsible for validating findings, connecting attack-chain steps and providing a clear, reproducible proof of concept.2 Bugcrowd requires human review and validation of vulnerability reports produced with GenAI assistance.4 Intigriti requires researchers to personally identify, test and understand findings and prohibits fabricated endpoints, generic exploit templates and misleading impact claims.5 YesWeHack has likewise framed unverified AI-generated hypotheses and program spamming as a serious quality problem.8

The wording differs by platform.

The operational principle does not:

AI can assist the work. It cannot inherit the researcher's accountability for the evidence.

Generation became cheap; validation did not

Generative AI changes the appearance of weak reports faster than it changes the evidence required to validate them. A report can now be long, grammatical and professionally structured while still containing an untested endpoint, a non-working payload, a missing prerequisite or an impact claim that was never demonstrated.

The system-level economics of that shift are analysed separately in AI Slop Is Not a Content Problem. It Is a Verification-Cost Problem.. This article stays at the report boundary: what must be true before a vulnerability hypothesis is ready to submit?

A hypothesis is not a finding

An AI-assisted research workflow benefits from explicit state transitions.

AI can help at every stage.

It can flag code that resembles a path traversal.

It can suggest a test.

It can generate candidate payloads.

It can cluster responses.

It can organise notes and draft the final report.

What it must not do is silently promote:

“this may be vulnerable”

into:

“this is vulnerable.”

YesWeHack published a useful 2026 example in which an LLM correctly identified a DOM-XSS condition in source code but produced a proof-of-concept payload that the application's filter blocked. The vulnerability hypothesis was right; the proposed proof was wrong. Human analysis was still needed to produce a working PoC.9

That distinction is the centre of this article.

signal / anomaly
      ↓
hypothesis
      ↓
verified observation
      ↓
reproduced security condition
      ↓
bounded impact claim
      ↓
report

Human validation does not mean “no automation”

A serious standard should not romanticise manual clicking.

Security research has always used automation:

  • fuzzers;
  • scanners;
  • crawlers;
  • static analysis;
  • symbolic execution;
  • custom scripts;
  • exploit harnesses;
  • large-scale differential testing.

AI expands that toolbox.

The correct requirement is therefore not:

“The vulnerability must have been found manually.”

It is:

“The submitter has personally verified and understands that the submitted evidence applies to the stated target and supports the stated claim.”

A deterministic script can be stronger evidence than ten manual screenshots.

A fuzzer can discover real memory corruption.

A crash can be valid without a weaponised exploit.

A race condition may require hundreds of attempts and a documented success rate rather than one click-by-click reproduction.

The important property is a verified causal chain, not the absence of automation.

This is an old PSIRT problem amplified by a new tool

FIRST's PSIRT Services Framework already recommends that organisations define and publish a minimum quality bar for vulnerability reports. Its baseline examples include a write-up, reproduction steps, tested platforms and a proof of concept. It separately treats vulnerability reproduction as a PSIRT function needed to validate and understand the conditions leading to a vulnerable state.1

AI did not create the need for reproducibility.

It made it much easier to manufacture the appearance of reproducibility.

That is why evidence needs to become more explicit, not more verbose.

The minimum evidence contract

I use one rule as the foundation:

Every material claim in a vulnerability report should trace to an observation, a reproducible step or an explicitly labelled assumption.

A practical minimum follows from that.

1. Scope and authority context

The report or its metadata should identify:

  • the applicable program or disclosure policy;
  • the tested asset;
  • any time or environment restrictions that matter;
  • material testing limitations imposed by the program.

AI-assisted testing does not expand scope.

HackerOne's current AI rules explicitly require autonomous and semi-autonomous tools to comply with program scope, automation limits, request-volume restrictions and rate limits.2

The same principle applies outside platforms.

If an AI agent discovers a plausible critical issue on an out-of-scope third-party asset, the phrase critical vulnerability does not create authority to keep exploiting it.

For the distinction between CVD, bug bounty and commissioned testing, see CVD, Bug Bounty and Penetration Testing Are Not the Same Thing.

2. Precise target identity

“example.com is vulnerable” is usually insufficient.

The recipient needs enough target identity to recreate the finding:

  • host, URL or API route;
  • mobile build;
  • software version;
  • repository/commit for source findings;
  • relevant feature flag or configuration;
  • account role or tenancy context;
  • test time where the system changes rapidly.

LLMs are good at filling missing context with typical architecture.

A vulnerability report needs the architecture you actually tested.

3. Preconditions

Many inflated reports hide important prerequisites.

Examples include:

  • an authenticated account;
  • administrator privileges;
  • victim interaction;
  • knowledge of an object identifier;
  • same-tenant access;
  • a specific enabled feature;
  • another vulnerability earlier in the chain;
  • a particular browser or deployment mode.

These conditions are not footnotes.

They are part of the vulnerability.

A report that hides them to preserve a stronger severity claim is less accurate, not more persuasive.

4. Reproduction steps

A technically competent recipient should be able to tell:

  1. where to start;
  2. exactly what action or request to make;
  3. what result to expect;
  4. how to distinguish successful exploitation from normal behaviour.

HackerOne's current submission standard emphasises step-by-step, reproducible impact and a strong PoC.2 Intigriti's reporting guidance similarly says that the information needed to reproduce a vulnerability should be in the report itself rather than existing only in a screenshot or video.6

A simple quality test is:

Can the triager reproduce this without guessing my missing step?

5. Minimal PoC or equivalent evidence

A proof of concept is not synonymous with exploit code.

Depending on the vulnerability, sufficient evidence may be:

  • a raw HTTP request/response pair;
  • two researcher-controlled accounts demonstrating broken authorisation;
  • a crash input plus stack trace or sanitizer output;
  • a minimal script;
  • an event log proving an unauthorised state change;
  • a race-condition harness with repeat count and observed success rate;
  • a video as supplemental evidence, while the textual reproduction remains complete.

Intigriti's 2026 triage standard describes the default as a written step-by-step attack demonstration and asks for the simplest possible demonstration that establishes exploitability and impact beyond reasonable doubt.7

The important word is simplest.

The PoC is not theatre.

It is a measurement instrument.

6. Actual versus expected behaviour

A report should state both.

Actual

Authenticated user A can retrieve object B belonging to user B by changing /api/orders/{id}.

Expected

Server-side authorisation should prevent user A from accessing objects outside the authorised account boundary.

Expected behaviour needs a basis.

That may be:

  • the product's role model;
  • API documentation;
  • program documentation;
  • an explicit tenant-isolation boundary;
  • consistent behaviour elsewhere;
  • a documented control.

If the expected security property is uncertain, say so.

“Best practice says this should be secure” is not a substitute for understanding the boundary.

7. Demonstrated impact versus possible impact

This is where AI-assisted reports often become fiction.

A useful report separates three layers.

Demonstrated impact

What the PoC actually establishes.

User A can read user B's invoice title and amount.

Supported additional impact

What follows reasonably from the evidence, with explicit assumptions.

The same server-side authorisation pattern may expose other invoice objects. Bulk enumeration was not performed in order to minimise data exposure.

Speculative downstream impact

Complete account takeover, RCE, regulatory penalties, full database compromise.

If there is no demonstrated chain to the third category, it should not be written as fact.

LLMs are extremely good at plausible worst-case narratives.

A vulnerability program needs bounded claims.

Severity describes the finding; it does not validate it

Many platforms ask the researcher to propose severity.

As of 21 September 2026, HackerOne requires a severity selection for submissions to programs that have enabled that requirement.3 Latvia's CERT.LV platform also requires use of an internationally recognised CVSS calculator for impact assessment in its reporting process.12

Neither changes the evidence rule.

A CVSS score cannot promote an unverified hypothesis into a vulnerability.

FIRST's CVSS v4.0 documentation explicitly distinguishes Base severity from organisational risk and provides separate Threat and Environmental metric groups for context.13

For AI-assisted reporting, a useful severity section therefore includes:

  • CVSS version;
  • vector string where CVSS is used;
  • short reasoning for disputed metrics;
  • a boundary between demonstrated and assumed impact;
  • the program's own scoring policy where it differs.

A triager changing the score later does not make the technical report invalid.

Repeatedly inflating the score beyond the evidence is a different quality problem.

Existing mitigations are evidence too

A model may identify a plausible attack path while ignoring a control that blocks it in the real deployment.

HackerOne's current standards explicitly ask reports to account for real-world mitigations and defence-in-depth, and require a practical exploit path where an existing control would otherwise defeat the claim.2

This does not mean a compensating control erases an underlying defect.

It means the report should record:

  • which mitigation was present;
  • whether it actually stopped exploitation;
  • whether the researcher bypassed it;
  • whether the defect exists only in an unsupported configuration;
  • whether another security boundary remains broken.

Context should not be outsourced entirely to triage.

Data minimisation is part of report quality

The strongest PoC is not the one that extracts the most data. If one controlled response proves an authorisation failure, collecting thousands of real records usually adds risk faster than it adds technical certainty.

Current platform rules similarly emphasise limiting exposure and collecting only what is needed to demonstrate the condition.512 The detailed GDPR, evidence-retention and breach-handling questions are treated in When Vulnerability Research Exposes Personal Data.

Confidential vulnerability data is not automatically safe LLM input

AI-assisted report writing creates another risk unrelated to hallucination.

A researcher may paste:

  • private program identifiers;
  • undisclosed endpoints;
  • session tokens;
  • source code;
  • customer data;
  • PoC material;
  • internal architecture details

into a third-party AI service.

The result may be a better paragraph and a worse security outcome.

Bugcrowd's GenAI rules explicitly require researchers to preserve the confidentiality and security of program-owner information and findings.4 Intigriti also restricts sharing confidential private-program and vulnerability-reproduction information with third parties.5

A professional AI workflow therefore needs a pre-prompt question:

Am I permitted to send this information to this model/provider?

Model choice is part of operational security.

Duplicate is a triage state, not a quality verdict

A technically excellent report can be a duplicate.

External researchers normally cannot see a program's private report history.

“Prove this is not a duplicate” is therefore not a realistic universal admission criterion.

The researcher can still check what is available:

  • their own prior reports;
  • public advisories;
  • known CVEs or vendor issues;
  • whether several endpoints are merely variants of one root cause under the program's rules.

AI makes it cheap to generate one report per endpoint for a repeated pattern.

That does not make each instance an independent vulnerability.

Program policy decides how variants and root causes are handled.

AI disclosure requirements are platform-specific

There is also no universal industry rule requiring every report to display an “AI-generated” label.

Intigriti currently requires researchers to disclose when and how AI was used in a submission.5

Other platforms phrase the obligation differently.

My minimum standard is therefore not:

“Every vulnerability report must carry an AI disclosure banner.”

It is:

Follow the actual program's disclosure requirement, and preserve enough working provenance to personally defend every claim in the report.

The important test is not an AI detector.

It is whether the researcher's name beneath the submission means:

I verified this and I understand what I am claiming.

A minimal report record

A compact evidence-oriented report can look like this:

reporter_attestation is not a legal oath.

Its meaning is operational:

Every endpoint, payload, request/response result and exploitation step stated as fact has been checked against the stated target; hypothetical continuation is labelled as hypothetical.

That single rule eliminates a remarkable amount of report slop.

report_title:
program_or_policy:
asset:
tested_version_or_context:
test_time:

preconditions:
attacker_position:

actual_behavior:
expected_security_boundary:

reproduction_steps:
minimal_poc:
evidence_refs:

demonstrated_impact:
supported_additional_impact:
not_tested_or_not_proven:

mitigations_observed:
limitations:

severity_method:
severity_vector_or_rationale:

sensitive_data_handling:
ai_assistance_disclosure_if_required:

reporter_attestation:

The pre-submission evidence gate

Before pressing Submit, I would ask:

This is a minimum gate, not a maximum report template.

The pre-submission evidence gate
QuestionIf the answer is no
Is the target in scope?stop or clarify authority
Does the condition exist on the stated target/version?do not submit it as a vulnerability
Can I reproduce it, or accurately describe its nondeterminism?gather better evidence
Do the steps include all material preconditions?add them
Does the PoC prove the exact claim in the title/impact?narrow the claim or fix the PoC
Are actual and expected behaviour explicit?define the security boundary
Is the impact demonstrated rather than narrated?rewrite the impact
Does severity describe the same condition I proved?rescore
Would further testing create unnecessary risk or data exposure?stop
Did AI invent any endpoint, feature, response, mitigation or attack step?verify each technical detail
Am I allowed to use this AI service with the report data?keep sensitive material out
Can I defend every sentence without saying “the model said so”?the report is not ready

What a report does not need

Quality standards can become bureaucracy too.

A valid report does not inherently need:

  • a long narrative;
  • perfect English;
  • a diagram;
  • an end-to-end compromise of the most valuable asset;
  • a CVE candidate;
  • a patch;
  • a polished AI-written executive summary;
  • a video;
  • multiple screenshots;
  • a speculative monetary-impact calculation.

OWASP's vulnerability-disclosure guidance focuses on sufficient reproducibility, supporting evidence and impact, not rhetorical volume.11

A short report that reproduces cleanly is better than two thousand words of possibility.

For program owners: enforce the evidence contract at intake

The programme-side implication is deliberately small: make the evidence contract explicit at intake. Define the minimum report fields, distinguish insufficient evidence from a demonstrated false positive, automate cheap preflight checks, and preserve human validation for the security conclusion.110

Broader questions about queue capacity, throttling, reputation, openness and verification economics belong to AI Slop Is Not a Content Problem, not to this reporting standard.

Conclusion

AI does not need to be banned from security research.

It needs to be placed correctly in the evidence chain.

AI can surface the anomaly.

It can propose the hypothesis.

It can suggest a test.

It can assist with code analysis.

It can draft the report.

But by the time the researcher presses Submit, the report should no longer be a model's hypothesis.

It should be a researcher-verified claim with reproducible evidence and a bounded impact statement.

In one sentence:

A vulnerability hypothesis is not a vulnerability finding.

And a good report is not the one that sounds convincing.

It is the one that can be checked.

Frequently asked questions

Can AI be used to write bug bounty or CVD reports?

It depends on the applicable program rules, but major platforms in 2026 generally do not prohibit AI assistance as such. Their emphasis is on researcher accountability, validation, reproducibility, scope and confidentiality.245

Does every vulnerability report need exploit code?

No. It needs sufficient reproducible evidence. Depending on the issue, that can be a request/response pair, crash reproducer, minimal script, controlled PoC or another evidence form that establishes the security condition and impact.

Must a vulnerability reproduce 100% of the time?

No. Race conditions, concurrency defects and probabilistic systems can be real at lower reproduction rates. The report should document conditions, number of attempts, observed success rate and enough evidence for the receiving team to reproduce the behaviour.

Is AI-use disclosure mandatory everywhere?

No. Requirements differ by platform and program. Intigriti, for example, currently requires disclosure of when and how AI was used in a submission.5 The current program policy is the authoritative rule.

Does a high CVSS score prove that the vulnerability is valid?

No. CVSS characterises vulnerability severity. It does not replace reproduction or evidence, and FIRST explicitly distinguishes CVSS Base severity from organisational risk.13

Is a duplicate a bad report?

Not necessarily. A real, well-documented vulnerability may already have been reported privately. Duplicate status is a coordination and triage outcome, not an automatic judgment that the research was technically poor.

Source status

Platform policies and technical sources were checked on 25 September 2026. Platform rules can change; researchers must follow the current policy of the specific bug bounty or VDP program. The “minimum evidence contract”, report record and pre-submission evidence gate in this article are the author's practical model, not a single official standard issued by FIRST, CERT.LV or any bug bounty platform.

This article analyses security-research and vulnerability-reporting practice. It is not individual legal advice.

Sources

  1. FIRST, PSIRT Services Framework v1.1, particularly Finder Report Quality and Vulnerability Reproduction · FIRST
  2. HackerOne, Code of Conduct, “AI-assisted Research & Submission Standards”, checked 25 September 2026 · hackerone.com
  3. HackerOne Help Center, Submitting Reports, checked 25 September 2026 · HackerOne
  4. Bugcrowd, Code of Conduct, “Responsible use of GenAI tools”, updated 25 November 2025, checked 25 September 2026 · bugcrowd.com
  5. Intigriti, Community Code of Conduct, 9 March 2026 · Intigriti
  6. Intigriti, How to write and submit a good report, 10 April 2026 · Intigriti
  7. Intigriti, Triage Standards, checked 25 September 2026 · Intigriti
  8. YesWeHack, YesWeHack Report 2026, section on AI, human-in-the-loop research and “program spamming and AI slop” · choose.yeswehack.com
  9. YesWeHack, “How to hack with LLMs, agentic CLIs, MCP servers”, 2026 · YesWeHack
  10. YesWeHack, “Scaling Bug Bounty triage in the AI era”, 19 May 2026 · YesWeHack
  11. OWASP Cheat Sheet Series, Vulnerability Disclosure Cheat Sheet, checked 25 September 2026 · cheatsheetseries.owasp.org
  12. CERT.LV, Vulnerability Reporting Platform Terms of Use, effective 1 August 2026, particularly sections 5.5–5.10 · CERT.LV
  13. FIRST, Common Vulnerability Scoring System v4.0 Specification and User Guide · FIRST