Introduction
Vibe Coding has created an interesting situation. People who previously lacked the programming knowledge to build an application independently can now create working software with the assistance of generative AI. For some, this has opened genuine opportunities to develop ideas, acquire technical skills and deliver useful products. Others have discovered that getting an application to run is rather different from understanding how to maintain, debug, extend and secure it. The first demonstration may look impressive. The difficult questions tend to arrive later, when something breaks, dependencies change, access controls fail or the application needs to support more users than originally anticipated.
A similar development is now visible in cybersecurity. AI tools can analyse source code, suggest attack scenarios, interpret scanning results, produce vulnerability reports and recommend remediation steps. Experienced security researchers use these capabilities, as do people entering the profession. Their differences in knowledge and experience are not necessarily apparent in the resulting documents. A technically sophisticated report can be produced without the person submitting it fully understanding the system architecture, the conditions required to exploit a vulnerability or the actual security implications of the finding.
I would describe this phenomenon as Vibe Cybersecurity: conducting security-related activities, making technical judgements or presenting professional expertise while relying heavily on AI-generated conclusions that the individual cannot adequately verify, explain or defend independently.
This is a proposed analytical term, not an established industry classification. It is not intended to describe everyone who uses AI for cybersecurity, nor does it imply that someone without formal qualifications cannot produce valuable security research.
The practical question is simpler: what can the person actually demonstrate when the AI-generated answer turns out to be wrong?
When working software creates false confidence in security
Functional correctness and security are different properties. An application may successfully authenticate users, process payments, store documents and pass routine functional tests while still containing broken access controls, insufficient tenant isolation or unsafe input handling. These weaknesses may remain invisible during ordinary use because they only emerge when someone deliberately tests assumptions the original developer never considered.
Veracode's Spring 2026 GenAI Code Security Update illustrates the distinction. The company examined AI-generated code across 80 programming tasks involving Java, JavaScript, C# and Python, covering four specified vulnerability categories. Its testing found that approximately 45% of generation cases contained a known security flaw under the study's conditions. The findings do not establish that 45% of all AI-generated software is vulnerable.5 They do, however, provide a concrete demonstration that improvements in generating functional code do not automatically deliver equivalent improvements in security.
Consider a small company that uses AI to develop a customer portal. Users can sign in, retrieve documents and update their account details. The interface works, authentication tests pass and the application appears ready for deployment. What nobody has tested is whether changing a document identifier in an API request allows one customer to retrieve another customer's records. Perhaps nobody has examined how data is isolated between business customers either.
The company then asks an AI assistant to generate security tests. The results look reassuring, and a polished assessment declares the application secure. But if the original development process failed to recognise an important trust boundary, the generated security assessment may overlook the same boundary. Both activities can appear complete while sharing an untested assumption.
This is an illustrative scenario, not a documented incident involving a particular company. Its relevance lies in the failure mechanism: using automated assessment as evidence of security without establishing whether the assessment covers the actual security properties that matter.
What does a vulnerability look like when it does not really exist?
Generative AI can be extremely useful for developing vulnerability hypotheses. A model may identify suspicious input handling, an unusual data flow, a potentially unsafe library call or a configuration worth investigating. A competent researcher then needs to establish whether the affected code is reachable, what preconditions apply, which protections are already present and whether exploitation would have meaningful security consequences.
Without that work, a report may describe a critical vulnerability that nobody can reproduce.
A static analysis finding might identify a dangerous function that is never reachable from an attacker-controlled input. A configuration may look unsafe in isolation but be protected by controls elsewhere in the deployment. A model may identify a genuine programming defect without demonstrating that the defect creates an exploitable security weakness. The report could still contain convincing technical language, a plausible attack narrative and a proposed patch.
The missing evidence matters more than the quality of the prose.
In March 2026, Google published changes to its Open Source Software Vulnerability Reward Program following a substantial increase in AI-generated submissions. It described reports containing incorrect explanations of how vulnerabilities could be triggered, as well as findings involving coding errors with negligible security impact or unreachable code paths. Google responded by introducing stronger evidence requirements for certain categories, including precise OSS-Fuzz reproduction steps or a merged patch in specified circumstances.4
HackerOne reported a related development in May 2026. According to the company's assessment, the wider industry had experienced an increase of more than 100% in vulnerability report volume following the arrival of more capable AI tools in February. Some submissions represented genuine discoveries; others were duplicates, unverifiable reports or findings without sufficient detail for meaningful action.3
Those numbers do not tell us how many people have become competent security researchers. Nor do they measure the proportion of practitioners whose expertise depends excessively on AI. They do show how quickly the capacity to produce vulnerability reports can grow, while the work needed to evaluate those reports remains substantial.
The curl story is more complicated than the headlines suggested
In January 2026, curl maintainer Daniel Stenberg announced that the project's monetary bug bounty programme would end. One of the central concerns was the growing burden of poor-quality security submissions. Historically, more than 15% of submitted reports had resulted in confirmed vulnerabilities. During 2025, that figure fell below 5%. Maintainers were spending increasing amounts of time investigating and responding to claims that did not establish genuine security problems.1
It would be tempting to stop there and present curl as proof that AI had damaged vulnerability research.
That would leave out the more interesting part of the story.
In April 2026, Stenberg reported a striking improvement. The project had returned to HackerOne in March as a vulnerability reporting channel, even though the original monetary reward programme had been discontinued. Submission volume was higher than before, but the quality of reports had also improved considerably. The confirmed vulnerability rate had returned to approximately 15–16%, matching or exceeding its earlier level.2
Stenberg also observed that AI assistance appeared to be involved in almost every report to varying degrees. Unlike the earlier wave of poor-quality submissions, however, many of these reports were genuinely useful.
This example deserves more attention than the usual argument about whether AI is good or bad for security research. The same broad class of tools can contribute both to an overwhelming volume of unsupported claims and to the discovery of real, technically significant vulnerabilities.
The existence of AI assistance tells us relatively little about the quality of an individual finding. What matters is whether the researcher can establish that the finding is real, relevant and properly supported.
Cybersecurity involves far more than finding vulnerabilities
The potential consequences of Vibe Cybersecurity extend well beyond bug bounty platforms. Similar failure mechanisms can appear in penetration testing, security consultancy, incident response, security operations and compliance assessments.
Imagine a consultant delivering a security audit containing references to recognised standards, risk classifications, technical findings and extensive remediation recommendations. The document may be well written and internally consistent. Yet if the consultant has not examined actual access permissions, system configurations, data flows or operational controls, its apparent sophistication may greatly exceed its evidential value.
Incident response creates another set of risks. An AI assistant can produce a plausible incident timeline from incomplete logs. If the analyst does not know which telemetry sources are missing, how timestamps should be interpreted or which activities were never recorded, a convincing narrative may be accepted as the root cause. Decisions based on that narrative can lead to ineffective containment, unnecessary configuration changes or incorrect public explanations.
The same issue appears in compliance work. AI can help map organisational controls to requirements in standards and regulations. But a control description is not evidence that the control operates as intended. If a security policy states that privileged access is reviewed regularly, I would want to see when the last review took place, what it found and whether the resulting issues were resolved. A perfectly written policy does not answer those questions.
NIST's Generative AI Profile addresses the risk of AI systems confidently producing incorrect information, including misleading explanations and references. It provides a useful foundation for examining this problem, although it does not establish a cybersecurity-specific definition of Vibe Cybersecurity or prescribe a single method for validating security findings.6
An organisation still has to decide what evidence is acceptable, how results will be independently checked and who is responsible for decisions made on the basis of those results.
Verification debt and the cost of premature confidence
Software engineering has long recognised the concept of technical debt. A shortcut may help a team deliver something quickly, but it can create additional work when the system needs to be maintained, extended or corrected.
I see a related problem emerging in AI-assisted security work. I would call it verification debt.
Verification debt accumulates when hypotheses, assessments or security conclusions are accepted and reused without sufficient validation. One person generates a list of suspected vulnerabilities. Another turns that list into a risk assessment. A third uses the assessment to prioritise remediation. The number of documents increases, but the original claims may still be unverified.
If one of those claims is eventually disproved, the organisation may need to revisit every decision that depended on it.
Sometimes the error is never discovered. Management receives a positive security assessment, the project moves forward and the assumptions underlying that assessment remain unchallenged. That can be more dangerous than openly acknowledging uncertainty, because confidence discourages further investigation.
Verification debt should not be treated as an inevitable consequence of using AI. Well-designed automation, reproducible tests, independent review and documented decisions can reduce the amount of work required to establish confidence. The problem arises when apparent certainty travels further than the evidence supporting it.
This is an analytical concept rather than an established measurement standard. It is useful because it identifies a cost that conventional productivity metrics may overlook: the downstream burden created by decisions based on inadequately tested assertions.
How can we distinguish expertise from its imitation?
I would not begin by asking which university someone attended, how many certifications they possess or which AI model they use. Such information may be relevant to a wider professional assessment, but it cannot establish the quality of a specific security finding.
I would select one significant finding and ask its author to walk through the complete evidence chain.
First, understanding the system. What exactly is being assessed? Which components are involved, how do they communicate and where are the relevant trust boundaries? Can the person explain the architecture without relying on a prepared summary?
Second, testing the hypothesis. What observations suggest a security weakness? What conditions are required for the suspected behaviour? What evidence would disprove the hypothesis? If an AI model highlighted suspicious code, can the researcher demonstrate an actual execution path?
Third, reproducibility. Can the result be reproduced safely within an authorised environment? Are the required software versions, assumptions, inputs and observations documented? Is it clear where direct observation ends and interpretation begins?
Fourth, actual security impact. What does the weakness allow an attacker to achieve? Which controls may prevent exploitation? Why does the behaviour constitute a security vulnerability rather than merely a software defect or a theoretical concern?
Finally, remediation and retesting. Does the proposed change resolve the identified weakness without introducing another problem? How has that been checked? Who accepts any remaining risk?
I propose this five-stage sequence as a practical evaluation method, not as an official NIST standard or professional certification scheme. It provides a way to assess the relationship between a researcher's conclusions and the evidence behind them.
Authorisation remains part of competent security practice throughout. Being capable of reproducing a vulnerability does not automatically provide permission to test someone else's systems. A responsible researcher must understand both how to investigate and when to stop.
| Stage | Question the author should be able to answer |
|---|---|
| 1. Understanding the system | What exactly is being assessed, where are the trust boundaries and how do the relevant components interact? |
| 2. Testing the hypothesis | What evidence suggests a security weakness, and what observation would show that the hypothesis is wrong? |
| 3. Reproducibility | Can the result be reproduced safely in an authorised environment with known preconditions and the same outcome? |
| 4. Actual impact | What does the weakness really allow, which controls constrain it and why does the behaviour create security risk? |
| 5. Remediation and retesting | Does the proposed change fix the problem without introducing another regression, and how was that verified? |
What AI can accelerate — and what it cannot prove
The practical boundary between productive AI use and Vibe Cybersecurity is not whether AI is used at all. It is the boundary between assistance with analysis and substitution for a conclusion.
AI can be extremely effective at generating hypotheses, locating related code paths, summarising large volumes of data, proposing test variations, comparing configurations and documenting results that have already been verified. But an AI-generated conclusion does not by itself establish reachability, exploitability, authorisation or actual business impact.
AI can accelerate the path to evidence. It should not become the evidence itself.
A professional security process should therefore preserve the provenance of its conclusions. Where did the hypothesis come from? Which data was used? What was tested? Which negative results were recorded? Who decided that the finding was sufficiently supported? The more work is automated, the more important that chain becomes.
What does it mean to be a cybersecurity professional in the age of AI?
I consider AI a valuable instrument for security professionals. Used well, it can accelerate source-code analysis, reveal relationships between vulnerabilities, assist with test development and help researchers examine alternative explanations. Someone with strong technical knowledge and the ability to challenge a model's conclusions can gain substantial advantages from these capabilities.
The difficulty lies in how professional credibility is assessed.
A technically sophisticated report once provided at least some indication that its author had invested meaningful effort. Today, the same level of presentation can be achieved much more quickly. That makes the quality of supporting evidence, reproducibility, understanding of limitations and responsibility for decisions increasingly important indicators of professional competence.
Newcomers should not be dismissed. Someone who uses AI to discover a real vulnerability, reproduces it safely and explains its impact has made a valuable contribution, regardless of their professional background. Equally, an impressive profile, a collection of certifications and dozens of polished reports cannot make unsupported conclusions reliable.
I would have more confidence in a researcher who can explain what they have not yet verified than in someone who appears to have an immediate answer to every technical question. Recognising uncertainty is often a sign that the person understands the limits of the available evidence.
Vibe Coding has demonstrated how quickly someone can reach a working software prototype. Vibe Cybersecurity describes the risk of reaching an apparently authoritative security conclusion just as quickly. In both cases, the quality of the result becomes clearer when somebody has to examine what lies beneath the presentation.
If someone tells me a system is secure, I want to know how they reached that conclusion.
And, just as importantly, what they did not test.
Frequently asked questions
Does using AI in cybersecurity mean a practitioner is less competent?
No. The use of AI tells us very little by itself about competence. What matters is whether the practitioner can independently verify, reproduce and explain the conclusions produced with AI assistance.
What is Vibe Cybersecurity?
In this article, Vibe Cybersecurity describes security work or professional expertise that relies heavily on AI-generated conclusions that the person cannot adequately verify or defend independently. It is an analytical term proposed by the author, not an established industry classification.
Is an AI-generated vulnerability report automatically low quality?
No. The 2026 curl experience shows that AI-assisted reports can be either poor or highly valuable. The relevant questions are whether the finding is reproducible, supported by evidence and associated with meaningful security impact.
How do you distinguish a possible defect from a demonstrated vulnerability?
A suspicious code fragment or scanner alert is not enough. At minimum, reachability, preconditions, existing controls, reproducibility and actual security impact need to be established.
What is verification debt?
Verification debt is the downstream burden created when insufficiently validated claims are reused in later assessments, decisions and remediation priorities. The more work that depends on an unverified starting point, the more expensive it becomes to correct the chain later.
Can AI fully automate a security audit?
Individual audit activities can be automated extensively, but conclusions about real-world security risk still require context, evidence and accountable judgement, particularly where business logic, authorisation boundaries or incomplete incident data are involved.
Which skill becomes more important in the age of AI?
Not merely finding a possible problem, but proving that it exists, understanding its limits and being equally clear about what has not yet been established.
References
- Daniel Stenberg, “The end of the curl bug-bounty”, 26 January 2026 · daniel.haxx.se · 2026-01-26
- Daniel Stenberg, “High-Quality Chaos”, 22 April 2026 · daniel.haxx.se · 2026-04-22
- HackerOne, “Navigating the AI Wave: How We're Keeping Security Research Meaningful”, 29 May 2026 · hackerone.com · 2026-05-29
- Google Bug Hunters, “Streamlining Google's OSS VRP: Key Rule Updates”, 19 March 2026 · bughunters.google.com · 2026-03-19
Updated in April 2026.
- Veracode, “Spring 2026 GenAI Code Security Update”, 2026 · veracode.com
- National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1, 26 July 2024 · NIST · 2024-07-26