← PortfolioHow software gets made

Vidoc Security

Securing AI Generated Code

Team
Klaudia KlocCEO
Dawid MoczadłoCTO
Founded
2021
Invested
2024
The problem

How do you find the real security holes in code nobody really wrote?

Click a service to open it to the internet or close it off. The warnings are checked again, and the short list changes with what an attacker can reach.On the left, every warning a scanner raises. In the middle, the system the code runs in: which services face the internet, what calls what, and where the data sits. Each warning is checked against that map, and only the ones an attacker could actually reach make the short list, each with its severity, the way in and a fix.An illustration, not real data.
How code gets broken into

Most security holes come from the same short list of mistakes, made over and over. OWASP keeps a Top 10 of the most critical risks to web applications, a broad consensus of what goes wrong. At the top sits broken . Access control is the part of an app that stops users acting outside their permissions, and every application tested had some form of it broken.

Injection is the other old favourite. Untrusted input reaches an interpreter, such as a database, a browser or the command line, and gets run as commands. is the famous version: someone types SQL into a form field and the database obligingly runs it, which can let them dump the data, change balances or make themselves the administrator.

Tools that hunt for these mistakes by reading code without running it are doing , which security people call . They scale well, can run on every build, and catch well-known classes such as buffer overflows and SQL injection flaws.

Further reading OWASP Top 10:2025 (OWASP)A01 Broken Access Control (OWASP)A05 Injection (OWASP)SQL injection (Wikipedia)Static program analysis (Wikipedia)Source Code Analysis Tools (OWASP)

Why it is hard
  1. i.

    No perfect checker

    There is a proof in the way. Rice's theorem says every non-trivial question about what a program does is undecidable, so nobody can build a tool that decides whether any given program runs without error. A real tool has to either overestimate or underestimate. Every scanner is paranoid or forgetful, and the vendor chose which.

  2. ii.

    Crying wolf

    Most tools pick paranoid, and it shows. Static tools are known for high numbers of , and it is hard for them to prove that a flagged issue is an actual vulnerability. They also catch only a relatively small share of security flaws. A team that gets two hundred warnings a week learns to ignore all of them, including the one that mattered.

  3. iii.

    Bugs that need context

    The worst bugs are often the ones a pattern can't see. Authentication problems and access control issues are hard to search for automatically, and configuration mistakes don't live in the code at all. Take , where an attacker gets a server to send requests on their behalf to internal systems. Whether that is a shrug or a disaster depends on what the server can reach, which no single file tells you.

  4. iv.

    Code nobody wrote

    Now add code written by AI. In one study, researchers had GitHub Copilot complete 89 security-relevant scenarios, producing 1,689 programs, and about 40% were vulnerable. In a user study, people with an AI assistant wrote significantly less secure code than those without one, and were more likely to believe it was secure. Models also invent software packages that don't exist, at least 5.2% of the time for commercial models and 21.7% for open-source ones, which is an open door for anyone who registers the name first.

Further reading Rice's theorem (Wikipedia)Source Code Analysis Tools (OWASP)Server-side request forgery (Wikipedia)Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions (arXiv)Do Users Write More Insecure Code with AI Assistants? (arXiv)We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs (arXiv)

What Vidoc Security is after

Vidoc wants a security reviewer that behaves like an engineer who already knows your repos, instead of one more scanner. It is built for context-aware vulnerability validation for AI-era software teams, where much of the code arrives from assistants.

The goal is fewer, better findings: connect a repository and get back a short, prioritized list of real AppSec issues, each with its severity, whether it is reachable, and a fix.

Further reading Product (Vidoc Security Lab)

How they go at it
  1. Step 1: Map the whole system

    Vidoc maps an organisation the way an architect would, every service and data store and how they connect. That is how it can tell what is exposed, what is internal, and what an attacker can actually reach.

  2. Step 2: Check before shouting

    It reviews every , then checks each critical finding for real exploitability before a person sees it. What survives comes with a PR-ready fix prompt for an engineer to review.

  3. Step 3: Learn from no

    When a finding doesn't apply, someone can say why in Slack or in the PR, in plain English. Vidoc remembers that per repo and per team, and every suppression is audit-logged.

  4. Step 4: Hackers on staff

    Behind the product is a research team that studies how systems break. Its finds include out-of-bounds reads in the kernel's netfilter decoders for H.323, a network-reachable bug that needed no authentication.

Further reading Product (Vidoc Security Lab)Security Research (Vidoc Security Lab)

Still open
  • How good are language models at finding real vulnerabilities?

    Honestly, nobody is sure. How well they spot complex vulnerabilities in real-world code remains poorly understood, partly because existing benchmark datasets don't look much like real security reviews.

  • What makes AI-assisted code safer?

    In the user study above, the people who trusted the AI less and put more work into their prompts, rephrasing and adjusting settings, wrote code with fewer security vulnerabilities. A healthy scepticism seems to help, though a lab task is not a production codebase.

Further reading Security Research (Vidoc Security Lab)Do Users Write More Insecure Code with AI Assistants? (arXiv)

About Vidoc Security

Vidoc is building an agentic security engineer designed to understand a company's codebase deeply and operate directly within modern CI/CD workflows. The product focuses on analyzing real code and its structure, rather than relying on surface-level scans or generic language-model reasoning.

As AI-generated code increasingly floods production systems, traditional AppSec tools struggle to separate real vulnerabilities from noise. Vidoc is designed to reason over codebases with an attacker's mindset, prioritizing high-signal issues that matter in practice.

Vidoc is built for the security engineers and DevSecOps teams who need something smarter than a scanner — a system that actually understands what the code does and flags the issues that matter.

Words used here
access control
The rules in an application that decide who may see or change what.
SQL injection
An attack where input typed into an app is run by its database as a command.
static analysis
Checking a program for problems by reading its code, without running it.
SAST
Static application security testing: static analysis aimed at security bugs.
false positives
Warnings about problems that turn out not to be real.
SSRF
Server-side request forgery: tricking a server into sending requests to places the attacker can't reach directly.
pull request
A proposed change to a codebase, reviewed before it is merged.
Sources