I have previously been asked what I find interesting about cyber/appsec or what drew me to it. My answer: no one knows what they’re doing. That is not to say that practitioners in the field are broadly incompetent. Rather, I mean to say that the field is underdeveloped in such a way that everyone is still figuring out what it is we should be doing. I see this as an opportunity because there is substantial room to solve novel problems and have an impact. Compared to more mature industries, it takes little time and effort to make a meaningful contribution.
Unfortunately, there is a competing trend in the industry: no one cares what they’re doing. I can often excuse ignorance; many people lack the opportunity for effective guidance, training, and feedback and they’re doing the best with what they know. However, the seemingly growing disinterest in doing good or worthwhile work is not just troubling, but it has had a substantial impact on my ability to construct this newsletter.
Because cyber is so immature, there are no great channels for continuing education for experienced practitioners in niche areas like application security. Practitioners adapt by individually aggregating various resources in different ways, including social media, blogs, industry reports/events, and newsletters. Newsletters can be great because you can offload the effort of identifying (determining what is important) and analyzing (determining what you need to know/do) new information, relying on a seasoned professional who works in your specialization.
I was originally motivated to start this newsletter because the existing newsletters I subscribed to began to fail in two ways. First, they began to expand the scope of the material, increasingly drifting into non-appsec and non-cyber topics. This obviously negates one of the benefits of the form, lowering the important signal-to-noise ratio. Second, the analysis deteriorated. More recently, it seems many newsletter authors are barely authors at all, instead relying on LLMs to summarize and construct opinions for them. What then is the point of the newsletter?
So I made my own. This is it. But lately I have found the effort of keeping on top of appsec content to be increasing laborious. Is it because there is now such a rapid pace of LLM-driven developments that humans can no longer keep up? No. It’s because the quantity of contributions people are making has increased substantially, but the quality has decreased as well.
LLMs are responsible, of course. Practitioners no longer have to write to publish. They no longer must code to build. These were technical barriers, but they also meant that you had to be thoughtful and selective about what you produce. It turns out this barrier to entry was an important filter that helped (at least to some degree) select for worthwhile output.
I am interested in the problem of automatically distinguishing human output from LLM output (often to save myself from reading slop), so when it came across my feed, I was excited to read Why AI-Text Detectors Disagree About Who Cheated. Well it was one of the worst technical write-ups I have ever in my life read. Why? It was obviously LLM generated, but the writing is not just bad because it contains the common LLM-y writing patterns; the structure, the explanations, and the logical process are all a mess. I think a great exercise for an aspiring academic would be to rewrite the page in a way that is intended for a human audience to both understand and care. The creators of this tool would do as much if they cared enough about the quality of the output.
This is just one example of the slop that I now encounter day to day. The time and effort that normally would go into research, writing, and coding is now offloaded to LLMs, and then to whoever wishes to assess whether what the LLMs spit out is worthwhile (me). Naturally, people are being rewarded for the slopification but not the human analysis. Combined with the existing industry incentives to build your own personal portfolio of work (as opposed to contributing to a greater collaborative project), the result is bad.
Interestingly, it does appear possible to produce writing using LLMs that is indistinguishable from human writing (at least in certain domains to certain audiences - the stories from the study are available for download and IMO there are clear signs of LLM writing). Similarly, it is possible to use LLMs to develop high quality software. Therefore, where I encounter barely legible LLM writing or nonsense project implementations, I use these as strong signals that whoever was behind the LLM really did not care about the quality of the output. More and more, it seems people just don’t care.
Well, I care. And if you are reading this, I hope you care too. I am not writing this for someone to AI it into a podcast and then for someone else to summarize the AI podcast. I am writing for human readers with the hope that some will enjoy and learn, and others will get mad, leading to self reflection and ultimately personal growth.
That’s not a threat.
It’s a commitment.
General Application Security
Prompt Injection
I have largely avoided writing about “prompt injection” or any of the many vulnerabilities that plague LLM-integrated applications and systems. Quite frankly, I find them uninteresting (with some exceptions). If you provide an LLM with the capability to do something and an attacker can influence the input to the LLM, then the attacker can do that thing. This is an unsolved problem and I don’t think anyone is surprised that these issues keep being reported.
A popular model for understanding risky integrations of LLMs is Simon Willison’s lethal trifecta, but I think even this model imposes too many requirements. As we have learned from the OpenAI/Hugging Face incident (and others), security incidents can occur without access to private data, the ability to externally communicate, or exposure to untrusted content as long as a vulnerability or misconfiguration exists. Technically, you could say that access to private data and the ability to externally communicate were ultimately present and the LLM was insufficiently sandboxed, but risk from LLM integration does exist without exposure to untrusted content.
These issues will continue to be identified and published and I will continue to be uninterested.
uBlock Origin Gives up on Facebook
I am always fascinated at the technological arms races that are occurring on small scales with relatively low stakes. One of those is the battle between ad blockers trying to block ads and ad delivery platforms trying to get around ad blockers. Most recently, uBlock Origin has given up on supporting Facebook, which is not unsurprising since you have an open source project battling with one of the biggest advertisers.
Our Industry Is Deeply Embarrassing
As a follow up to the opening screed of this newsletter, I have to admit that I am increasingly embarrassed by apparent lowering standards and lack of professionalism in our industry.
Most recently, I completed a round of reviews for the next OWASP Global AppSec event. I have always enjoyed being a reviewer; you can see novel work early and you have the opportunity to steer the direction of an event based on your expertise. Lately, the process is painful. In addition to the consistent difficulty prospective presenters have in following basic requirements (like anonymizing their submission), there is an increasing proportion of slop submissions. And they’re bad.
While you shouldn’t rely on LLMs to write your submissions, I do understand the draw for people who are not effective writers, or submitting in a secondary language. That said, these were not bad submissions simply because they were LLM generated or LLM-assisted; they were bad because they were low effort, which was enabled by LLMs. Either way, if you are making me read your submission with the hopes of attending the conference, you should at least make an effort to write it.
OWASP conferences are not the only target of slop submissions. DEF CON 34 recently published presentation material from last week’s event. Despite my consistent criticism of how other people use LLMs, I do work and build with them regularly to augment offensive security work, so naturally the talk “Taming the Swarm Hard Architectural Lessons from Building a Deterministic Agentic Web Pentesting System” had me interested. Too bad it’s slop.
You don’t need to be an LLM writing bloodhound as I am to immediately identify the presentation material as heavily AI written and influenced.

“It is a horoscope” - this doesn’t even make sense.
Now, this content could still be worthwhile. There could be valuable technical work. The presentation itself (I have not watched the recording yet) could be interesting and engaging. OK. But for me, I would be deeply embarrassed to include this type of obvious LLM drivel. So from my perspective, the author either does not care, or does not recognize it. In fairness, the author appears to not be English-native, so perhaps the signal is not as strong for non-native speakers.
But “music is the universal language of mankind” and the obviously AI generated music that accompanies the demo video with this presentation is one of the worst things I have ever listened to (not to mention visuals of the tachycardic DONATE button). But, OK, maybe writing and music taste are not prerequisites for producing outstanding technical work for offensive security testing. So even with these concerning signals, I still felt compelled to evaluate the work (again: see opening screed).
Perhaps I need to introspect and reckon with my compulsions. Almost every time I open a fresh write-up or tool release with an earnest effort to learn about some new approach to improve security testing using LLMs, I am disappointed. Even well-known industry veterans apparently do not possess the same sense of shame that I do and are more than willing to publish absolute slop write-ups of what they are doing.
OK. I took a break. I watched the DEFCON presentation. Is it good? Is it worth your limited time? No. No it is not. And look: I don’t enjoy being critical of people who are building things in an effort to advance the field. The intent of this newsletter is not to find flaws in every project. I am legitimately spending my time trying to learn how people are moving the state of the art in security testing and when presenters get up on stage with a clearly LLM generated project and a clearly LLM generated slide deck and they aren’t even presenting novel ideas, what respect is due when I am investing my time in this? I will give kudos to the presenter for playing a video demo live, blasting their AI song over the speakers. The song is awful, but the A/V worked, which is a rare conference W.
What did I learn about the framework from the presentation? Almost nothing, so let’s look at it on GitHub. My first thought: there is so much LLM generated code and content here, who would feel comfortable deploying this without a substantial review effort? Even the documentation is LLM generated, which - in my opinion - is a very important piece to have a human touch (and preferable only a human). Basically, one guy vibe coded this.
So how does it work? It uses agents to run tools. Groundbreaking. It would take us far too much time to review all the code, so I am going to share a strategy I have started employing when examining these types of “agentic” pentesting solutions. Choose a single vulnerability class or test technique where you are knowledgeable. Then, search across the repo to identify the relevant logic for these test cases.
You will often find that these solutions make an effort to reinvent tooling for each domain. As a result, it’s usually quite bad. Cross-site scripting is a great area to examine because “xss” is a short string that typically will uniquely match relevant functionality. I used this technique when evaluating the somewhat popular T3MP3ST tool (side note: for this repo and others that embed payloads, it’s even easier to find relevant logic by searching common payloads). Take a look at the code behind its XSS tests:
{
name: 'xss_scan',
description: 'Test for XSS vulnerabilities (real requests with payload reflection check)',
category: 'vuln',
parameters: [
{ name: 'url', type: 'string', description: 'URL with parameter to test', required: true },
{ name: 'param', type: 'string', description: 'Parameter name to test', required: true },
],
handler: async (context) => {
const baseUrl = context.parameters.url as string;
const param = context.parameters.param as string;
const payloads = [
{ payload: '<script>alert(1)</script>', name: 'Basic script tag' },
{ payload: '<img src=x onerror=alert(1)>', name: 'IMG onerror' },
{ payload: '"><svg onload=alert(1)>', name: 'SVG onload breakout' },
{ payload: "'-alert(1)-'", name: 'JS string breakout' },
{ payload: '<body onload=alert(1)>', name: 'Body onload' },
{ payload: '{{constructor.constructor("alert(1)")()}}', name: 'Template injection' },
];
const results: { payload: string; name: string; reflected: boolean; encoded: boolean }[] = [];
const vulnerable: string[] = [];
for (const { payload, name } of payloads) {
try {
// Build URL with payload
const testUrl = new URL(baseUrl);
testUrl.searchParams.set(param, payload);
const response = await targetFetch(testUrl.toString(), {
signal: AbortSignal.timeout(5000),
});
const body = await response.text();
// Check if payload is reflected
const reflected = body.includes(payload);
const encoded = body.includes(encodeURIComponent(payload)) ||
body.includes(payload.replace(/</g, '<').replace(/>/g, '>'));
results.push({ payload, name, reflected, encoded });
if (reflected) {
vulnerable.push(name);
}
} catch {
results.push({ payload, name, reflected: false, encoded: false });
}
}
const output = results.map(r => {
const status = r.reflected ? '⚠️ REFLECTED (potentially vulnerable)' :
r.encoded ? '✓ Encoded/Filtered' : '✓ Not reflected';
return `[${r.name}] ${status}`;
}).join('\n');
return {
success: true,
output: `XSS scan on ${baseUrl} (param: ${param}):\n${output}`,
findings: vulnerable.length > 0 ? [{
title: 'Potential XSS Vulnerability',
severity: 'high',
details: `Parameter "${param}" reflects unencoded payloads: ${vulnerable.join(', ')}. Manual verification required.`,
}] : undefined,
};
},
},I will let you, brilliant reader, be the judge of that (no I won’t: it’s bad!).
Returning to our repo of interest (BugTraceAI), there is a dedicated agent directory containing a number of Python scripts. This includes a subfolder for XSS. It’s a mess. There is substantial hard-coded logic and a lack of structure that lends to modular extension. Let’s look at a specific script as an example: parameter discovery. Here are just some basic design issues from a quick glance:
In addition to being a very limited (presumably LLM generated) list, the parameters are hard-coded and embedded as Python constants. They are not externalized, versioned, organized into categories, tagged by technology, and so on…
From a general design perspective, IMO it would make more sense to have a generic parameter discovery capability coupled with whatever manages surface area. There will ultimately be duplication of functionality.
The actual detection logic is very limited. Existing tooling (like extensions in Burp) provides substantially more advanced capabilities compared to this.
There is a lack of configuration. You get the shallow hard-coded tests.
I want to focus on that last part. If human testers make an effort to discover HTTP parameters across an application, they should outperform this agentic platform. This is especially the case for any variation in the application’s interface that the script is not capable of adapting to. So for in-depth penetration tests, my question is this: what is the point of this tool? If it cannot be relied on to even meet the efficacy of human testers, it is of little use to pentesters, who will have to conduct these tests anyways.
Maybe some of the other tools within this platform provide effective automation/capabilities, but I don’t have high hopes. The presenter did not conduct (or at least did not share) any form of experiment/analysis/comparison. They vibed up a tool, their presentation, and probably their submission to DEFCON if I had to guess. Ultimately, this tool is not novel and is certainly less capable compared to existing commercial tools and probably also some other agentic open source platforms. Hell, I bet Claude Code running with the prompt “You are THE GENERAL” and a skill to use a web browser could outperform this harness.
Burp Agentic Testing (AT) Is Here
PortSwigger has released their latest iteration of LLM-integrated testing into Burp Suite. But who is this for? Just like their previous disappointing integration of LLMs into Burp, the system forces users into using PortSwigger-controlled LLM services using PortSwigger-controlled tokens. After evaluating requirements for some of our key clients for integrating LLMs into the testing lifecycle for penetration tests and code assessment, I can conclude that the PortSwigger solution is non-viable for a number of reasons. At the same time, bug bounty hunters appear to largely be migrating to Caido, which already had LLM-integrations available and appears to much more effectively suit testers who have been driving their process with existing harnesses like Claude Code. So who is Burp AT for?
I do intend to evaluate the capabilities, but here are my roadblocks:
I have to dedicate personal time to do so. I cannot simply evaluate these capabilities on a client engagement when the platform necessitates that I send test data to PortSwigger.
Even though PortSwigger initially provided some tokens to accounts for using their AI service, these stupid tokens expire in a year, so I no longer have any tokens banked to experiment with. This isn’t a huge block (it’s easy enough to purchase some), but it’s annoying.
Please don’t get the impression that I am not rooting for PortSwigger. My life would be substantially easier if they built a top tier LLM-assisted testing solution that allowed Pentesters to stay competitive. From everything I have seen from the design and initial results (or lack thereof), they did not. As a result, I have to do what ever other asshole in this forsaken industry is doing: build my own AI tooling and brand myself as a Offensive AI Engineer (no, sorry, I am hearing now that this is now “Offensive AI Architect”).
Be Careful Not To Extrapolate
LLMs are kind of an incredible technology. Like, considering how they are constructed, it is impressive that they can accomplish the things that they can. That said, I do think the right mental framing of the technology takes some of the magic away.
LLMs are built on the largest compressed sets of digital information. They use this information to probabilistically predict (or continue) tokens in a sequence. Yes, we already know all of this. My conceptualization is this: LLMs are effectively like fuzzers. Instead of mutating simple patterns, LLMs are complex enough to generate human language and code that are syntactically and (mostly) semantically valid (or at least semantically coherent).
This is why concepts like Loop Engineering have taken off. If you can provide the right preconditions (context!) and verifiers, LLMs can produce output probabilistically on a loop until your success criteria are met. This is true for least for certain classes of problems, but we have also had difficulty in understanding how to classify problems and how to determine which ones are in the domain of LLMs.
Even so, I often come across LLM use cases that I find baffling. Here is an example from Soroush Dalili that is part of a skill for evaluating web security research:
This seems unlikely to succeed, no? To me, the challenge of identifying worthwhile techniques and research across the listed categories feels entirely outside of the capabilities of LLMs to effectively do and I can think of many ways where this would go wrong based on my experience with LLMs.
Soroush pretty consistently produces interesting work, so I’ll probably take a deeper look at this in the future when I continue to work on appsec curriculum content, but my instinct is that this is not the way.
Don’t Install Zoom
Every time I hear about a major Zoom vulnerability, I remember that people are actively installing it on their Windows systems. Don’t do this. Windows is insecure. Their app is insecure. Browser at least offer some additional sandboxing. Just run it there.
Yes, Windows is Insecure
One of my favorite stories is actually using a Razor mouse to gain admin access to a Windows system (a point-of-sale machine). Well it seems like core issue in Windows remained. Are you surprised?
Jailbreaks Are Back
Many said it would not happen, but iOS 26 has a jailbreak via Dopamine 3.0. I have not yet looked into the implications for security testing with physical devices.
Burp And Server-Sent Events (SSE)
If you’re working with Burp and a web service that uses SSE (increasingly popular for interactive chat-based applications that integrate LLMs), you should know that there are at least two bugs due to the way Burp handles it:
You cannot send request/response pairs to organizer.
You cannot use Intruder.
While it appears someone built an extension to overcome these limitations, I have found it to be quite brittle against different implementations. So what did I do? I vibe coded a simple web interface tailored to a target’s SSE implementation that let me interact with the service and automate the sending of payloads.
XBOW Vs. Akido
This feels like ages ago now, but Doyensec conducted an analysis of two leading fully automated AI security testing tools. I think the most interesting part is seeing what these tools offer as far as capabilities for configuring scans. The options are limited yet there are so many potential issues that these tools could encounter.
Even Security Companies Don’t Know What They Are Doing
Improved Acoustic Side Channel
You have probably heard of the technique where audio from keyboard typing can be used to recreate the data typed. The technique was recently improved.
Handy Hackvertor Feature
I doubt any of us will ever recall this feature when a use cases arises, but it seems nifty.
Vuln Summer
There certainly are classes of vulnerabilities that LLMs are attuned towards identifying. A small sample of vulnerabilities identified (or exploited) with LLM assistance:
Some major WordPress vulnerabilities
Some crypto issues
And this is just a small sample I have had time to read…
Connect
Respond to this email to reach me directly.
Connect with me on LinkedIn.
Follow my YouTube.
RSS feed here.




