Mythos Detected 23,000 Vulnerabilities Across 1,000 OSS Projects 74
wiredmikey shares a report from SecurityWeek: Anthropic says its Claude Mythos model discovered thousands of severe vulnerabilities across more than 1,000 open source software (OSS) projects. According to the AI giant, Mythos Preview has identified more than 23,000 potential vulnerabilities. Of these, 1,900 have been reviewed by external security firms, and 1,726 have been confirmed, including over 1,000 rated "high" or "critical" severity.
The findings are still being reviewed, and Anthropic estimates that nearly 3,900 critical and high-severity vulnerabilities will be confirmed based only on current findings. As the scans are ongoing, the company believes the number of severe vulnerabilities may reach 6,200. Anthropic says more than 1,100 unverified findings have been reported to vendors, and 75 issues with a critical or high severity rating have been patched. Vendors have published 65 security advisories. "The number of patches is still relatively low for three reasons. First, we're still early in the 90-day window that's set out in our Coordinated Vulnerability Disclosure policy: we expect many more patches to land soon," the AI company explained.
"Second, we are likely to be undercounting patches because some vulnerabilities are patched without a public advisory: in those cases, we're reliant on scanning for the patches ourselves using Claude. Third, the low volume of patches reflects a genuine problem: even at our relatively slow pace of disclosures, Mythos Preview is adding to an already-overloaded security ecosystem," it added.
The findings are still being reviewed, and Anthropic estimates that nearly 3,900 critical and high-severity vulnerabilities will be confirmed based only on current findings. As the scans are ongoing, the company believes the number of severe vulnerabilities may reach 6,200. Anthropic says more than 1,100 unverified findings have been reported to vendors, and 75 issues with a critical or high severity rating have been patched. Vendors have published 65 security advisories. "The number of patches is still relatively low for three reasons. First, we're still early in the 90-day window that's set out in our Coordinated Vulnerability Disclosure policy: we expect many more patches to land soon," the AI company explained.
"Second, we are likely to be undercounting patches because some vulnerabilities are patched without a public advisory: in those cases, we're reliant on scanning for the patches ourselves using Claude. Third, the low volume of patches reflects a genuine problem: even at our relatively slow pace of disclosures, Mythos Preview is adding to an already-overloaded security ecosystem," it added.
Caveat... (Score:5, Interesting)
already-overloaded security ecosystem
This is true, but in part because a lot of 'security' reports are pretty bogus, even if they get CVEs and 'security researchers' call it a vulnerability, others may be inclined to roll their eyes. For example, the curl project had a write up:
https://daniel.haxx.se/blog/20... [daniel.haxx.se]
So LLM findings I anticipate to be similar, but just a firehose of stuff to dig through to separate the real findings from the innocuous ones.
We likely will never have a grip on that, as it's generally easiest to patch the report and not think about whether it *really* was a security risk. The patch may confirm incorrect behavior being acknowledged, but not whether it was realistically a 'security' risk or not.
Re: (Score:3, Insightful)
Hence the need for PoC exploits (Score:1)
Re: (Score:2, Insightful)
There are too many idiots deep in denial as to what LLMs actually can do. Hence your completely accurate comment gets voted down to -1.
I also love how these vulnerabilities are both "severe" and "potential" in the story. There is no honesty in these people, just greed.
FOSS developer retirements expected (Score:2)
Speculating that lots of FOSS project maintainers will retire from their project based on maintainer age and other demographics.
What should come out this is what code structure, such as a 10 +/- line window around the bug, is common so that compiler, syntax checkers, code suggestion tools, etc. can recommend, compiler warn or fail to compile those language patterns.
And then again, some languages, like C / C++ can be changed that a free() or dispose call set the variable to Null instead of letting it refere
Re:Caveat... (Score:5, Informative)
Yep. CVE slop. Reported as 23'000 "severe" vulnerabilities, but if you look closer it is really just "potential" ones. Lies, lies and more lies. LLM pushers at work. And of course, the security companies that verified some of these would never overstate things.
Re: (Score:1)
Please waste more mod-points here...
Re: (Score:2)
It's a concept called defense in depth, and perhaps also defensive programming. It's good practice. You do not want to hold things off at the gate exclusively, because that relies entirely on your gate defense. This shouldn't be a difficult concept to understand.
Yes, it's potentially more difficult to exploit, but if it's known, a clever exploit can still be fashioned to expose it. This is being seen increasingly with AI driven exploits. You don't need a kernel RCE to gain full system access - you need 3 or
It's a crock of shit like their "acc compiler" (Score:5, Insightful)
I don't doubt that they've found some bugs, but they keep claiming thousands. Extraordinary claims require extraordinary evidence and they simply haven't produced every 1/10th of the evidence they need to substantiate their claims.
Show your cards or SHUT THE FUCK UP, Anthropic. You guys are annoying and I suspect you are giant liars, too. You sure as fuck lied about the C compiler. It was broke as fuck with no assembler and a useless non-working linker.
Re:It's a crock of shit like their "acc compiler" (Score:5, Insightful)
Reminds me of those recent "271 zero days in Firefox!". But in the patch notes, it was 3, two of them "use after free" that just shows shoddy engineering and not using techniques that prevent these and not using the tools that find these without any use of "AI" at all. And only one of these rated "high". At the same time there were about 20 reports from people. Hence Mythos found a small part of the problems. Also note that people will keep finding more while Mythos is probably done after one go.
The LLM pushers are shamelessly lying. Looking at their abysmal business numbers, it becomes clear why they do. They need to improve revenue by about a factor of 10x very soon or they will just go bankrupt.
Re: (Score:3)
CVE-wise
Most bugs do not get CVEs. That's like saying "lets judge how many people in the population are speeding by counting only the fines issued by the police".
To be clear your premise is correct. Most of these bugs won't be critical. But I'd stop short of gaslighting them simply because you couldn't count the CVEs. Also this isn't Anthropic claiming Mythos discovered and reported them, this is Anthropic claiming security researchers looked at Mythos's output and confirmed they were in fact real bugs.
Re: (Score:1)
There is no point in claiming 23,000 bugs were discovered when in fact only 1000 were rated important (by who?). That's a 4% discovery rate across all bugs for starters.
Now according to the curl project maintainer (who at least we can trust as he has skin in the game), only 1 real vulnerability was found in his project. Extrapolating across 1000 OSS projects as stated in the summary, we get 1 important bug per project on average, but over 23,000 bugs total claimed. That's an abysmal 22 false positives and
Re: (Score:2)
No one claimed what you said. The claim is that Mythos identified 23000 potential vulnerabilities. The summary goes to great lengths to explain that 23000 is not the actual number of bugs.
Now according to the curl project maintainer
One project with one coding style provides data that can only be used to conclude how Mythos works on their code and theirs alone. Even then your entire evidence is wrong as it's stuck in the wrong time. Mythos didn't even exist when the Curl developer made that statement back mid last year.
You're basically saying EVs don'
Re: (Score:2)
That is trivial. I can easily find 50000 potential vulnerabilities in any project whose code base is at least 50000 lines of code. It's junk reporting, which is why I pointed to the recall metric. Anyone in AI knows this stuff, it is basic.
You're wrong [daniel.haxx.se]
Re:It's a crock of shit like their "acc compiler" (Score:5, Interesting)
Given that the whole topic is *security* bugs then counting CVEs seems extremely apt, and is still grading on easy, as CVEs are actually a pretty low bar. E.g. a 'medium' CVE for ncurses exists (CVE-2023-50495). The tic compiler can segfault with malicious input. Fine, a bug, but... what is the security angle? It has a CVE despite not being a vulnerability. Then you have scenarios where someone finds a component with a bug and describes 6 different ways of making the bug misbehave and get 6 distinct CVEs for what is a common code fix. Example that comes to mind is that VIM had a bunch around it's script interpreter where malicious scripts can run arbitrary code (which is not pitched as a sandboxed environment and explicitly allows arbitrary commands already).
Also getting security researchers to agree something is a security problem is similarly easy. I have an 'advisory' here that tried and failed to get a CVE but a security company granted it a special advisory. Digging in the issue is, under certain circmstance, a person trying to make it misbehave actually *fails* to get permission to something they should have had permission to... Denying service *only* to the attacker... The deeper analysis shows this is the *only* way it could misbehave, it could only fail to acquire privilege under deliberate abuse. There's zero appetite in the industry for pushing back against pretty dumb "security" findings, so they err on the side of accepting everything could maybe be a security issue if someone says it could be.
The issue is an AI twist on a long standing problem particularly in the security industry: people standing on counts of CVEs and handing them out like candy resulting in 'vulnerability slop'. There are real issues out there, but you say 'advisory' or 'cve' and I'm not inclined to think one way or another until I look deeper, and I can't afford to look deeper into the sea of CVEs I already have to wade through.
Re: (Score:2)
Given that the whole topic is *security* bugs then counting CVEs seems extremely apt
Not really. CVEs are 100% optional and manually generated without oversight from any authority. Not all bugs (even security bugs) get CVEs, they explicitly rely on someone participating in the process of creating a CVE. The vast majority of bugs in the world (even security ones) simply get patched, some times they even get a line written in the release notes, other times they are one of many that are lumped together with "* Fixed bugs"
Re: (Score:2)
Great point. It's also really unclear how many Mythos found that are not/would not be found by open weight models.
This is pure marketing FUD, which is one of Anthropic's most advanced forms of "intelligence".
Re: (Score:3, Informative)
Also funny how apparently the "severe" vulnerabilities are really only "potential" vulnerabilities. No lies here, nope.
Re: (Score:1, Flamebait)
And the cult-like LLM believers moderating down, again, anything that does not fit their deranged world-view. Pathetic.
Re: (Score:2, Funny)
Please waste more mod-points here....
Re: Caveat... (Score:2)
90 days (Score:5, Funny)
And....go!
---
Yes, it's an average of 23 per project. But still.
Re: (Score:3, Insightful)
Here's 23,000 vulnerabilities that we found ...
Vulnerability. They keep using that word. I do not think it means what they think it means.
A real vulnerability can be exploited. Flaws in code, may be vulnerabilities, or they may not be. Until they are proven exploitable, they are lower priority. Still worth fixing... as time allows.
Re: (Score:2)
Re: (Score:2)
Problem is that ship sailed *years* ago, claim a bug is a 'vulnerability' and you'll probably get a CVE regardless of merit.
Been there, done that, in two weeks (Score:2)
Here's 23,000 vulnerabilities that we found, and we're giving you 90 days to fix them.
Not a problem. I once lint'ed a major software product once the code was locked down as release approached, only bug fix changes were allowed. I spent two week reading through lint output. Three honest to god crashing bugs were found. 90 days sounds quite generous. :-)
Re: (Score:3)
These guys...SMH. Lint missed the bugs ...
You misunderstand. Its not about the tool, its about going through a long list of reported source code problems that has an extremely high mount of false reports.
Re: (Score:2)
Re: (Score:1)
Indeed. By the recent CURL story, something like 80% false positives. Probably worse in general. Such a tool has a term for it: useless. This is just more AI slop.
Re:Been there, done that, in two weeks (Score:4, Insightful)
I'm sure attackers would be fine taking the time to sift through 20 bugs to exploit the one critical one.
cURL is probably a bad example in general since it's relatively small and has had considerable development to review it
Re: (Score:2)
A 20% positive rate isn't all that bad for a tool. My experience with static analysis tools is that the number of genuine problems found is about 1% (but often it is still worth the effort sifting through to find such problems).
Of course, most people understandably do not like doing the tedious tasks. There has to be the motivation to do so (as Ksevio points out, the attackers have such motivation).
Re: (Score:2)
The problem with that is that the static analysis tools typically also have a quite low false negative rate. Hence you get more for working through the reports.
I expect that is not true for LLMs.
Re: (Score:2)
It is 23000 issues with 1000 projects, or about 23 per project (obviously there will be some projects with a lot more, some with less).
Some bugs are easy to fix with low risk of regression (e.g. an off by one calculation or a missed null-check), some can be much more challenging and risky (like race conditions). It is feasible to fix the former type of problem in that timeframe.
For what it's worth, I think that finding 3 serious bugs in 2 weeks isn't a bad score (I understand that it was tedious). With a ma
Re: How about... (Score:2)
Crashing bugs are the easier categories of bugs to find and fix. Things like timing attacks and race conditions are more difficult, and can take more than 90 days to find and fix.
Death of security (Score:5, Interesting)
When the pace of bug discovery overwhelms the capacity to patch, and the discovery tools are available to... well, everybody... doing any business online is fraught with peril. You can't even triage trust by the integrity of the company. You might trust that "Valerie's Dog Treats" is legit, but their payment dependancy might be using compromised packages.
How in hell are we going to hold this thing together?
Re: (Score:3, Funny)
How in hell are we going to hold this thing together?
With more duct tape .. same way the internet has been held together for decades now
Re: (Score:2)
Relax it's not a problem (Score:4, Insightful)
Basically this is just another shitty AI hype press release.
Now it's not impossible this will change in 6 months to a year but if AI gets that good then we're going to have bigger problems because we're going to be looking at something like 25 to 40% unemployment. Security is going to be the least of your problems. You're going to have roving bands of bandits within about 6 months. I mean the last thing we're going to do is just give people food and shelter because that would be theft.
About the same problem as using lint (Score:2)
What they're finding is a bunch of theoretical problems that in practice cannot be exploited.
Much like lint output over many decades of use
Re: (Score:2)
Thanks for confirming what I already suspected. This is worthless, or rather negative-worth, slop. It nicely illustrates again that LLMs come with impressive recognition of localized patterns, but absolutely no insight whatsoever.
Obviously, good code will have redundancy in the form of defense-in-depth, minimal privilege, input validation, privilege separation and so on. But this mythically stupid system cannot understand any of that. And hence it flags non-problems and makes us all less secure by wasting d
AI scan becomes part of CI/CD, code commits (Score:2)
When the pace of bug discovery overwhelms the capacity to patch, and the discovery tools are available to... well, everybody... doing any business online is fraught with peril.
The bug discovery process becomes part of the automated Continuous Integration and Continuous Delivery system. The bug discovery process gets run on a module before code is committed. We find the bugs as we create them, they are discovered earlier, their cost is lessened.
There is only a crisis now because were are doing all this scanning retroactively.
Re: (Score:2)
That's it I'm moving back to OpenVMS.
Re: (Score:2)
https://vmssoftware.com/ [vmssoftware.com]
Runs on X86 and virtual.
I run it on my Intel Mac in VirtualBox and UTM (which is easy QEMU) but VMware and KVM also supported.
Retired: less digital, more analog
Re: (Score:2)
How in hell are we going to hold this thing together?
We are not going to. Forst, technological debt just became due for payment. And people cannot pay. Second, this will be a lot of AI slop, i.e. false positives (note how the story uses "severe" and "potential" for the same findings....). This is a gigantic DoS attack on the developers.
I think we need to forbit the LLM pushers to do this form of marketing now. Or things will come crashing down.
Re: (Score:2)
When the pace of bug discovery overwhelms the capacity to patch, and the discovery tools are available to... well, everybody... doing any business online is fraught with peril.
Kinda sounds like the online businesses need to start being financial contributors to ensure they are not relying on flawed software.
Besides, bugs are finite.
Mythos found only one low-severity vulnerability in Curl, with experts debating whether that is a failure of the AI model or a testament to the open source data transfer tool’s maturity.
Re: (Score:2)
The existing software, with people slapping together 3rd party components such that hello-world app ends up being 800MB in size due to dependencies for modularity and future proofing, has been get
Re: (Score:2)
This ideal state sounds good but it is prohibitively expensive for all but deep pockets.
Re: (Score:2)
Re: (Score:2)
How in hell are we going to hold this thing together?
By turning programming into an actual engineering discipline? I dunno. Might be more effective than seat of the pants programming that we encourage now. But wait, yet another language will make it easy to program again.
Lazy and undisciplined. What do you think will happen? Exactly what we are seeing?
Re: Death of security (Score:2)
Alex I'll take "Wromg Priorities" for $600. (Score:2)
...the low volume of patches reflects a genuine problem: even at our relatively slow pace of disclosures...
Yes, and that problem is development time focused on visible, feature-driven work gets more attention than bug-fixing or security coding. This is in corporate profit-driven products as well as open-source projects, too (remember all those Firefox bugs you reported 10 years ago?)
Almost like all these layoffs happening as a result of AI's changes in workplace are a mistake -- also because of AI's changes in the workplace. Why lay off software devs when there is suddenly a bunch of patching needed?
I see this as an indictment of Mythos (Score:4, Informative)
Automation is a great tool for doing repetitive tasks and automation is all these companies have.
It takes the Intelligence part of their AI(Artificial Intelligence) business to actually layout a properly constructed, documented and tested fix that should go along with any submission.
And that is beyond the capabilities of the current Massive Automation Machines they are selling today as AI.
Re: (Score:2)
It takes the Intelligence part of their AI(Artificial Intelligence) business to actually layout a properly constructed, documented and tested fix that should go along with any submission.
And that is beyond the capabilities of the current Massive Automation Machines they are selling today as AI.
Indeed. I have two students currently testing this out in a limited set-up and the proposed fixes are mostly between "crap" and "causes more vulnerabilities".
Dealing with the backlog? (Score:2)
Sounds like they're scanning all the most prominent repos and reporting results. Fine. But this is old code, a backlog that has never been looked at by LLMs before. Say those projects find your reports useful and patch then bugs. The backlog goes to near zero. What then? What are you going to use to impress your investors and sell your services on a regular basis?
Can I run it myself? (Score:3)
So like for CURL? (Score:1)
Where 80% were false positives and it found a whopping one (!) true vulnerability in a large project that must have more? Pretty bad. A human reviewer would get fired for this.
I am sick and tired of this mindless cheerleading.
no more secrets (Score:2)
reminds me of that sneakers chip from the movie. picture it marty, all code is cracked and exploited. all systems are compromised. no more secrets.
Re: (Score:2)
Patch up or shut up (Score:5, Interesting)
They talk a big game, now it's time for trial by fire.
Tick up for how long? (Score:2)
I think the interesting number to keep an eye on is how long does the uptick in reasonable AI generated reports continue?
Like any new tool, will it follow a bell curve where things found drops off after the class of bugs found are addressed?
After all, how many more "new" classes of bugs can it find? That number isn't infinite.
Wash, Rinse, & etc. (Score:2)
Is this just it finding the same 23 bugs in a common dependency in 1k different projects and counting those repeatedly?
bugs or not, might be interesting. (Score:2)
I would not mind getting 23 reports (bugs, vulnerabilities, issues, call them what you will) on my projects on github. They would be interesting even if they are just ai slop.
Translation: They Trained AI on 1000 OSS Projects (Score:2)