The Era of Software Quality, or the Era of Ostriches?

(blogs.gnome.org)

60 points | by birdculture 5 hours ago ago

41 comments

  • asjq178 an hour ago ago

    The era of software quality has always been there like in projects like postfix. Gnome could reduce features and increase testing and auditing.

    But let that not waste the opportunity to promote AI.

    • HPsquared 37 minutes ago ago

      AI is well-suited to the tedious work of testing and QA.

      • centuryfall 13 minutes ago ago

        This same statement has been said countless times, replacing “AI” with whatever trend is big at the time. It has also been wrong in every case where that thing is said to replace QA.

      • hungryhobbit 32 minutes ago ago

        "of testing and QA" ... for a certain percentage of "testing and QA". The rest still needs humans.

        • itishappy 13 minutes ago ago

          The article makes a strong point that the slice that needs humans is dwindling quickly. The article seems to suggest our ability to write prose for other humans is our main differentiator.

      • convolvatron 22 minutes ago ago

        just from first principles its really not. given its propensity to just make shit up and cheat, ai is really useful when there is some kind of formal or exhaustive checking as a wall for it to throw crap at. so if you rigorously define success and spend a bunch of tokens, you could easily save time and money. but if you ask the ai 'is this correct' and it says 'yes!', or even 'no!', then you've really learned nothing. this isn't a can you can kick arbitrarily far down the road.

        I think writing tests is a great use of ai, but only if the tests themselves are throughly reviewed or are themselves validated by statements in a formal system.

        • itishappy 10 minutes ago ago

          > given its propensity to just make shit up and cheat

          That's more-or-less how I define QA work. The goal is not proving overall correctness, it's surfacing individual issues.

  • someonebaggy 4 hours ago ago

    Everything coming from GNOME about software quality should be taken with a Strategic Petroleum Reserve of salt.

    While other projects rewrite things in memory-safe ways, GNOME's response is to ask a chatbox if there are any memory vulnerabilities.

    • tom_ 4 hours ago ago

      Say what you like about the random word sequence generation machine, but it does actually seem to be usefully good at generating sequences of random words that correspond to problems in your software. And if you're inclined to write the code by hand, it's probably going to be easier to fix up your existing shit than rewrite it all. (And if you're going to use AI, then you're hardly going to listen to me.)

    • troyvit 4 minutes ago ago

      I'm a KDE guy through and through, don't get me wrong, but statements like this deserve attribution, otherwise it's just FUD.

    • jeremyjh 4 hours ago ago

      The purpose of this is to address a practice common in certain open source ideologies of banning all AI contributions - including vulnerability reports. Basically they are choosing to ignore security vulnerabilities because they had to read too much slop last year.

      • lelanthran 3 hours ago ago

        > Basically they are choosing to ignore security vulnerabilities because they had to read too much slop last year.

        That's not only an uncharitable take, it's also wrong.

        If 999 out of every 1000 "reports" from a specific source is wrong, then it is not irrational to disregard all 1000, especially when they can be generated faster than you can read.

        I mean, it's just probabilities, right? If you're okay trusting output from an LLM, you should be okay with using statistics in general as a source for informing decision-making.

        • jeremyjh 3 hours ago ago

          The TFA is making the assertion that most vulnerability reports by AI in 2026 are valid. They reference the fact that the curl maintainer agrees. That has been my experience as well. Are you arguing something different?

          • lelanthran 2 hours ago ago

            > The TFA is making the assertion that most vulnerability reports by AI in 2026 are valid. They reference the fact that the curl maintainer agrees. That has been my experience as well. Are you arguing something different?

            Right, but that assertion does not contradict what I said: there's a difference between the articles premise (AI reports are mostly valid) and what I said (3rd-party submitted AI-reports are mostly invalid).

            See my reply to a sibling poster who also implies I did not read the article.

        • saghm 2 hours ago ago

          I think the disagreement here is whether this constitutes a "specific source" or not. I've seen some people produce things with incredibly quality and others produce useless slop all with the same AI tools and models, so defining that as one single source doesn't seem like a very good way of viewing things.

          • lelanthran an hour ago ago

            Lets assume your argument is valid, and further assume hat only the high-skill people submit AI-reports.

            Out of, say, 40 bugs that SOTA models can find, you're still going to have to sift through all the hopeful wannabes who each submit that same list of 40, but differently worded, differently explained and with different PoC code.

            The problem still remains when welcoming AI-reports from the world: you could potentially spend all your time on examining and then discarding these reports without even getting to any new bugs in those reports.

            • saghm an hour ago ago

              You could make the same argument for rejecting all reports from third-parties in a pre-AI world though. It just seems like you're picking an arbitrary property to extrapolate a trend from when there are plenty of other similarly arbitrary properties you could extrapolate similar trends from. If I noticed that most bug reports that come on Tuesdays are low quality, I don't think it would make sense to have any bugs that get reported on Tuesdays get auto-closed, regardless of the statistical trend.

              • wonnage 35 minutes ago ago

                Projects were already doing this by making you jump through hoops to submit an issue. It filters out the low effort chaff. I think asking for people (instead of their agents) to submit reports is on the same level.

        • ethersteeds 3 hours ago ago

          If you won't read TFA, source being the RedHat employee paid to triage Gnome vulnerability reports since 2020:

          > Have you heard that most AI bug reports are “slop?” Not so in 2026. That was true for most of 2025, but the quality of AI-generated vulnerability reports has drastically improved. That is not to say that we no longer have problems with bad vulnerability reports, but in general, nowadays most of them are pretty good. (Daniel Stenberg reports the same pattern for curl.)

          • lelanthran 2 hours ago ago

            > If you won't read TFA, source being the RedHat employee paid to triage Gnome vulnerability reports since 2020:

            I read the article very carefully, including the bit that you quoted. Here's what I read:

            > Red Hat’s scan of GLib found 118 vulnerabilities. Or at least, it claimed to. However, due to the way we ran the scans, several of these are actually unnecessary duplicates of each other, which we have not fully deduplicated yet, so the number I report is not entirely trustworthy. Moreover, 46 of these “vulnerabilities” are bugs in gobject-introspection, mostly in the typelib support, which is evidently not very robust. A typelib controls how your program calls libraries; it is effectively calling convention, so it must inherently be fully trusted: a malicious typelib would be able to induce vulnerabilities even without any bugs! I would expect an AI ought to have been able to figure that out, but apparently not. These bugs are still real problems that we ought to fix, but all maintainers agree they are not security vulnerabilities, so let’s count all of them as false positives. That alone creates a 40% false positive rate. Ouch.

            And that's with them running the scanner, not with submitted reports by 3rd-parties! When you welcome AI reports, everybody is going to submit the same report, just differently ordered and differently worded.

            I mean, he even said:

            > I requested that the bug bounty program end because I was overwhelmed with incoming AI-generated issue reports. The final issue was reported on February 23, 2026. Here are the results:

            Sure, he attributes it to a financial incentive, but it's clear that submitted AI reports will overwhelm, and the only way they got to a measly 40% real-bugs was by do the scanning themselves.

            (Also, I wish all these sibling posters implying that I did not very carefully and thoroughly read the article would, themselves, read the article!)

            • qarl an hour ago ago

              Yes. They will need to use AI to process the increased load of AI reports.

              I hope you understand that software engineering is now going to require the use of AI. In the same way that software engineering requires the use of compilers. Sure, some people refuse... but...

              • bigstrat2003 20 minutes ago ago

                "use the tool to fix issues created by using the tool" is not a valid solution. The solution is to stop using bad tools.

                > I hope you understand that software engineering is now going to require the use of AI. In the same way that software engineering requires the use of compilers.

                When LLMs actually can reliably do their jobs (which compilers do), then they might be an essential tool. Not before. For now, they are slop machines used by people who care more about going fast than getting things correct.

                • qarl 16 minutes ago ago

                  > When LLMs actually can reliably do their jobs

                  LLM agents do a fantastic job of finding exploitable bugs in code. MUCH better than humans.

                  They would also do a fantastic job isolating the duplicate reports as described above.

                  So what's your issue?

        • boxed 3 hours ago ago

          You are out of touch unfortunately. Your logic was correct last year, but things have changed radically and super fast. You need to re-evaluate, and this article by a core GNOME developer specifically doing security work should have made you do that re-evaluation!

          • juped 3 hours ago ago

            The last 1000 times someone has said "nah bro AI was bad last year but this year it's good trust" have been wrong, I am also comfortable being informed by evidence.

            • qarl 3 hours ago ago

              > I am also comfortable being informed by evidence.

              Then look. If you can't judge, then trust the experts. This article is written by experts.

              • someonebaggy 2 hours ago ago

                They did this the last 999 times and always came to the same conclusion: that the AI promoters were talking nonsense. At what point should you stop listening to people who only talk nonsense so far, to avoid getting DoS attacked? Must the villagers look for the wolf every time the boy cries?

                • qarl an hour ago ago

                  No. But when the wolf experts announce that there are a dangerous number of wolves - and you ignore it - that's a problem.

                  The people writing this article are experts. They cite other experts.

                  If you can cite real data from the last few months that still claims there are no wolves - and it's not just insane anti-wolf propaganda - I'd love for you to show me.

                  • someonebaggy an hour ago ago

                    What if the last 999 times wolf experts announced there were a dangerous number of wolves, no wolves were found?

                    • qarl an hour ago ago

                      You're going to need to leave the hyperbole behind and use actual facts if you want to continue.

                      Cite data from the last few months.

                      • someonebaggy 4 minutes ago ago

                        But you get to use hyperbole to defend your side? No, I'm not interested in fighting Brandolini's Law.

                  • lelanthran an hour ago ago

                    > The people writing this article are experts. They cite other experts.

                    "Economists have predicted 18 of the last 2 recessions".

                    I mean, c'mon! You have never read that?

                    Besides, when "expert in $FOO" means "familiar with $FOO that's only 6 months old", then it's not unreasonable to be skeptical.

                    In other fields, an expert is someone who's studied the specific field $FOO for decades. Here we're talking about a skill level that is not distinguishable between "1 weeks experience" and "two years experience".

                    • qarl an hour ago ago

                      Your argument boils down to "I'm not trusting that bridge! Have you seen how many mistakes astrologers make!!"

                      I've found that software engineers are good at analyzing the public reports they receive for their own projects.

                      Let's not over generalize, shall we?

  • nottorp 3 hours ago ago

    Hey does that mean they'll use "AI" to allow users to customize the desktop environment again?

    • WD-42 an hour ago ago

      You can do this by vibe coding extensions, and it appears many people are doing this. Ironically this probably makes GNOME the easiest DE to customize at this point.

      • type0 28 minutes ago ago

        > GNOME the easiest DE to customize at this point.

        It's not even false, you do know that KDE exists

    • someonebaggy 2 hours ago ago

      No - non-customizability is a design goal of the GNOME desktop, since they want the experience to be identical for everyone.

  • Sevii 2 hours ago ago

    They should be using an automated AI agent to validate vulnerability reports. Have it pull up the code base, confirm the bug exists, try to reproduce, then update the ticket. You might even run another AI agent pass to clean up the text before humans look at it. AI writing quality improves a lot with multiple passes.

    • saghm 2 hours ago ago

      "Vibe reviewing" is an underrated way of dealing with vibe coded PRs. There's going to be a lot of low-quality noise that's probably better to just close without a human looking at it, but it would be unfortunate to ignore the stuff that might be worthwhile because of that.

    • jayd16 2 hours ago ago

      Sounds really expensive as far as the work you're willing to let the public invoke on your project.

    • tonyedgecombe 2 hours ago ago

      > AI writing quality improves a lot with multiple passes.

      Or you end up with an AI version of Chinese whispers.