AI agents in a stand-up call

(atoll92.github.io)

47 points | by jonbaer 12 hours ago ago

53 comments

  • piterrro 9 hours ago ago

    Take a stepback and understand why standup was invented in the first place. The reason was: information flow bottleneck - with agents information flow should not be a bottleneck because at any given time the agent can grab as much context as they need, including talking to other agents (althought that nit necessary imo since agents should work visibly through issues and PRs)

    • stillpointlab 8 hours ago ago

      One useful part of standup is to elicit connections where one wasn't expected. For example, someone might say "I've been working on X" and another person might say "Oh, I worked on X a couple of days ago and you should know ..."

      It is an opportunity for things people didn't realize they needed to communicate to get stated publicly in a way that can create connections neither side would have initiated.

    • necovek 8 hours ago ago

      That's why people "meet".

      However, a "standup" was invented to solve a particular problem: people, in isolation, do not always understand that they are stuck, and talking it through lets fellow humans identify their challenge.

      The stand up part is mostly to encourage keeping it concise and short, so you only focus on what is actionable between all participants (waiting on a review, I am looking at this for the 3rd day and I think I am ok, anyone wants to brainstorm this...).

      With agents, they get similarly stuck (completely or in a loop), and with each keeping a different context at any given moment, they might still be able to help each other.

      Obviously, their memories are not like humans' (there is no "I touched that two months ago, let's discuss after"), but there might be a different sweet spot where having them sync regularly (every 1M tokens instead of daily?) might be useful as they are also parallelizing work.

    • spacebanana7 9 hours ago ago

      Couldn’t the same be said for skilled humans working together on an effective team?

      • piterrro 4 hours ago ago

        It could, that's why there are teams that don't have a standup and work very well (I've been with both and I much prefer the ones that dont have a standup OR they meet on a standup to chit-chat and keep up the connections rather than discuss the actual work - simply because they have autonomy and can do it async)

      • ranguna 5 hours ago ago

        Unfortunately I cannot peek at the brains of my fellow colleagues to have their context at any given time.

      • faithlv 3 hours ago ago

        Yes it can, see Conway's Law.

      • newsicanuse 8 hours ago ago

        no

      • darkwater 8 hours ago ago

        Yes but good luck finding the skilled humans that do both things - developing and communicating - well. They do exist, but they are not so common.

      • eru 8 hours ago ago

        No, why? Humans need time to get into concentration, and they aren't awake 24/7, and they can get distracted.

  • coubri 10 hours ago ago

    Its unbelievable how much everything changed in less then 5 years.

    Just imagine seeing this in 2021

    • pmg101 8 hours ago ago

      It's great. I remember bemoaning how little things had really changed between 2006 and 2016. Given that in theory software is infinitely malleable I found it depressing how it felt like we were just doing the same things. It's exciting we've finally unblocked progress and are finding amazing new ways to work, although I never would have foreseen this being the way! 2026 sure feels nothing like 2016.

      • eru 8 hours ago ago

        Well, we've unlocked something. How much of it is progress remains to be seen. Definitely interesting times.

        A lot of the AI hype these days reminds me of the infamous pets.com; but, of course, the general ideas of the dot-com boom did come to pass. (And if all ideas people had tried back then had turned out to be sound and successful, that would have been a surefire sign that they weren't trying out enough crazy stuff.)

    • CurleighBraces 8 hours ago ago

      I would argue really it's less than 9 months. At the start of the year then models/AI really weren't capable of shipping production code etc without hand holding...

      Todo's left in code, mock implementations instead of fully working, people doing Ralph loops etc

      It's an entirely different situation now

  • Kwpolska 8 hours ago ago

    > Made by a bored dev and his robots. This dev was replaced by his own agents.

    Maybe if you hadn't outsourced your job to Claude, you wouldn't be bored. And if you're forced to do so by manglement, why would you also outsource this toy project?

    • N_Lens 7 hours ago ago

      Manglement, heheheh. Thanks for that one, I’m stealing it.

  • donkey_brains 2 hours ago ago

    Congratulations, you’ve built the Torment Nexus from the classic sci-fi story, “Don’t Build The Torment Nexus”.

  • cyberrock 9 hours ago ago

    For a long time I've avoided anthropomorphizing agents like this, but it really is hard to ignore the effectiveness of using them with human processes and systems. Are the processes (standup, sprint, etc.) effective because they are using our languages? Or are the processes just universally good?

  • holdupagain 10 hours ago ago

    This is what my life's work is replaced by. I'm way over the initial shock and honestly these little guys are cute.

    • K0balt 10 hours ago ago

      I know right?

      I’m running little offices of 5-10 persistent agents with defined roles and it’s crazy productive and code quality has never been better. I’ve been working on management of physical systems as well, and so far it’s been pretty solid.

      • my-next-account 9 hours ago ago

        Do you basically know that quality of code hasn't declined based on the number of incidents that are occurring? How do you even know what is happening?

        • duttish 9 hours ago ago

          Whenever I stop reviewing code I later find yanky shit or a small but important misunderstanding in the plan or...

          Claude has let me build really cool things but so far I'm not letting it run on it's own.

        • Olscore 9 hours ago ago

          About 25-30% of the things shipped by agents need fixes.

          • eru 8 hours ago ago

            How does that compare to human shipping?

            • necovek 8 hours ago ago

              I'd instead say that for both it's really at 100%, and not any number in between.

              Problem is that we can't define precisely what is good enough or when software is "finished": I mean, we are trying to do that with human languages, so it should not be a surprise.

              Yes, LLMs are now similarly aware of the average context a human would be aware of, but for anything specific to the situation a human has better chances of resolving the ambiguity.

      • saturn8601 10 hours ago ago

        Is there any great documentation you have followed to build your workflow or has it been experimentation? I've been experimenting a lot and gotten some cool things done but every time I read up on what everyone else is doing I realize I don't have much imagination and so I miss out.

        • seer 10 hours ago ago

          Honestly Claude code with agent teams and opus 5.5 is good enough for this by default no need for a custom harness / skill

      • technocratius 10 hours ago ago

        When you say, persistent agents, how does that work in practice for you?

      • GolfPopper 8 hours ago ago

        How do your users feel about the code?

  • woggy 9 hours ago ago

    I would replace the AI avatars with Daedalus, Icarus, Helios and Morpheus avatars from Deus Ex

  • lelanthran 10 hours ago ago

    What is the point of an AI agents having a standup?

    You don't need the self-policing enforcement of micromanagement that humans use standups for, after all.

    I mean, why not just write the harness so that a management agent constantly reads those sub-agent's files and course-correct the subs?

    • gwt4life 9 hours ago ago

      Its a gimmick.

  • apt-apt-apt-apt 7 hours ago ago

    Love the promo video song, how did you make it?

  • Olscore 9 hours ago ago

    It's cute, but my goal is building serious projects autonomously.

    https://github.com/pullboard-dev/pullboard

    Been using the Pullboard workflow to build software that MUST be correct, and recently open sourced it. There was another HN link today involving agents role-playing and I can't take that seriously, no offense. (Love fun projects like this, seriously, it's cool.)

    From my perspective, agents are great, they code better than most of us. But like all of us, they also make many mistakes or can't see how their changes impact the codebase. But that is a test-driven development mitigation along with some sort of scrutinizing process, "prove it". I'm actively working on this, and by actively working on it, I mean building serious apps that hit correctness walls, figuring out why Pullboard didn't catch it. To me, this is all about the process.

    • adamddev1 8 hours ago ago

      Tests are inadequate to catch bugs. Tests are like putting thin net in some little sections of a window, but property based testing is like something more like a decent net but not a solid barrier. And agents will lie, misunderstand or hallucinate things when asked to "prove it." And none of this fights off the bloat, performance issues, and decrease in maintainability. People keep thinking the process will save them from actually having to think through and put things together correctly. It's too bad that people are missing the satisfaction from personally building solid, correct things. And users are getting increasingly worse software because of this trend.

      • CurleighBraces 8 hours ago ago

        Entirely agree, also things can be logically correct and well tested but not the behaviour the user intended.

        Even if that's down to a bad initial prompt, or lack of data for the agent to notice the edge case I don't see how you can ever engineer a better agentic solution unless as a human you're monitoring the output.

        • adamddev1 7 hours ago ago

          > I don't see how you can ever engineer a better agentic solution unless as a human you're monitoring the output.

          I also totally agree with your response. But my point was that you can't engineer a better agentic solution. I'm arguing to never use agentic development because it is fundamentally flawed and inferior.

        • Olscore 8 hours ago ago

          You are correct that it can all be correct in one sense, and wrong in the product (“what the user wants”) sense. Of course you need to monitor the output, that’s a given. You need to question the entire system. That is the frontier for engineers.

          “How is the agent lying to me?” Etc.

      • undefined 8 hours ago ago
        [deleted]
  • bazza451 9 hours ago ago

    I hope when the bubble pops people will reference this and go “yeah it got so stupid that people started making teams of cartoons you could work with” - our industry is a complete clown car right now

    • hlynurd 8 hours ago ago

      Do you think the technology will vanish if and when a financially bubble pops?

      • grey-area 8 hours ago ago

        I don’t think it will disappear as the tools can be useful (finding bugs, searching large stores of info etc), but obviously bogus things like this meant to impress the gullible will disappear yes.

        Sadly the internet has already been ruined by bots pretending to be human, and there’s no way back from that.

      • bazza451 8 hours ago ago

        LLM’s - no. Cartoon colleagues - yes.

        Bring this into a professional workplace and watch yourself get laughed out of the building

      • lelanthran 3 hours ago ago

        > Do you think the technology will vanish if and when a financially bubble pops?

        The internet didn't disappear after dot-bomb, but the cue-cat did.

      • blamestross 8 hours ago ago

        To a certain extent yes, the subsidies that enable dumb use cases will go away. The capacity will remain, but who would bother?

    • undefined 8 hours ago ago
      [deleted]
  • teaearlgraycold 8 hours ago ago

    This makes me want to get a job just to be able to run this with my coworkers.

  • Gigachad 10 hours ago ago

    All of this stuff feels like it will be obsoleted by just running one instance faster.

  • RexHuang an hour ago ago

    [flagged]

  • danielsorok 7 hours ago ago

    The dependency tree and "who is blocked" view feels more useful than the call itself. When a few agents run in parallel, the failure mode I keep hitting is not missing status updates, it is silent waits on a shared file or API that nobody surfaces. Curious whether the standup format actually changed how you unblocked them, or whether the value was mostly having that tree visible.

  • alyssassan 6 hours ago ago

    [flagged]

  • senectus1 9 hours ago ago

    [dead]