Do you have a solution for degradation in accuracy when compiling larger amounts of llm-produced text?
I am also building LLM knowledge/memory systems and I've been surprised how bad LLMs are, even SOTA models, at summarizing non-trivial input batches of text. They get things wrong, distort the underlying meaning or data, etc.
Divide and conquer essentially, is what I've found so far to work best. Split things into smaller and smaller chunks to independently be verified, double-check everything, then coalesce upwards with verified summarizations. Have benchmarks for every single task and sub-task that will happen everywhere a LLM is involved, so you can measure improvements. Takes a ton more effort and tokens in the system itself obviously, but if you're not paying per token, it seems to work pretty well, albeit feels slightly over-engineered already.
A constant challenge. Don't have a perfect solution for it yet, but importantly every change to any article logs who did it, what it did and the reasoning behind it. So I have enough data to work with as I continue to improve things.
It’s very encouraging to see serious attempts at addressing continuity between agents’ outputs. Right now, everyone seems to be figuring out their own way of maintaining consistency across sessions without endlessly over-contextualizing each new one.
This feels like an important layer of the emerging agent stack, and I think this Show HN will be useful to a lot of people working through exactly that problem.
Would you consider this a different type product/benefit than all the "memory" things we have seen popping up everywhere?. Is it different just because it lives in the cloud? To me it feels like a different thing than memory.
I think they’re similar but different. Memory is usually single user focused, a summary of specific facts or instructions. OzBrain is everything that was captured, reasoned, promoted, etc. and auditable record of what the latest thinking is and why. When agents are the primary user you need a way to capture why they ended up at where they ended up. Memory is more like the filtered summary of everything it has saved as important.
I’m not sure I get the value over Obsidian, can you explain the static file issue? Only happens at 7K+ individual files? What happens if you always conjoin files?
Obsidian would fall into the same answer as how is this different than gBrain. Obsidian is a powerful and configurable... but requires more work to maintain and sharing knowledge with others is more difficult. I want this to be a super easy way for agents to connect to my knowledge, and make it easy for me to share my chunks of knowledge with my teammates and their agents. Just MCP in and have your agents get to work.
I'd say gBrain is also more structured to the priorities of a VC and less of a general-purpose graph-based memory. There are tons of hard-coded regexes that may work beautifully for Garry's priorities and workflow that don't fit other use cases, and are probably really brittle even if you're a VC who uses different words in your notes.
Yeah, close. But all handled so a user can just focus on connecting the MCP and doing work. Goes to my answer elsewhere here about different between gbrain / llm-wiki… those are highly configurable and powerful, but require technical knowledge and time. I think the vast majority of tech workers will need “that for dummies” ie this just works for them.
Almost all of this is stuff I have indeed "frankenstein[ed]" for myself, so consider this comment a +1 on market fit, there!
That also gives me a reason to pause, tho; the pitch in general is as solid as it can be on a site with markdown turned off (why, lord, why), but as a format minutiae megafan, I was left a little dissapointed. Where do you/OzBrain stand on Markdown formats? Could I use Sphinx with this, in rST and/or native MyST? Can it generate plain PDFs, fancy PDFs, or even animated static sites? etc. etc. etc. Not trying to gotcha, just curious to hear your thoughts & dreams on the topic!
It seems like some subculture(s) of SWE/SV/YC/AI has landed on obsidian-ish markdown with lots of wikilinks as the presumed default, which makes sense. So I'm assuming it's the same here. But also, your 'OzBrain vs. Obsidian' page does describe one difference as 'Markdown export anytime' vs. 'Markdown on disk' -- presumably that's just a hedge about hosting paradigm rather than a comment on the persistent format?
P.S. You're likely aware but there's at least one other company using Oz -- Warp's coding agent. Have you considered renaming this to something unimpeachable like DeepReasoningBrain? ;)
P.P.S. Holy hell your `eng-flow` thing is incredible. Maybe I'm behind the times, but... I mean, has anyone else processed how close we are to Minority Report and Iron Man?!
Along with my conversations with the 75 founders, there were two other distinct camps that I saw... those that care about the display and UX for them to see their .md files... and those (of which I am in) that don't care, I don't want to see the markdown, it is for my agents to see and relay to me what is important.
OzBrain is text only, and basic markdown formatting (OKF). No fancy PDFS, etc. BUT that is what your agent is for. If you want to generate a fancy PDF, have an agent reference the relevant articles and generate what you want. I've considered a lot of additional services like automatic ingestion, ie every call transcript auto-ingested as a source article in the brain... but you can do that with routines or Zapier. My focus is on the infra of the data getting to and from your agents.
I think Karpathy's llm-wiki was a boost to the wiki-markdown club, and it makes a ton of sense why a technical person would adopt that. They're already veru comfortable with github and moving files with terminal commands. They're not my target audience right now, but maybe when we have more robust brain maintenance they'll decide it's jsut easier to use OzBrain and not maintain their own thing.
P.S. In true move fast break things mode... I spent a good 15min working on the name. :) It works for now and if it really works well for people and they love it, the name won't matter so much.
P.P.S. Thanks, it's either genius or incredibly stupid... hard to know these days as stuff moves so quickly and your AI tells you you're so smart. It's what I HAD to build, because I don't speak any of the current languages. I lost my coding skills long ago, so at a point where it might make sense for a smart engineer to review or approve something... I needed to insert an agent that actually knows what it's talking about and understands how to make a good decision. That thing is constantly improving or breaking... I tried to automate one more step a few days ago and have been paying the price and bug fixing my thing that builds things, instead of just BUILDING THE THINGS! ugh.
P.P.P.S. I mean I used OzBrain to build OzBrain and make that eng flow work. If you think something in there would be useful, I'd just point an agent at it that is connected to your whole workspace/flow and ask it what is useful/dumb.
*Bonus point. Because my whole brain is in OzBrain and it knows what I'm building, why, how... I can take a talk transcript like Garry's from Startup School and just ask an agent "Save this transcript in the brain as a source and then review it and show me where this validates or invalidates some of my thinking. And what else would be interesting for me to consider in my broader work."
While I agree with you on most of what you’re saying, I think there is still an audience outside of your viewpoint that may see value. But thanks for sharing your insights! Super helpful.
Do you have a solution for degradation in accuracy when compiling larger amounts of llm-produced text?
I am also building LLM knowledge/memory systems and I've been surprised how bad LLMs are, even SOTA models, at summarizing non-trivial input batches of text. They get things wrong, distort the underlying meaning or data, etc.
Divide and conquer essentially, is what I've found so far to work best. Split things into smaller and smaller chunks to independently be verified, double-check everything, then coalesce upwards with verified summarizations. Have benchmarks for every single task and sub-task that will happen everywhere a LLM is involved, so you can measure improvements. Takes a ton more effort and tokens in the system itself obviously, but if you're not paying per token, it seems to work pretty well, albeit feels slightly over-engineered already.
A constant challenge. Don't have a perfect solution for it yet, but importantly every change to any article logs who did it, what it did and the reasoning behind it. So I have enough data to work with as I continue to improve things.
It’s very encouraging to see serious attempts at addressing continuity between agents’ outputs. Right now, everyone seems to be figuring out their own way of maintaining consistency across sessions without endlessly over-contextualizing each new one.
This feels like an important layer of the emerging agent stack, and I think this Show HN will be useful to a lot of people working through exactly that problem.
This is interesting.
Would you consider this a different type product/benefit than all the "memory" things we have seen popping up everywhere?. Is it different just because it lives in the cloud? To me it feels like a different thing than memory.
I think they’re similar but different. Memory is usually single user focused, a summary of specific facts or instructions. OzBrain is everything that was captured, reasoned, promoted, etc. and auditable record of what the latest thinking is and why. When agents are the primary user you need a way to capture why they ended up at where they ended up. Memory is more like the filtered summary of everything it has saved as important.
I have been thinking about this idea for a while. Cool work.
I’m not sure I get the value over Obsidian, can you explain the static file issue? Only happens at 7K+ individual files? What happens if you always conjoin files?
Obsidian would fall into the same answer as how is this different than gBrain. Obsidian is a powerful and configurable... but requires more work to maintain and sharing knowledge with others is more difficult. I want this to be a super easy way for agents to connect to my knowledge, and make it easy for me to share my chunks of knowledge with my teammates and their agents. Just MCP in and have your agents get to work.
I'd say gBrain is also more structured to the priorities of a VC and less of a general-purpose graph-based memory. There are tons of hard-coded regexes that may work beautifully for Garry's priorities and workflow that don't fit other use cases, and are probably really brittle even if you're a VC who uses different words in your notes.
This is a fascinating place in the stack to sit. I need to dig in a lot more. Nice start.
Good luck with the showing!
So is this cloud sync for my Md files? Who pays for the diffing and versioning?
In a way yes. A hosted llm-wiki, where I handle the diffing, versioning and audit log of what was changed, by what agent and why.
So a git repo + md files?
Yeah, close. But all handled so a user can just focus on connecting the MCP and doing work. Goes to my answer elsewhere here about different between gbrain / llm-wiki… those are highly configurable and powerful, but require technical knowledge and time. I think the vast majority of tech workers will need “that for dummies” ie this just works for them.
Almost all of this is stuff I have indeed "frankenstein[ed]" for myself, so consider this comment a +1 on market fit, there!
That also gives me a reason to pause, tho; the pitch in general is as solid as it can be on a site with markdown turned off (why, lord, why), but as a format minutiae megafan, I was left a little dissapointed. Where do you/OzBrain stand on Markdown formats? Could I use Sphinx with this, in rST and/or native MyST? Can it generate plain PDFs, fancy PDFs, or even animated static sites? etc. etc. etc. Not trying to gotcha, just curious to hear your thoughts & dreams on the topic!
It seems like some subculture(s) of SWE/SV/YC/AI has landed on obsidian-ish markdown with lots of wikilinks as the presumed default, which makes sense. So I'm assuming it's the same here. But also, your 'OzBrain vs. Obsidian' page does describe one difference as 'Markdown export anytime' vs. 'Markdown on disk' -- presumably that's just a hedge about hosting paradigm rather than a comment on the persistent format?
P.S. You're likely aware but there's at least one other company using Oz -- Warp's coding agent. Have you considered renaming this to something unimpeachable like DeepReasoningBrain? ;)
P.P.S. Holy hell your `eng-flow` thing is incredible. Maybe I'm behind the times, but... I mean, has anyone else processed how close we are to Minority Report and Iron Man?!
P.P.P.S. Is any part of that/this OS?
Along with my conversations with the 75 founders, there were two other distinct camps that I saw... those that care about the display and UX for them to see their .md files... and those (of which I am in) that don't care, I don't want to see the markdown, it is for my agents to see and relay to me what is important.
OzBrain is text only, and basic markdown formatting (OKF). No fancy PDFS, etc. BUT that is what your agent is for. If you want to generate a fancy PDF, have an agent reference the relevant articles and generate what you want. I've considered a lot of additional services like automatic ingestion, ie every call transcript auto-ingested as a source article in the brain... but you can do that with routines or Zapier. My focus is on the infra of the data getting to and from your agents.
I think Karpathy's llm-wiki was a boost to the wiki-markdown club, and it makes a ton of sense why a technical person would adopt that. They're already veru comfortable with github and moving files with terminal commands. They're not my target audience right now, but maybe when we have more robust brain maintenance they'll decide it's jsut easier to use OzBrain and not maintain their own thing.
P.S. In true move fast break things mode... I spent a good 15min working on the name. :) It works for now and if it really works well for people and they love it, the name won't matter so much.
P.P.S. Thanks, it's either genius or incredibly stupid... hard to know these days as stuff moves so quickly and your AI tells you you're so smart. It's what I HAD to build, because I don't speak any of the current languages. I lost my coding skills long ago, so at a point where it might make sense for a smart engineer to review or approve something... I needed to insert an agent that actually knows what it's talking about and understands how to make a good decision. That thing is constantly improving or breaking... I tried to automate one more step a few days ago and have been paying the price and bug fixing my thing that builds things, instead of just BUILDING THE THINGS! ugh.
P.P.P.S. I mean I used OzBrain to build OzBrain and make that eng flow work. If you think something in there would be useful, I'd just point an agent at it that is connected to your whole workspace/flow and ask it what is useful/dumb.
*Bonus point. Because my whole brain is in OzBrain and it knows what I'm building, why, how... I can take a talk transcript like Garry's from Startup School and just ask an agent "Save this transcript in the brain as a source and then review it and show me where this validates or invalidates some of my thinking. And what else would be interesting for me to consider in my broader work."
Meh
While I agree with you on most of what you’re saying, I think there is still an audience outside of your viewpoint that may see value. But thanks for sharing your insights! Super helpful.