> Protobuf now has modern IDE support for the first time
Weird post. I built IntelliJ protobuf support [1] while at Google like 10 years ago, and it started shipping by default with IntelliJ in ~2021. Maybe that's not considered "modern".
Sure, but having a generic LSP that works with all editors vs. an editor specific integration is still a very different thing.
It's a bit odd for the parent commenter to say "Weird post" and implying that having an editor specific feature would somehow make having an LSP that works for everyone be any less interesting.
The difference is that Buf's LSP is fully spec-complete since it's built on the same tooling that powers the buf CLI. I've seen several attempts at Protobuf LSPs over the years, but none that actually conform to the Protobuf specification, including the one you linked. That project is built on `emicklei/proto`, which readily accepts broken Protobuf and always has. The author says so in the README: "Current parser implementation is not completely validating .proto definitions." Maybe Buf should update the headline to "Protobuf finally has LSP support that works."
The bizarre "you're welcome" in the title is so obnoxious that I was sure it would be related to some sort of funny twist in the blog entry. Nope, they really thought that was appropriate.
Fair enough - feedback taken, we don't have a second career in comedy! Appreciate the feedback - we're renaming the post once GitHub is back online, and are adding the following to it for transparency:
> Edit: In full transparency, this blog post was originally titled "Protobuf finally has LSP support. You're Welcome." We thought this might generate some interest, but unfortunately this post sat unloved for months. Then, out of the blue, this post decided to make front page HN, and we received some choice feedback on our use of "You're Welcome" and our poor comedic abilities. Fair enough, we appreciate you speaking up - and we of course *thank you* for your continued support of our work!
Make sure to take reporting negativity bias into account. People will speak up when they don't like something, while stay quiet when everything is fine.
I, personally, enjoy Buf. It has solved many frustrations I've had with The Way it was previously done, manually installing protoc plugins, writing arcane flags into scripts nobody wants to touch, working around various quirks, like how protoc produces different output based on which directory you invoke it in, et cetera... I don't quite like the remote plugin model and rate limiting, but using local plugins is easy enough that it doesn't matter.
Buf has made my life easier, and I've actually been using `buf_ls` for a while now in NeoVim. I understand where the criticism of the title comes from, but I don't personally feel it in this particular case. In fact, I didn't even notice it.
I don't even mind the monetized properietary cloud thing, as long as Buf CLI and the related toolchain remains free software, serving the community. Hope you won't do the ol' switcharoo any time soon.
Totally, and we appreciate the sentiment. But the title was meant to be lighthearted, and instead it clearly struck a nerve. That's not our intent, and it was easy to change, so we just swallowed our pride. We rather the conversation be around Protobuf tooling, and not our poor comedy.
And the Buf CLI will always be free and OSS, don't worry!
I’ve seen similar wording from colleagues on the ASD (with Asperger’s).
It can be hard for some people to navigate the nuances between usage of a tongue-in-cheek usage of the phrase which is more socially acceptable or humorous from the version that reads as arrogant
It sounds like you're having a bad day or something. Putting "you're welcome" in the title of your own blog, on a free software release, is fine. It's not pushed into anyone's face. It loses some points for being a corporate ad but that would happen regardless of wording.
That's a shared chore. Do you not see the obvious differences in the situation? If they said "you're welcome" about an obligation like paying taxes and also made it very personal, then it would be close in feeling to your scenario. But it's not. If you offer a genuine gift, in a situation where there is no expectation of payback or debt, you can say you're welcome.
It’s impolite and weird to say “you’re welcome” when offering a gift for which the recipient did not first express gratitude. It follows “thank you;” it does not precede it.
Try this instead: give a stranger a piece of pre-chewed gum and say “you’re welcome!” Hell, even if it’s not pre-chewed, it’s just presumptuous and rude. You have no idea if they value it or not. They might not even like it.
Even if my barista handed me the coffee I ordered and said “you’re welcome” before I thanked them first, I’d think they’re a little off kilter.
Sounds like a you problem. If someone did me a service and told me I was welcome, my take would be entirely based on the delivery of the phrase. I don’t even know what to do with the gum analogy because that’s just nonsense.
Saying "you're welcome" without a "thank you" preceding will sound passive aggressive to quite a lot of people; it's the type of thing that a parent does to chide a small child who forgets their manners in the same way they might prompt a child with "What's the magic word?" if they grabbed at something without saying please.
If this doesn't make sense to you, I'm not sure there's much anyone can do to convince you, because it's purely social convention. It's certainly possible that the author didn't realize it either, but having to deal with reactions about word choice when putting out writing explicitly for public consumption kinda comes with the territory.
This goes for Hacker News comments too of course, but as someone with no horse in this race, your interference about their word choice seems like much more of a stretch than theirs about the author. Not recognizing that a title to a blog post will evoke a negative reactions in a lot of readers does genuinely seem like information about the author socially in a way that having that common negative reaction does not.
It literally can't precede it, that's the entire point that's being made; it sounds rude in a blog post because the social expectations are being violated. When social convention says a phrase is rude in certain circumstances, and those circumstances are not (or cannot) be met, then social convention say that the phrase is rude.
You seem incredulous at the idea that some things that might be socially acceptable in conversational circumstances would not be perceived similarly in published writing, but I don't really understand why. There are all sorts of things that might not cause someone to bat an eye in a spoken conversation that are much more difficult to use in writing.
You say that expectation is violated here, I say it doesn't apply here.
I don't think we're going to resolve this difference in view.
But please stop making wild guesses about how I handle broad categories of communication when I'm commenting on one specific case. I don't say it doesn't make sense. I'm not incredulous. That's really obnoxious.
> How is a thank you supposed to precede it in a blog post?
That's a pretty clear distinction between a blog post and conversations.
Now you're saying:
> But please stop making wild guesses about how I handle broad categories of communication when I'm commenting on one specific case. I don't say it doesn't make sense. I'm not incredulous. That's really obnoxious.
I honestly have no clue what you're even trying to say at this point. You don't seem to want to say it directly, and yet you're now claiming I'm being obnoxious for trying to figure it out.
> I don't think we're going to resolve this difference in view.
That's the one thing we can agree on, since from my perspective, you're being a lot more obnoxious by pretty much treating people the exact same way you've been complaining about the entire time.
You took what I said about a single situation and tried to extrapolate it into entire categories of your choosing. That's why you're confused. Don't do that and you won't be confused.
I have been very direct. You're trying to make it more indirect when you try to generalize like this. And throwing out "if this doesn't make sense to you", or "you seem incredulous" about the entire existence of certain social norms that apply in some situations but not others, is a bad way of trying to figure out what I mean; it's rude.
Edit: An analogy: I have a critique about a type of waffle and you imply I don't understand the purpose of baked goods in general. This is not a good way to converse and the confusion is not my fault.
You're making it sound like suggesting that "blog posts" are a category of "writing" is some sort of random arbitrary choice on my part, and yet you've still refused to elaborate on how or why you think they're different. You're continuing to take issue with oddly specific parts of how I've phrased things and yet continue to use plenty of phrasings that sound at least as rude to me as anything I've said.
I don't think you've been direct at all about anything relating to the actual original conversation; the only things you've been direct about are in telling me or others that we're wrong and pointing out the things you think we've been doing wrong (along with a random assumption about someone else's life based on the comment they initially made), and then proceeded to take issue with every instance where someone else does the exact same things.
It's pretty much impossible phrase this point because you have gotten upset any time I happen to incidentally phrase something using the second person pronoun, but I'll try my best: when person A says something, and person B doesn't understand it, that's not sufficient evidence to conclude that person A or person B is at fault. Communication doesn't exist in only one direction. Sometimes it might be because person B did a bad job understanding, but sometimes it's also because person A did a bad job communicating. In the scenarios where person A happened to do a bad job, it's pretty common for them not to realize, because if they did realize, they would follow up by clarifying. It's pretty much impossible for person A to know whether they communicated clearly or not without some other person to interpret it, because communication that isn't interpreted in some way by someone else isn't actually communication.
> You're making it sound like suggesting that "blog posts" are a category of "writing"
This is so far away from anything I meant to say that I have no idea what happened.
> you've still refused to elaborate on how or why you think they're different
I don't?
I think blog posts are different from conversations but that should be self-explanatory.
> along with a random assumption about someone else's life based on the comment they initially made
"Bad day or something" is suggesting they're being too harsh. It's not suggesting they fail to understand the situation, the way you did with me.
> I don't think you've been direct at all about anything relating to the actual original conversation; the only things you've been direct about are in telling me or others that we're wrong and pointing out the things you think we've been doing wrong
What do you want me to be more direct about? The answer to whether I understand these broad categories of social norms exist is yes. I just don't think the blog post falls into the category where you need a thank you first. Is there anything else you want to ask?
I don't think I can say anything useful about your last paragraph.
No, I don't think I have anything else to ask you, because I'm honestly still struggling to understand any sort of consistency between what you've said overall. I still maintain that your initial comment was telling someone else they were wrong and that you haven't given any explanation that makes sense to me about why other than just asserting it as self-evidently true.
Looking back though before writing this comment, I did notice that the comment seems to have been flagged, and that doesn't seem necessary to me. On the off chance that it might have been a factor in your subsequent responses, I went ahead and vouched for it because I think I would be frustrated if I were flagged for a comment that was similarly benign (even if it was wrong).
They are doing us a favor, there's no denying that. Just like dozen of other providers did them a favor by providing them with the free software pieces they relied on to build their product.
Of course it's poor taste to add the "You're welcome". But poor taste is par for the course for such startup, unfortunately.
I meant it in a broader sense. Even for paid products, corporations view themselves as doing customers a favor. That's the sort of antagonistic relationship they define.
Despite interesting legal definitions, companies are just groups of people.
I also read the title as having an air of self-importance or maybe passive-aggression. If you read it a little more dispassionately then it's just an awkward way of stating that this is a gift.
Why is this a gift? The company does not need to give it to you, though it arguably benefits from doing so. It did so with some definition of "free."
I looked at the dependencies and noticed that it wasn’t using an existing Protobuf parser which means they reimplemented the parser from scratch. Perhaps due to a lack of error recovery in the existing implementations? I don’t have the energy right now for further inspection.
It is definitely best to reuse the parser for the runtime when implementing an LSP but to do so properly means implementing the parser itself as a standalone library. Even better is shipping the semantic analysis as well!
Implementation drift is definitely an issue.
But great project anyways, just wanted to put my thoughts on the matter into the conversation!
It absolutely can be done [1] but it is a lot of effort to do high quality error recovery. It's easier in languages with natural "synchronization points" [2], harder in languages which don't have them.
Having built a lot of protobuf tooling, I'd estimate that protobuf largely falls into the former camp; most of your time working with .proto files is operating on fields which terminate using `;` at the end of the line (though multi-line is also possible).
Adding to this, rustc's parser is both correct and fault-tolerant, yet rust-analyzer ended up building its own for different reasons (needing a CST and not an AST).
I settled on starting with a CST using rowan, doing semantic analysis, having an agnostic editor services, exposing both an LSP interface with tower, and a WASM interface (for Monaco in the browser [non-LSP!!]) and having the runtime binary doing the interpretation/compilation with the same pathway, with a slight jog, as the editor services. Phew!
It is my understanding that basically all LSP implementations use tree-sitter because of its incremental rebuild capability, and this requires re-implementing the parser.
That's just not true. rust-analyzer does not use tree-sitter. Neither does gopls. Nor Pylance. Or tsc. In fact I don't know if there is a single popular LSP that uses it (but there probably is).
Many editors use tree-sitter, but that is separate from the LSP.
Lots of naysayers here, but an advantage of protobuf is that proto files are hand-writeable, and therefore having an LSP for that could be useful.
That said, proto itself dissuades or forbids the kind of common things you might do with a LSP, such as renaming.
Renaming fields is a big no-no [edit: this isn't true, see corrections below], as is doing things like re-ordering fields.
A core idea of proto is that versions are strictly compatible with with previous versions. This itself has limitations and challenges for migrations, but encourages good practice about compatibility that usually gets ignored or hand-waved away in most ecosystems.
I accept however that it's often easy to offload both the re-structuring and the checking of version compatibility to an LLM and let them go at it.
Renaming is forbidden though (because JSON and textproto). In Google, it's a documented antipattern to try to make protobuf look "nice" by changing field indices, rearranging fields, etc. — the common ground is that it's better to not do it.
As long as you know the use cases of your fields, renaming is just fine. My team regularly does it. We also maintain our own serializer and deserializer json, XML, and fixed with formats. The json one we handle serialization using an annotation to say how it should be exported.
Of course, the whole concept of a breaking change does not really apply if all usage is within the controlled code. The problems start to appear when you have external users with old versions; then protobuf starts having a bunch of weird limitations. My favorite is that it's forbidden to move a field into or out of a `oneof`: it's a compatible change in a sense of wire format and JSON, but breaks the generated Golang code.
Breaking changes matter even with controlled code, because you can have requests that straddle upgrade boundaries, it's fiction to believe that all services are upgraded at the same moment, and pretending that is the case is the sort of thing that leads to quiet data corruption or mysterious bugs that can never seem to get reproduced.
Even if you upgrade with coordinated downtime across your entire service stack ( a bit old-school, but still happens more than you might imagine. ), then you still have to occasionally deal with requests that get persisted somewhere, possibly for support purposes, and it's much handier if the wire format remains compatible, at least between immediate versions.
Yes, breaking changes need to be staged carefully/compatibly across versions.
Simple renaming/reordering is not a breaking change between code versions. The wire format knows numeric ids, not names. It breaks code compilation until the renames are put into effect. There is a subtle breakage where someone renames a field (think: field -> old_field) and then later adds "field" to mean something else; software that isn't recompiled in this window might not recognize that "field" is potentially different. Uncompiled languages may suffer this even more subtly. -- All of this to say, breaking compilation is not the end of the world but there are more dragons as the scope grows.
Stop-the-world is usually only needed when someone has made an unplanned/incompatible change with versions that are still running. If your infrastructure+development is done right (hah) this should never happen.
While not a direct competitor to protobufs, if you are working in the video game space where struct versioning is not needed, there is an alternative language called "schema" that supports C, C++, C#, Golang, Rust and JavaScript.
Multiplayer games typically deploy both client and server at the same time, and refuse to connect a client if it doesn't speak the exact same protocol as the server.
Thus all the versioning overhead of protobufs is not needed for this wire protocol.
(Yes, games still use versioning everywhere else where it makes sense: save games, asset data, config etc...)
You can do this but if it's too granular (like you have no concept of version compatibility) then it can heavily split your matchmaking.
Plus the headaches of keeping many out of date builds up to date enough to deploy.
Even if you don't care about in game compatibility, all your servers still talk to some centralized data store and that will likely want a single deploy that handles old clients
The promise of Protobuf/gRPC for seamless communication across tech stacks just wasn't delivered. Third party projects like Buf or BetterProto for Python try their best but it's still a nightmare trying to integrate with gRPC from something remotely modern, like Python with uv and async/await. Plus there's a lot of little pains like not being human-readable, always requiring HTTP/2 and thus HTTPS, and so on.
JSON isn't "optimal" but after transfer compression it's not that bad, and the support for JSON Schema is much better across the stacks I use, e.g. Zod in JS/TS, Pydantic in Python, etc. And it's fully usable "by hand" without having the schema, for one-off scripts and such -- compare with Protobuf where you need the full definition to even parse a chunk of data.
In the web dev world, I’ve noticed there are a lot of people who think that JSON is the ultimate modern data format. Typically those same people have barely heard of JSON Schema.
This has been driving me crazy ever since I started using protobuf/grpc and realized more major tech companies (in the cloud/infra/data world at least, and a lot of other SAAS) were using it or something similar (eg capn proto) internally than not.
It feels like we’re in some sort of deadlock where each of them think “proto/grpc are too niche to support for external users, better just use JSON”, keeping it unfamiliar for an Average Web Developer. But if every company using it just exposed it to third parties/added it to their public APIs, it would immediately be common (and trendy) enough for every web developer to learn it and start using it.
If Google added the missing HTTP/2 streaming support to browsers (blocking native bidi grpc streaming) it would have an instant killer app in making it easier to implement websocket-like client/server applications. It makes absolutely no sense that full duplex bidi was added to the HTTP/2 but remains unimplemented in browsers.
The role MCP, OpenAPI, and JSON schema fill all would be a million times simpler if they were based on protobuf instead of JSON. I can forgive OpenAPI/JSON but it honestly pisses me off that we ended up with MCP and JSON-RPC + JSON schema, and people think these are cool/good tools, and actively adopting them. Just piling on the slop
I pushed for us to adopt this setup at my previous company and it was wildly successful. I thought frontend devs would resist but they absolutely loved the generated types and clients. It quickly seemed silly to do anything else.
This is the way! It's great to just point frontend developers to the Buf documentation for new features and they easily have all the types and generated clients available in their projects too.
Unpopular opinion -- LSPs are bad because they introduce latency into coding. For every keystroke, my IDE needs to make request/response with LSP, instead of using its own internal parser/colorer/autocomplete.
LSP isn't used for syntax highlights usually, that's still the job of the editor. This is where something like treesitter usually comes in. LSP is only used in this case for errors/annotations/etc.
You can just not use an LSP? My understanding is that LSPs aren't made to be the fastest (code completion|syntax highlighter|code analysis), but easily integrated into all sorts of IDEs/text editors.
As far as I understand, IDEs ara way less eager to spend resources onto their own integration if there's an LSP. E.g. last time I tried to code in Zig there was either LSP or nothing.
Protobuf was never particularly a Java thing. You could more accurately say it was a C++ thing. It was initially developed in C++, and used mainly in C++ services at Google. The implementation and early ecosystem were C++-centric.
Later, Go became one of the major Protobuf ecosystems, and today it would be understandable to think it was a Go thing.
For me that is how I know something like protobuf is good. It is a nuisance to manage and distribute the definitions, adds a build step even to languages with no build step normally, is slower than almost every alternative, and artificially restricts you from doing lots of common things. It's so good!
And look at the code quality of the implementation! It's like a team of interns wrote it while drunk. It is a complete spaghetti mess, but has tons of super convoluted micro optimizations that are slower than just doing the most obvious thing, but make the implementation confusing and indirect. It's trash code.
What are good alternatives when you need a common "single source of truth" schema shared between multiple languages? We use protobuf between c# and Python.
I quite like the look of Typespec though I haven't used it much.
I always thought Thrift was waaay better than any of the alternatives, but it always had terrible documentation and I think it died mainly because of that.
The goal is not to have a single source of truth schema. That is a means to some other goal, and it's not even a good means.
If you never change the schema then you don't have to worry about it, get things working and never look back.
If you do change your schema from time to time, you need testing between the two systems. If you have good tests again a single source of truth is fully redundant, both systems are talking just fine. If you don't have tests things can and will break all the time even using protobuf.
> The goal is not to have a single source of truth schema. That is a means to some other goal, and it's not even a good means.
It’s about data transmission. Being able to encode and decode in a type safe manner between different languages (and so, different platforms) is a goal that makes a lot of sense.
> If you do change your schema from time to time, you need testing between the two systems
Or you could just use a defined format that doesn’t require testing. I rarely use protobuf but I can see why people do. The guaranteed backwards compatibility is huge for people who can’t just publish a new web frontend at the drop of a hat.
If you understand how to evolve protobuf schema definitions, then you don’t really need testing. You instinctively know how the parser works when it is parsing data with a different schema from what it expects. And that’s a powerful thing. If your things break even when using protobuf then you don’t grok protobuf.
It’s probably not an exaggeration to say that being able to avoid tests between different systems who have different versions of the schema is a core goal of protobuf. Why? These two different systems are probably owned by different teams, and introducing explicit tests between different versions of them increases coupling between them.
Protobuf only allows you to add optional fields after the initial version, right? Because otherwise it would not be backwards compatible. Protobuf does not allow you to define contingent logic between fields. Optional fields are always nullable (or you must provide a default). This forces you to know all of this and use a method to see if it was actually set or just defaulted. So you have a ton of nullable/defaulted fields that likely are required to be filled in or not filled in, contingent on other values in the same struct. For example if the charge type is "purchase" then price must be not null and > 0. That sort of thing. This is so common you should just assume your app has a million little rules like this that are assumed. Protobuf does not help you here at all. You need some other logic to validate the data on top of protobuf. once you have that anyway protobuf's value is that it's expensive, requires build steps, is actually not fast, and forces you to distribute the schema between different apps somehow.
If it's not obvious yet, lets say you are sending your charge structs and you have a bug where sometimes you don't set price. It's defaulted to 0 or -1 or whatever nonsense value the default is, or null, it doesn't matter it's not correct sometimes. That is why you need tests my guy. Protobuf can't fix this. If you use json the tests make sure everything protobuf does for you is done too. In a world where you have to write tests because protobuf can't force you to correctly set fields, you have tests already, and protobuf didn't help you at all. Whether you call it tests or input validation or whatever, protobuf definitions are insufficient, and when you have what is sufficient it 100% covers everything protobuf does 'for you'.
If you don't get it at this point then lets just agree that you will never get it.
> Protobuf only allows you to add optional fields after the initial version, right?
That’s not all. For example, you can also change fields from optional to repeated or vice versa. For another example, you can also delete fields.
> Protobuf does not allow you to define contingent logic between fields.
You are just saying that protobuf is not a data validation library. That’s true; it only handles data serialization. You need to write validations yourself. And you will most likely need tests to test your validation code. But those are entirely tests within a single process; they are not tests involving a producer and a consumer of a protobuf message, which is the wrong kind of tests.
> For example if the charge type is "purchase" then price must be not null and > 0. That sort of thing.
That sort of thing sounds like you are not using `oneof` appropriately. The charge type shouldn’t be a field. It should be a submessage called Purchase.
You don’t get what they say. It’s not about about how efficient it is after encode, it’s about how fast encode is. They are not spewing bs, they’re focusing on a single point. The question is: do you send it over the wite more often than performing encode/decode.
That's what the person you replied to is talking about, and they're right. Putting aside the final byte size (where protobuf also wins), protobuf is faster at both encoding and decoding than json. There are numerous benchmarks you can find that show this.
The advantages of json are not related to performance.
If you're writing JS you cannot beat JSON.parse, because you're running the most optimized C++ implementation of JSON which will outcompete any decoder written JS itself.
It is generally correct. When both implementations are completely in the same language Protobuf will win. Instead of a pure-JS implementation one can also make a FFI module for Protobuf.
The comparison that's being made is equivalent to implementing quick sort in Python and bubble sort in C++ and then declaring bubble sort is the better alternative.
Anyone with a healthy understading of computer science and programming experience will not make bullshit claims like that.
No, protobuf is not faster, not on Python, because google's implementation is poor. Python is the #1 language in the world right now. Don't use protobuf.
Buf's offering of protobuf registries and codegen SDKs for microservices seems less necessary in the LLM era.
I'm starting to question many of protobuf's advantages (perhaps not the wire format). Add to that monorepos and other fads of the 2010s given the rise of LLMs.
I used to be a big believer in this stuff, but I'm quickly having my core assumptions change out from under me.
A big differentiator is whether one imagines an LLM in-band with most/all future software. If there is, and we’re deferring until very late parts of a program that world have been load bearing, and we’re able to programmatically ands reliably squint and say “eh i know what you meant” … then yeah formalizations seem superfluous-to-counterproductive.
OTOH if LLMs are to write, but not supplant, much of software, then boundaries, delegation to deterministic layers, good compilers to bonk miscreant models on the head with error message seem essential.
At one point it would have been shocking to assert that the compiler would live in-band with the program too. and yet JS eats the world. It seems shocking today that we could have a universal prior over the world operating in the ms/us nJ/pJ range required. And yet … ?
The real argument shouldn't be about protocols becoming obsolete, but programming languages that are "less efficient" could eventually become obsolete in favor of highly scalable and performant languages due to LLMs when the main gap becomes knowing how the tech works at a high level, and not the syntax, why code in one language over another if you don't need to worry about messing up on syntax, only about reviewing logic for sanity and correctness against business rules as well as validating that it is stable code.
Anybody who thinks that you can just chuck unstructured data into LLM and YOLO the app is an idiot.
This works up to a point, and then it doesn't. And you're left with tons of inconsistently formatted data.
My company is built on protobufs from ground up :) We use it in the database, for remote calls, on the frontend, etc. The protobuf language is not great, but it's about the right balance between too expressive and too restricting.
And the best thing is that it's compact, compared to OpenAPI.
I feel exactly that way about REST. The assumption that LLMs make schemas obsolete misses how structured outputs actually work in production. When you have probabilistic models generating code, strict contracts become more critical, not less. It is no coincidence that several major LLM platforms rely on ConnectRPC and Protobuf for their own APIs.
You're right, but the adeptness of models to spin up clients and behaviors on the fly is remarkable. They're capturing the semantics of behavior at a deeper level.
If we do strict schemas, I'd like to see less ceremony around them. Tool calls instead of brittle build steps and protocol registries.
Hm... Maybe. In my view, an IDL is part of the input that you absolutely want humans to author or carefully review at least. In my experience, the ceremony around generating code is also performed very well by LLMs. But I do agree, there's definitely some changes that are needed to integrate Protobufs better. Some languages have built-in tooling to make it seamless, but it's definitely not universal.
That's not Google. Buf.build is, hilariously enough, a completely separate company that tries to make sense of Google's protobuf, and it's doing a decent job from what I can see.
They reluctantly acknowledged that as a mistake by adding an actual "optional" modifier for field presence. proto3's previously so-called optional were really all present with a zero default value; proto2 could have supported that for required fields at the code level if they had wanted. In the end, all they've changed is how they frame the feature.
Aren't LSP dying (and eventually IDE, at least in their current form) as everybody use LLM to code. I know some big tech companies redirected teams supporting them to new efforts (ie. tool integration with AI).
> Protobuf now has modern IDE support for the first time
Weird post. I built IntelliJ protobuf support [1] while at Google like 10 years ago, and it started shipping by default with IntelliJ in ~2021. Maybe that's not considered "modern".
[1] https://github.com/jvolkman/intellij-protobuf-editor
I wouldn't consider an IDE-specific integration very modern (in IDE terms) in the age of LSP.
Conversely, LSP is quite limited in comparison to what can be done with a real IDE integration in JetBrains product suite, so it is not so clear cut. See this post for some details: https://matklad.github.io/2023/10/12/lsp-could-have-been-bet...
Sure, but having a generic LSP that works with all editors vs. an editor specific integration is still a very different thing.
It's a bit odd for the parent commenter to say "Weird post" and implying that having an editor specific feature would somehow make having an LSP that works for everyone be any less interesting.
Yeah, it would be like someone complaining for a new language getting invented because clang implemented a frontend for it years ago
"IntelliJ is for Java, and Java is legacy!"
Probably the author
What an oddly arrogant post, there's been a Protobuf LSP available for years: https://github.com/lasorda/protobuf-language-server
The difference is that Buf's LSP is fully spec-complete since it's built on the same tooling that powers the buf CLI. I've seen several attempts at Protobuf LSPs over the years, but none that actually conform to the Protobuf specification, including the one you linked. That project is built on `emicklei/proto`, which readily accepts broken Protobuf and always has. The author says so in the README: "Current parser implementation is not completely validating .proto definitions." Maybe Buf should update the headline to "Protobuf finally has LSP support that works."
The bizarre "you're welcome" in the title is so obnoxious that I was sure it would be related to some sort of funny twist in the blog entry. Nope, they really thought that was appropriate.
Fair enough - feedback taken, we don't have a second career in comedy! Appreciate the feedback - we're renaming the post once GitHub is back online, and are adding the following to it for transparency:
> Edit: In full transparency, this blog post was originally titled "Protobuf finally has LSP support. You're Welcome." We thought this might generate some interest, but unfortunately this post sat unloved for months. Then, out of the blue, this post decided to make front page HN, and we received some choice feedback on our use of "You're Welcome" and our poor comedic abilities. Fair enough, we appreciate you speaking up - and we of course *thank you* for your continued support of our work!
Make sure to take reporting negativity bias into account. People will speak up when they don't like something, while stay quiet when everything is fine.
I, personally, enjoy Buf. It has solved many frustrations I've had with The Way it was previously done, manually installing protoc plugins, writing arcane flags into scripts nobody wants to touch, working around various quirks, like how protoc produces different output based on which directory you invoke it in, et cetera... I don't quite like the remote plugin model and rate limiting, but using local plugins is easy enough that it doesn't matter.
Buf has made my life easier, and I've actually been using `buf_ls` for a while now in NeoVim. I understand where the criticism of the title comes from, but I don't personally feel it in this particular case. In fact, I didn't even notice it.
I don't even mind the monetized properietary cloud thing, as long as Buf CLI and the related toolchain remains free software, serving the community. Hope you won't do the ol' switcharoo any time soon.
Totally, and we appreciate the sentiment. But the title was meant to be lighthearted, and instead it clearly struck a nerve. That's not our intent, and it was easy to change, so we just swallowed our pride. We rather the conversation be around Protobuf tooling, and not our poor comedy.
And the Buf CLI will always be free and OSS, don't worry!
I’ve seen similar wording from colleagues on the ASD (with Asperger’s).
It can be hard for some people to navigate the nuances between usage of a tongue-in-cheek usage of the phrase which is more socially acceptable or humorous from the version that reads as arrogant
surely we're not a community so unfamiliar with lacking social skills and this aberration isn't at all important to the topic at hand
It sounds like you're having a bad day or something. Putting "you're welcome" in the title of your own blog, on a free software release, is fine. It's not pushed into anyone's face. It loses some points for being a corporate ad but that would happen regardless of wording.
I think it makes the author sound like a dick.
Test for yourself: next time you take out the garbage, loudly tell your partner or roommate, “I took out the garbage. You’re welcome.”
That's a shared chore. Do you not see the obvious differences in the situation? If they said "you're welcome" about an obligation like paying taxes and also made it very personal, then it would be close in feeling to your scenario. But it's not. If you offer a genuine gift, in a situation where there is no expectation of payback or debt, you can say you're welcome.
It’s impolite and weird to say “you’re welcome” when offering a gift for which the recipient did not first express gratitude. It follows “thank you;” it does not precede it.
Try this instead: give a stranger a piece of pre-chewed gum and say “you’re welcome!” Hell, even if it’s not pre-chewed, it’s just presumptuous and rude. You have no idea if they value it or not. They might not even like it.
Even if my barista handed me the coffee I ordered and said “you’re welcome” before I thanked them first, I’d think they’re a little off kilter.
Remember that this is a post on their website. They're not pushing it onto anyone. So it's not like the gum.
> Even if my barista handed me the coffee I ordered and said “you’re welcome” before I thanked them first, I’d think they’re a little off kilter.
Jeez...
But also that's a conversation where you can thank them first. That part isn't applicable to a post like this.
Sounds like a you problem. If someone did me a service and told me I was welcome, my take would be entirely based on the delivery of the phrase. I don’t even know what to do with the gum analogy because that’s just nonsense.
I don't think it's just a "me" "problem"; there are many others who agree with my perspective in this discussion. Take a look around.
Saying "you're welcome" without a "thank you" preceding will sound passive aggressive to quite a lot of people; it's the type of thing that a parent does to chide a small child who forgets their manners in the same way they might prompt a child with "What's the magic word?" if they grabbed at something without saying please.
If this doesn't make sense to you, I'm not sure there's much anyone can do to convince you, because it's purely social convention. It's certainly possible that the author didn't realize it either, but having to deal with reactions about word choice when putting out writing explicitly for public consumption kinda comes with the territory.
This goes for Hacker News comments too of course, but as someone with no horse in this race, your interference about their word choice seems like much more of a stretch than theirs about the author. Not recognizing that a title to a blog post will evoke a negative reactions in a lot of readers does genuinely seem like information about the author socially in a way that having that common negative reaction does not.
How is a thank you supposed to precede it in a blog post?
The social convention is for conversations.
It literally can't precede it, that's the entire point that's being made; it sounds rude in a blog post because the social expectations are being violated. When social convention says a phrase is rude in certain circumstances, and those circumstances are not (or cannot) be met, then social convention say that the phrase is rude.
You seem incredulous at the idea that some things that might be socially acceptable in conversational circumstances would not be perceived similarly in published writing, but I don't really understand why. There are all sorts of things that might not cause someone to bat an eye in a spoken conversation that are much more difficult to use in writing.
You say that expectation is violated here, I say it doesn't apply here.
I don't think we're going to resolve this difference in view.
But please stop making wild guesses about how I handle broad categories of communication when I'm commenting on one specific case. I don't say it doesn't make sense. I'm not incredulous. That's really obnoxious.
Verbatim, you said:
> The social convention is for conversations.
> How is a thank you supposed to precede it in a blog post?
That's a pretty clear distinction between a blog post and conversations.
Now you're saying:
> But please stop making wild guesses about how I handle broad categories of communication when I'm commenting on one specific case. I don't say it doesn't make sense. I'm not incredulous. That's really obnoxious.
I honestly have no clue what you're even trying to say at this point. You don't seem to want to say it directly, and yet you're now claiming I'm being obnoxious for trying to figure it out.
> I don't think we're going to resolve this difference in view.
That's the one thing we can agree on, since from my perspective, you're being a lot more obnoxious by pretty much treating people the exact same way you've been complaining about the entire time.
You took what I said about a single situation and tried to extrapolate it into entire categories of your choosing. That's why you're confused. Don't do that and you won't be confused.
I have been very direct. You're trying to make it more indirect when you try to generalize like this. And throwing out "if this doesn't make sense to you", or "you seem incredulous" about the entire existence of certain social norms that apply in some situations but not others, is a bad way of trying to figure out what I mean; it's rude.
Edit: An analogy: I have a critique about a type of waffle and you imply I don't understand the purpose of baked goods in general. This is not a good way to converse and the confusion is not my fault.
You're making it sound like suggesting that "blog posts" are a category of "writing" is some sort of random arbitrary choice on my part, and yet you've still refused to elaborate on how or why you think they're different. You're continuing to take issue with oddly specific parts of how I've phrased things and yet continue to use plenty of phrasings that sound at least as rude to me as anything I've said.
I don't think you've been direct at all about anything relating to the actual original conversation; the only things you've been direct about are in telling me or others that we're wrong and pointing out the things you think we've been doing wrong (along with a random assumption about someone else's life based on the comment they initially made), and then proceeded to take issue with every instance where someone else does the exact same things.
It's pretty much impossible phrase this point because you have gotten upset any time I happen to incidentally phrase something using the second person pronoun, but I'll try my best: when person A says something, and person B doesn't understand it, that's not sufficient evidence to conclude that person A or person B is at fault. Communication doesn't exist in only one direction. Sometimes it might be because person B did a bad job understanding, but sometimes it's also because person A did a bad job communicating. In the scenarios where person A happened to do a bad job, it's pretty common for them not to realize, because if they did realize, they would follow up by clarifying. It's pretty much impossible for person A to know whether they communicated clearly or not without some other person to interpret it, because communication that isn't interpreted in some way by someone else isn't actually communication.
> You're making it sound like suggesting that "blog posts" are a category of "writing"
This is so far away from anything I meant to say that I have no idea what happened.
> you've still refused to elaborate on how or why you think they're different
I don't?
I think blog posts are different from conversations but that should be self-explanatory.
> along with a random assumption about someone else's life based on the comment they initially made
"Bad day or something" is suggesting they're being too harsh. It's not suggesting they fail to understand the situation, the way you did with me.
> I don't think you've been direct at all about anything relating to the actual original conversation; the only things you've been direct about are in telling me or others that we're wrong and pointing out the things you think we've been doing wrong
What do you want me to be more direct about? The answer to whether I understand these broad categories of social norms exist is yes. I just don't think the blog post falls into the category where you need a thank you first. Is there anything else you want to ask?
I don't think I can say anything useful about your last paragraph.
No, I don't think I have anything else to ask you, because I'm honestly still struggling to understand any sort of consistency between what you've said overall. I still maintain that your initial comment was telling someone else they were wrong and that you haven't given any explanation that makes sense to me about why other than just asserting it as self-evidently true.
Looking back though before writing this comment, I did notice that the comment seems to have been flagged, and that doesn't seem necessary to me. On the off chance that it might have been a factor in your subsequent responses, I went ahead and vouched for it because I think I would be frustrated if I were flagged for a comment that was similarly benign (even if it was wrong).
"You're welcome" is absolutely hilarious to read from a company post.
Your HN comment has a reply. You're welcome.
What can I say, except you're welcome.
It's gotten out of hand. Companies making money think they do everyone a favor.
They are doing us a favor, there's no denying that. Just like dozen of other providers did them a favor by providing them with the free software pieces they relied on to build their product.
Of course it's poor taste to add the "You're welcome". But poor taste is par for the course for such startup, unfortunately.
I meant it in a broader sense. Even for paid products, corporations view themselves as doing customers a favor. That's the sort of antagonistic relationship they define.
Despite interesting legal definitions, companies are just groups of people.
I also read the title as having an air of self-importance or maybe passive-aggression. If you read it a little more dispassionately then it's just an awkward way of stating that this is a gift.
Why is this a gift? The company does not need to give it to you, though it arguably benefits from doing so. It did so with some definition of "free."
Wow 18 hours in and they haven't changed the title of their post. What jerks.
We're terrible comedians - we've added something of a mea culpa to the post!
I looked at the dependencies and noticed that it wasn’t using an existing Protobuf parser which means they reimplemented the parser from scratch. Perhaps due to a lack of error recovery in the existing implementations? I don’t have the energy right now for further inspection.
It is definitely best to reuse the parser for the runtime when implementing an LSP but to do so properly means implementing the parser itself as a standalone library. Even better is shipping the semantic analysis as well!
Implementation drift is definitely an issue.
But great project anyways, just wanted to put my thoughts on the matter into the conversation!
I think its worth considering but I don’t agree it’s always best to use the language parser.
A language parser should be correct. A LSP parser should be fault tolerant. My understanding is you cant have both.
It absolutely can be done [1] but it is a lot of effort to do high quality error recovery. It's easier in languages with natural "synchronization points" [2], harder in languages which don't have them.
Having built a lot of protobuf tooling, I'd estimate that protobuf largely falls into the former camp; most of your time working with .proto files is operating on fields which terminate using `;` at the end of the line (though multi-line is also possible).
[1] source: I've done it for SQLite SQL at https://github.com/LalitMaganti/syntaqlite/
[2] e.g. SQL naturally has this at statement and expression boundaries which covers almost all of the cases people care about.
Adding to this, rustc's parser is both correct and fault-tolerant, yet rust-analyzer ended up building its own for different reasons (needing a CST and not an AST).
I settled on starting with a CST using rowan, doing semantic analysis, having an agnostic editor services, exposing both an LSP interface with tower, and a WASM interface (for Monaco in the browser [non-LSP!!]) and having the runtime binary doing the interpretation/compilation with the same pathway, with a slight jog, as the editor services. Phew!
It is my understanding that basically all LSP implementations use tree-sitter because of its incremental rebuild capability, and this requires re-implementing the parser.
That's just not true. rust-analyzer does not use tree-sitter. Neither does gopls. Nor Pylance. Or tsc. In fact I don't know if there is a single popular LSP that uses it (but there probably is).
Many editors use tree-sitter, but that is separate from the LSP.
buf is a parser / compiler as an alternative to protoc https://buf.build/docs/migration-guides/migrate-from-protoc/
Ah, so does the LSP just call out to the Buf binary?
The LSP is one of the subcommands provided by the buf binary. My LSP configuration for protobuf is `buf lsp serve`.
Makes sense, thanks for the clarification.
The lsp is a subcommand of buf, from the page: `buf lsp serve`
Lots of naysayers here, but an advantage of protobuf is that proto files are hand-writeable, and therefore having an LSP for that could be useful.
That said, proto itself dissuades or forbids the kind of common things you might do with a LSP, such as renaming.
Renaming fields is a big no-no [edit: this isn't true, see corrections below], as is doing things like re-ordering fields.
A core idea of proto is that versions are strictly compatible with with previous versions. This itself has limitations and challenges for migrations, but encourages good practice about compatibility that usually gets ignored or hand-waved away in most ecosystems.
I accept however that it's often easy to offload both the re-structuring and the checking of version compatibility to an LLM and let them go at it.
Actually renaming and reordering fields is completely fine, as long as the field ID and type stay the same
Considered breaking change because JSON serialization will change, and JSON is used quite often.
Reordering fields without changing their ID is indeed a non-breaking change from what I understand, albeit quite a useless one :)
> a useless one
If you have a (long) list of fields and you want to keep them lexicographically sorted, being able to reorder is quite useful if you rename a field.
Renaming is forbidden though (because JSON and textproto). In Google, it's a documented antipattern to try to make protobuf look "nice" by changing field indices, rearranging fields, etc. — the common ground is that it's better to not do it.
As long as you know the use cases of your fields, renaming is just fine. My team regularly does it. We also maintain our own serializer and deserializer json, XML, and fixed with formats. The json one we handle serialization using an annotation to say how it should be exported.
Of course, the whole concept of a breaking change does not really apply if all usage is within the controlled code. The problems start to appear when you have external users with old versions; then protobuf starts having a bunch of weird limitations. My favorite is that it's forbidden to move a field into or out of a `oneof`: it's a compatible change in a sense of wire format and JSON, but breaks the generated Golang code.
Breaking changes matter even with controlled code, because you can have requests that straddle upgrade boundaries, it's fiction to believe that all services are upgraded at the same moment, and pretending that is the case is the sort of thing that leads to quiet data corruption or mysterious bugs that can never seem to get reproduced.
Even if you upgrade with coordinated downtime across your entire service stack ( a bit old-school, but still happens more than you might imagine. ), then you still have to occasionally deal with requests that get persisted somewhere, possibly for support purposes, and it's much handier if the wire format remains compatible, at least between immediate versions.
There is a lot of ambiguity in this thread.
Yes, breaking changes need to be staged carefully/compatibly across versions.
Simple renaming/reordering is not a breaking change between code versions. The wire format knows numeric ids, not names. It breaks code compilation until the renames are put into effect. There is a subtle breakage where someone renames a field (think: field -> old_field) and then later adds "field" to mean something else; software that isn't recompiled in this window might not recognize that "field" is potentially different. Uncompiled languages may suffer this even more subtly. -- All of this to say, breaking compilation is not the end of the world but there are more dragons as the scope grows.
Stop-the-world is usually only needed when someone has made an unplanned/incompatible change with versions that are still running. If your infrastructure+development is done right (hah) this should never happen.
Depends, if you use textproto then renaming is problematic too
Shit, you're right, I'm trying to remember the one that used to trip us up all the time, maybe it was removing fields?
There was definitely one that kept tripping up the checks and it was something people like to do.
renaming is actually pretty fine, if you don't do stuff like json or text encodings. only renumbering fields causes problems.
"You're welcome." is such a smug verbal tic.
Protobuf getting LSP support before Codex implements LSPs is wild.
What would it mean for codex to "implement LSPs"
LSP server support, like Claude Code does.
While not a direct competitor to protobufs, if you are working in the video game space where struct versioning is not needed, there is an alternative language called "schema" that supports C, C++, C#, Golang, Rust and JavaScript.
https://github.com/mas-bandwidth/schema
Of all names, they pick schema? That's like calling a programming language "language."
My favorite programming language is called "A Programming Language".
Reminds me of xkcd tattoo that says in Chinese, "It's what my tattoo says."
> video game space where struct versioning is not needed
Save files? Looser than exact version multiplayer?
Multiplayer games typically deploy both client and server at the same time, and refuse to connect a client if it doesn't speak the exact same protocol as the server.
Thus all the versioning overhead of protobufs is not needed for this wire protocol.
(Yes, games still use versioning everywhere else where it makes sense: save games, asset data, config etc...)
Good luck deploying a client to the mobile app stores together with your backend :pain:
But then again, those barely count as games, I guess.
Wouldn’t it make sense to also have the old version of the game live and gradually roll out the new version for mobile
You can do this but if it's too granular (like you have no concept of version compatibility) then it can heavily split your matchmaking.
Plus the headaches of keeping many out of date builds up to date enough to deploy.
Even if you don't care about in game compatibility, all your servers still talk to some centralized data store and that will likely want a single deploy that handles old clients
We do it just fine (not for games). You submit a new binary for review in advance and then do a coordinated release once the review completes.
Then players are locked out of multiplayer until they download a possibily large content patch.
This already happens on PC games too.
I think both iOS and Android support separate downloads for large assets?
This feels like normal on mobile? Games force you to update to keep playing.
Checked the contributor list, please disclose that you're a primary contributor.
Buf does great work at fixing Protobuf to the point of being just barely usable. A godsend if you're stuck with Protobuf/gRPC on a legacy project.
What world are you living in where Protobuf/gRPC are considered "legacy"? ..what?
The promise of Protobuf/gRPC for seamless communication across tech stacks just wasn't delivered. Third party projects like Buf or BetterProto for Python try their best but it's still a nightmare trying to integrate with gRPC from something remotely modern, like Python with uv and async/await. Plus there's a lot of little pains like not being human-readable, always requiring HTTP/2 and thus HTTPS, and so on.
JSON isn't "optimal" but after transfer compression it's not that bad, and the support for JSON Schema is much better across the stacks I use, e.g. Zod in JS/TS, Pydantic in Python, etc. And it's fully usable "by hand" without having the schema, for one-off scripts and such -- compare with Protobuf where you need the full definition to even parse a chunk of data.
In the web dev world, I’ve noticed there are a lot of people who think that JSON is the ultimate modern data format. Typically those same people have barely heard of JSON Schema.
This has been driving me crazy ever since I started using protobuf/grpc and realized more major tech companies (in the cloud/infra/data world at least, and a lot of other SAAS) were using it or something similar (eg capn proto) internally than not.
It feels like we’re in some sort of deadlock where each of them think “proto/grpc are too niche to support for external users, better just use JSON”, keeping it unfamiliar for an Average Web Developer. But if every company using it just exposed it to third parties/added it to their public APIs, it would immediately be common (and trendy) enough for every web developer to learn it and start using it.
If Google added the missing HTTP/2 streaming support to browsers (blocking native bidi grpc streaming) it would have an instant killer app in making it easier to implement websocket-like client/server applications. It makes absolutely no sense that full duplex bidi was added to the HTTP/2 but remains unimplemented in browsers.
The role MCP, OpenAPI, and JSON schema fill all would be a million times simpler if they were based on protobuf instead of JSON. I can forgive OpenAPI/JSON but it honestly pisses me off that we ended up with MCP and JSON-RPC + JSON schema, and people think these are cool/good tools, and actively adopting them. Just piling on the slop
I use Go + Buf + protobuf-es + Connect on my web projects. There is no better stack, I'll die on this hill.
I pushed for us to adopt this setup at my previous company and it was wildly successful. I thought frontend devs would resist but they absolutely loved the generated types and clients. It quickly seemed silly to do anything else.
This is the way! It's great to just point frontend developers to the Buf documentation for new features and they easily have all the types and generated clients available in their projects too.
Unpopular opinion -- LSPs are bad because they introduce latency into coding. For every keystroke, my IDE needs to make request/response with LSP, instead of using its own internal parser/colorer/autocomplete.
LSP isn't used for syntax highlights usually, that's still the job of the editor. This is where something like treesitter usually comes in. LSP is only used in this case for errors/annotations/etc.
You can just not use an LSP? My understanding is that LSPs aren't made to be the fastest (code completion|syntax highlighter|code analysis), but easily integrated into all sorts of IDEs/text editors.
As far as I understand, IDEs ara way less eager to spend resources onto their own integration if there's an LSP. E.g. last time I tried to code in Zig there was either LSP or nothing.
Maybe if you have an extremely slow computer?
Even my little chromebook can provide completions and errors effectively instantly to me, even in medium sized rust projects.
Or are you saying (probably maximum) ~20ms delay is blocking your ability to code?
It's not 20ms, 1-2 seconds for sure, for Zig LSP.
96GB RAM, 12 vCPU, Linux, desktop.
I've always thought of protobuf as a Java thing.
It worked great in Java, but not so well in JS, especially with Kafka.
Protobuf was never particularly a Java thing. You could more accurately say it was a C++ thing. It was initially developed in C++, and used mainly in C++ services at Google. The implementation and early ecosystem were C++-centric.
Later, Go became one of the major Protobuf ecosystems, and today it would be understandable to think it was a Go thing.
Watch as I don't use protobuf because it is horrible.
....
Tada!
If you patch clients to google services in Python to use json instead of grpc they get faster and more reliable. A lot faster. Benchmark it!
For me that is how I know something like protobuf is good. It is a nuisance to manage and distribute the definitions, adds a build step even to languages with no build step normally, is slower than almost every alternative, and artificially restricts you from doing lots of common things. It's so good!And look at the code quality of the implementation! It's like a team of interns wrote it while drunk. It is a complete spaghetti mess, but has tons of super convoluted micro optimizations that are slower than just doing the most obvious thing, but make the implementation confusing and indirect. It's trash code.
What are good alternatives when you need a common "single source of truth" schema shared between multiple languages? We use protobuf between c# and Python.
Json schema?
I quite like the look of Typespec though I haven't used it much.
I always thought Thrift was waaay better than any of the alternatives, but it always had terrible documentation and I think it died mainly because of that.
The goal is not to have a single source of truth schema. That is a means to some other goal, and it's not even a good means.
If you never change the schema then you don't have to worry about it, get things working and never look back.
If you do change your schema from time to time, you need testing between the two systems. If you have good tests again a single source of truth is fully redundant, both systems are talking just fine. If you don't have tests things can and will break all the time even using protobuf.
> The goal is not to have a single source of truth schema. That is a means to some other goal, and it's not even a good means.
It’s about data transmission. Being able to encode and decode in a type safe manner between different languages (and so, different platforms) is a goal that makes a lot of sense.
> If you do change your schema from time to time, you need testing between the two systems
Or you could just use a defined format that doesn’t require testing. I rarely use protobuf but I can see why people do. The guaranteed backwards compatibility is huge for people who can’t just publish a new web frontend at the drop of a hat.
afavour it's been a long day already, maybe I can explain why you still need tests another day.
You could always try being less condescending when replying to users and you might find there was no need for anyone to reply in the first place.
If you understand how to evolve protobuf schema definitions, then you don’t really need testing. You instinctively know how the parser works when it is parsing data with a different schema from what it expects. And that’s a powerful thing. If your things break even when using protobuf then you don’t grok protobuf.
It’s probably not an exaggeration to say that being able to avoid tests between different systems who have different versions of the schema is a core goal of protobuf. Why? These two different systems are probably owned by different teams, and introducing explicit tests between different versions of them increases coupling between them.
"you don't really need testing"
Protobuf only allows you to add optional fields after the initial version, right? Because otherwise it would not be backwards compatible. Protobuf does not allow you to define contingent logic between fields. Optional fields are always nullable (or you must provide a default). This forces you to know all of this and use a method to see if it was actually set or just defaulted. So you have a ton of nullable/defaulted fields that likely are required to be filled in or not filled in, contingent on other values in the same struct. For example if the charge type is "purchase" then price must be not null and > 0. That sort of thing. This is so common you should just assume your app has a million little rules like this that are assumed. Protobuf does not help you here at all. You need some other logic to validate the data on top of protobuf. once you have that anyway protobuf's value is that it's expensive, requires build steps, is actually not fast, and forces you to distribute the schema between different apps somehow.
If it's not obvious yet, lets say you are sending your charge structs and you have a bug where sometimes you don't set price. It's defaulted to 0 or -1 or whatever nonsense value the default is, or null, it doesn't matter it's not correct sometimes. That is why you need tests my guy. Protobuf can't fix this. If you use json the tests make sure everything protobuf does for you is done too. In a world where you have to write tests because protobuf can't force you to correctly set fields, you have tests already, and protobuf didn't help you at all. Whether you call it tests or input validation or whatever, protobuf definitions are insufficient, and when you have what is sufficient it 100% covers everything protobuf does 'for you'.
If you don't get it at this point then lets just agree that you will never get it.
> Protobuf only allows you to add optional fields after the initial version, right?
That’s not all. For example, you can also change fields from optional to repeated or vice versa. For another example, you can also delete fields.
> Protobuf does not allow you to define contingent logic between fields.
You are just saying that protobuf is not a data validation library. That’s true; it only handles data serialization. You need to write validations yourself. And you will most likely need tests to test your validation code. But those are entirely tests within a single process; they are not tests involving a producer and a consumer of a protobuf message, which is the wrong kind of tests.
> For example if the charge type is "purchase" then price must be not null and > 0. That sort of thing.
That sort of thing sounds like you are not using `oneof` appropriately. The charge type shouldn’t be a field. It should be a submessage called Purchase.
Both sibling comments say one type of assurance makes the other irrelevant, but I would wager they cover different territory.
Is that a recent-ish improvement? I feel like HTTP/2 would be roughly the same performance for JSON and protobuf, so maybe this is HTTP/2 vs HTTP/3?
I think the overhead is protobuf itself but I can't check.
Comparing a wrapped C++ gRPC backed stack with an httpx/requests backed one is like comparing apples to elephants.
Protobuf simply encode things way more efficient that JSON can define a single object. You're quite frankly spewing bullshit in this whole thread.
You don’t get what they say. It’s not about about how efficient it is after encode, it’s about how fast encode is. They are not spewing bs, they’re focusing on a single point. The question is: do you send it over the wite more often than performing encode/decode.
That's what the person you replied to is talking about, and they're right. Putting aside the final byte size (where protobuf also wins), protobuf is faster at both encoding and decoding than json. There are numerous benchmarks you can find that show this.
The advantages of json are not related to performance.
doesn't seem universally true, https://github.com/protobufjs/protobuf.js/issues/2114#issue-...
If you're writing JS you cannot beat JSON.parse, because you're running the most optimized C++ implementation of JSON which will outcompete any decoder written JS itself.
Which is not a very generalizable situation.
Ok but that's the point. You are just explaining why the 'protobuf is faster' claim is generally false.
It is generally correct. When both implementations are completely in the same language Protobuf will win. Instead of a pure-JS implementation one can also make a FFI module for Protobuf.
The comparison that's being made is equivalent to implementing quick sort in Python and bubble sort in C++ and then declaring bubble sort is the better alternative.
Anyone with a healthy understading of computer science and programming experience will not make bullshit claims like that.
No, protobuf is not faster, not on Python, because google's implementation is poor. Python is the #1 language in the world right now. Don't use protobuf.
That's not the lesson I'm taking away from this. Imagine how much faster the code would be if you didn't use Python.
Is this because load is lighter on their JSON endpoints that their gRPC ones?
so how do you save data over the cable when it's needed?
Eventually it’s all bytes. Where do you want to save them?
zstd?? Obviously?
JSON. /s
Buf's offering of protobuf registries and codegen SDKs for microservices seems less necessary in the LLM era.
I'm starting to question many of protobuf's advantages (perhaps not the wire format). Add to that monorepos and other fads of the 2010s given the rise of LLMs.
I used to be a big believer in this stuff, but I'm quickly having my core assumptions change out from under me.
A big differentiator is whether one imagines an LLM in-band with most/all future software. If there is, and we’re deferring until very late parts of a program that world have been load bearing, and we’re able to programmatically ands reliably squint and say “eh i know what you meant” … then yeah formalizations seem superfluous-to-counterproductive.
OTOH if LLMs are to write, but not supplant, much of software, then boundaries, delegation to deterministic layers, good compilers to bonk miscreant models on the head with error message seem essential.
At one point it would have been shocking to assert that the compiler would live in-band with the program too. and yet JS eats the world. It seems shocking today that we could have a universal prior over the world operating in the ms/us nJ/pJ range required. And yet … ?
The real argument shouldn't be about protocols becoming obsolete, but programming languages that are "less efficient" could eventually become obsolete in favor of highly scalable and performant languages due to LLMs when the main gap becomes knowing how the tech works at a high level, and not the syntax, why code in one language over another if you don't need to worry about messing up on syntax, only about reviewing logic for sanity and correctness against business rules as well as validating that it is stable code.
Anybody who thinks that you can just chuck unstructured data into LLM and YOLO the app is an idiot.
This works up to a point, and then it doesn't. And you're left with tons of inconsistently formatted data.
My company is built on protobufs from ground up :) We use it in the database, for remote calls, on the frontend, etc. The protobuf language is not great, but it's about the right balance between too expressive and too restricting.
And the best thing is that it's compact, compared to OpenAPI.
You're wrong on every count: given the small context window of LLMs, the need for interfaces and scehmas is even greater.
I feel exactly that way about REST. The assumption that LLMs make schemas obsolete misses how structured outputs actually work in production. When you have probabilistic models generating code, strict contracts become more critical, not less. It is no coincidence that several major LLM platforms rely on ConnectRPC and Protobuf for their own APIs.
You're right, but the adeptness of models to spin up clients and behaviors on the fly is remarkable. They're capturing the semantics of behavior at a deeper level.
If we do strict schemas, I'd like to see less ceremony around them. Tool calls instead of brittle build steps and protocol registries.
Perhaps we need new tools for this going forward.
Hm... Maybe. In my view, an IDL is part of the input that you absolutely want humans to author or carefully review at least. In my experience, the ceremony around generating code is also performed very well by LLMs. But I do agree, there's definitely some changes that are needed to integrate Protobufs better. Some languages have built-in tooling to make it seamless, but it's definitely not universal.
why? are humans going to stop using services?
From protobuf to quantum computing, Google just keeps solving problems nobody has...
That's not Google. Buf.build is, hilariously enough, a completely separate company that tries to make sense of Google's protobuf, and it's doing a decent job from what I can see.
so you're saying protobuf has no purpose?
i mean...
Protobuf should at least get credit for enforcing fields be optional.
They reluctantly acknowledged that as a mistake by adding an actual "optional" modifier for field presence. proto3's previously so-called optional were really all present with a zero default value; proto2 could have supported that for required fields at the code level if they had wanted. In the end, all they've changed is how they frame the feature.
Aren't LSP dying (and eventually IDE, at least in their current form) as everybody use LLM to code. I know some big tech companies redirected teams supporting them to new efforts (ie. tool integration with AI).