Protobuf has LSP support. You're welcome

(buf.build)

62 points | by theanonymousone 2 hours ago

6 comments

  • williamcotton 3 minutes ago
    I looked at the dependencies and noticed that it wasn’t using an existing Protobuf parser which means they reimplemented the parser from scratch. Perhaps due to a lack of error recovery in the existing implementations? I don’t have the energy right now for further inspection.

    It is definitely best to reuse the parser for the runtime when implementing an LSP but to do so properly means implementing the parser itself as a standalone library. Even better is shipping the semantic analysis as well!

    Implementation drift is definitely an issue.

    But great project anyways, just wanted to put my thoughts on the matter into the conversation!

  • eterm 7 minutes ago
    Lots of naysayers here, but an advantage of protobuf is that proto files are hand-writeable, and therefore having an LSP for that could be useful.

    That said, proto itself dissuades or forbids the kind of common things you might do with a LSP, such as renaming.

    Renaming fields is a big no-no, as is doing things like re-ordering fields.

    A core idea of proto is that versions are strictly compatible with with previous versions. This itself has limitations and challenges for migrations, but encourages good practice about compatibility that usually gets ignored or hand-waved away in most ecosystems.

    I accept however that it's often easy to offload both the re-structuring and the checking of version compatibility to an LLM and let them go at it.

    • arccy 0 minutes ago
      renaming is actually pretty fine, if you don't do stuff like json or text encodings. only renumbering fields causes problems.
    • wazzaps 4 minutes ago
      Actually renaming and reordering fields is completely fine, as long as the field ID and type stay the same
  • gafferongames 40 minutes ago
    While not a direct competitor to protobufs, if you are working in the video game space where struct versioning is not needed, there is an alternative language called "schema" that supports C, C++, C#, Golang, Rust and JavaScript.

    https://github.com/mas-bandwidth/schema

    • satvikpendem 20 minutes ago
      Of all names, they pick schema? That's like calling a programming language "language."
      • max-privatevoid 11 minutes ago
        My favorite programming language is called "A Programming Language".
        • andai 7 minutes ago
          Reminds me of xkcd tattoo that says in Chinese, "It's what my tattoo says."
  • ltbarcly3 1 hour ago
    Watch as I don't use protobuf because it is horrible.

    ....

    Tada!

    If you patch clients to google services in Python to use json instead of grpc they get faster and more reliable. A lot faster. Benchmark it!

      def get_json_client() -> CloudLoggingQueryClient:
          """Client for log queries (JSON transport, avoids gRPC overhead)."""
          client = google.cloud.logging.Client.from_service_account_info(...)
          client._use_grpc = False
          return client
    
    
    For me that is how I know something like protobuf is good. It is a nuisance to manage and distribute the definitions, adds a build step even to languages with no build step normally, is slower than almost every alternative, and artificially restricts you from doing lots of common things. It's so good!

    And look at the code quality of the implementation! It's like a team of interns wrote it while drunk. It is a complete spaghetti mess, but has tons of super convoluted micro optimizations that are slower than just doing the most obvious thing, but make the implementation confusing and indirect. It's trash code.

    • tomtom1337 52 minutes ago
      What are good alternatives when you need a common "single source of truth" schema shared between multiple languages? We use protobuf between c# and Python.
      • throw1234567891 2 minutes ago
        Json schema?
      • ltbarcly3 49 minutes ago
        The goal is not to have a single source of truth schema. That is a means to some other goal, and it's not even a good means.

        If you never change the schema then you don't have to worry about it, get things working and never look back.

        If you do change your schema from time to time, you need testing between the two systems. If you have good tests again a single source of truth is fully redundant, both systems are talking just fine. If you don't have tests things can and will break all the time even using protobuf.

        • andai 2 minutes ago
          Both sibling comments say one type of assurance makes the other irrelevant, but I would wager they cover different territory.
        • afavour 40 minutes ago
          > The goal is not to have a single source of truth schema. That is a means to some other goal, and it's not even a good means.

          It’s about data transmission. Being able to encode and decode in a type safe manner between different languages (and so, different platforms) is a goal that makes a lot of sense.

          > If you do change your schema from time to time, you need testing between the two systems

          Or you could just use a defined format that doesn’t require testing. I rarely use protobuf but I can see why people do. The guaranteed backwards compatibility is huge for people who can’t just publish a new web frontend at the drop of a hat.

        • kccqzy 18 minutes ago
          If you understand how to evolve protobuf schema definitions, then you don’t really need testing. You instinctively know how the parser works when it is parsing data with a different schema from what it expects. And that’s a powerful thing. If your things break even when using protobuf then you don’t grok protobuf.

          It’s probably not an exaggeration to say that being able to avoid tests between different systems who have different versions of the schema is a core goal of protobuf. Why? These two different systems are probably owned by different teams, and introducing explicit tests between different versions of them increases coupling between them.

    • onei 51 minutes ago
      Is that a recent-ish improvement? I feel like HTTP/2 would be roughly the same performance for JSON and protobuf, so maybe this is HTTP/2 vs HTTP/3?
      • ltbarcly3 47 minutes ago
        I think the overhead is protobuf itself but I can't check.
        • okanat 15 minutes ago
          Protobuf simply encode things way more efficient that JSON can define a single object. You're quite frankly spewing bullshit in this whole thread.
    • pastel8739 54 minutes ago
      Is this because load is lighter on their JSON endpoints that their gRPC ones?
    • itsthecourier 58 minutes ago
      so how do you save data over the cable when it's needed?
  • echelon 2 hours ago
    Buf's offering of protobuf registries and codegen SDKs for microservices seems less necessary in the LLM era.

    I'm starting to question many of protobuf's advantages (perhaps not the wire format). Add to that monorepos and other fads of the 2010s given the rise of LLMs.

    I used to be a big believer in this stuff, but I'm quickly having my core assumptions change out from under me.

    • kristjansson 38 minutes ago
      A big differentiator is whether one imagines an LLM in-band with most/all future software. If there is, and we’re deferring until very late parts of a program that world have been load bearing, and we’re able to programmatically ands reliably squint and say “eh i know what you meant” … then yeah formalizations seem superfluous-to-counterproductive.

      OTOH if LLMs are to write, but not supplant, much of software, then boundaries, delegation to deterministic layers, good compilers to bonk miscreant models on the head with error message seem essential.

      At one point it would have been shocking to assert that the compiler would live in-band with the program too. and yet JS eats the world. It seems shocking today that we could have a universal prior over the world operating in the ms/us nJ/pJ range required. And yet … ?

    • giancarlostoro 1 hour ago
      The real argument shouldn't be about protocols becoming obsolete, but programming languages that are "less efficient" could eventually become obsolete in favor of highly scalable and performant languages due to LLMs when the main gap becomes knowing how the tech works at a high level, and not the syntax, why code in one language over another if you don't need to worry about messing up on syntax, only about reviewing logic for sanity and correctness against business rules as well as validating that it is stable code.
    • cyberax 1 hour ago
      Anybody who thinks that you can just chuck unstructured data into LLM and YOLO the app is an idiot.

      This works up to a point, and then it doesn't. And you're left with tons of inconsistently formatted data.

      My company is built on protobufs from ground up :) We use it in the database, for remote calls, on the frontend, etc. The protobuf language is not great, but it's about the right balance between too expressive and too restricting.

      And the best thing is that it's compact, compared to OpenAPI.

    • sudorandom 1 hour ago
      I feel exactly that way about REST. The assumption that LLMs make schemas obsolete misses how structured outputs actually work in production. When you have probabilistic models generating code, strict contracts become more critical, not less. It is no coincidence that several major LLM platforms rely on ConnectRPC and Protobuf for their own APIs.
      • echelon 1 hour ago
        You're right, but the adeptness of models to spin up clients and behaviors on the fly is remarkable. They're capturing the semantics of behavior at a deeper level.

        If we do strict schemas, I'd like to see less ceremony around them. Tool calls instead of brittle build steps and protocol registries.

        Perhaps we need new tools for this going forward.

        • sudorandom 1 hour ago
          Hm... Maybe. In my view, an IDL is part of the input that you absolutely want humans to author or carefully review at least. In my experience, the ceremony around generating code is also performed very well by LLMs. But I do agree, there's definitely some changes that are needed to integrate Protobufs better. Some languages have built-in tooling to make it seamless, but it's definitely not universal.