6 ms·
Show HN: Skir – like Protocol Buffer but better
Why I built Skir: https://medium.com/@gepheum/i-spent-15-years-with-protobuf-then-i-built-skir-9cf61cc65631 https://medium.com/@gepheum/i-spent-15-years-with-protobuf-t...
Quick start: npx skir init
All the config lives in one YML file.
Website: https://skir.build https://skir.build
GitHub: https://github.com/gepheum/skir https://github.com/gepheum/skir
Would love feedback especially from teams running mixed-language stacks.
- poly2it 7mo agoI would recommend exploring OpenRPC for those who have not yet seen it. It brings protocol-buffer-like definitions (components), RPC definitions and centralised error definitions.
- vvern 7mo agoNotably missing both Go and Rust
- gepheum 7mo agoAbsolutely, I am planning to add these 2 as well as C# by June. Working on it now.
- rubenvanwyk 7mo agoDo you have a newsletter or how should we know when C# will be available?
- dewey 7mo ago> Skir is a universal language for representing data types, constants, and RPC interfaces. Define your schema once in a .skir file and generate idiomatic, type-safe code in TypeScript, Python, Java, C++, and more. Maybe I'm missing some additional features but that's exactly what https://buf.build/plugins/typescript https://buf.build/plugins/typescript does for Protobuf already, with the advantage that you can just keep Protobuf and all the battle hardened tooling that comes with it.
- deleted 7mo ago[deleted]
- nine_k 7mo agoThe entire original post, it seems, is dedicated to explaining why Skir is better than plain Protobuf, with examples of all the well-known pain points. If these are not persuasive for you, staying with Protobuf (or just JSON) should be a fine choice.
- waynesonfire 7mo ago[flagged]
- nine_k 7mo agoIf you are fine enough with protobufs so that you're not actively looking for alternatives, maybe you should not spend the effort.
- gepheum 7mo ago+1 Copying from blog post [https://medium.com/@gepheum/i-spent-15-years-with-protobuf-then-i-built-skir-9cf61cc65631?postPublishedType=repub https://medium.com/@gepheum/i-spent-15-years-with-protobuf-t...]: """ Should you switch from Protobuf? Protobuf is battle-tested and excellent. If your team already runs on Protobuf and has large amounts of persisted protobuf data in databases or on disk, a full migration is often a major effort: you have to migrate both application code and stored data safely. In many cases, that cost is not worth it. For new projects, though, the choice is open. That is where Skir can offer a meaningful long-term advantage on developer experience, schema evolution guardrails, and day-to-day ergonomics. """
- gepheum 7mo agoSkir has exactly the same goals as Protobuf, so yes, that sentence can apply to Protobuf as well (and buf.build). I listed some of the reasons to prefer Skir over Protobuf in my humble opinion here: https://medium.com/@gepheum/i-spent-15-years-with-protobuf-then-i-built-skir-9cf61cc65631?postPublishedType=repub https://medium.com/@gepheum/i-spent-15-years-with-protobuf-t... Built-in compatibility checks, the fact that you can import dependencies from other projects (buf.build offers this, but makes you pay for it), some language designs (around enum/oneof, the fact that adding fields forces you to update all constructor code sites), the dense JSON format, are examples.
- drathier 7mo agohttps://capnproto.org/ https://capnproto.org/ has been my goto since forever. Made by the protobuf inventor
- k_g_b_ 7mo ago*Made by the proto2 implementor, Kenton Varda https://news.ycombinator.com/user?id=kentonv https://news.ycombinator.com/user?id=kentonv
- jauntywundrkind 7mo agoI've been dabbling with the newer Cap'n Web, whose nicely descriptive README's first line says: > Cap'n Web is a spiritual sibling to Cap'n Proto (and is created by the same author), but designed to play nice in the web stack. It's just JSON, which has up and down sides. But things like promise pipelining are such a huge upside versus everything else: you can refer to results (and maybe send them around?) and kick off new work based on those results, before you even get the result back. This is far far far superior to everything else, totally different ball-game. I've been a little rebuffed by wasm when I try, keep getting too close to some gravitational event horizon & get sucked in & give up, but for more data-throughput oriented systems, I'm still hoping wrpc ends up being a fantastic pick. https://github.com/bytecodealliance/wrpc https://github.com/bytecodealliance/wrpc . Also Apache Arrow Flight, which I know less about, has mad traction in serious data-throughput systems, which being adjacent to amazingly popular Apache Arrow makes sense. https://arrow.apache.org/docs/format/Flight.html https://arrow.apache.org/docs/format/Flight.html
- jeffbee 7mo agoObligatory dense field numbers seems like a massive downside, the problems of which would become evident after a busy repo has been open for a few days.
- gepheum 7mo agoIt's not obligatory. Basically Protobuf gives you a choice between (1) binary format, (2) readable JSON. Skir gives you a choice between (1) binary format, (2) readable JSON, (3) dense JSON. It recommends dense JSON as the "default choice", but it does not force it. The reason why it's recommended as the default choice is because it offers a good tradeoff between efficiency (only a bit less compact than binary), backward compatibility (you can rename fields safely unlike readable JSON) and debuggability (although it's definiely not as good as readable JSON, because you lose the field numbers, it's decent and much better than binary format)
- refulgentis 7mo agoI spent some time in the actual compiler source. There's real work here, genuinely good ideas. The best thing Skir does is strict generated constructors. You add a field, every construction site lights up. Protobuf's "silently default everything" model has caused mass production incidents at real companies. This is a legitimately better default. Dense JSON is interesting but the docs gloss over the tradeoff: your serialized data is [3, 4, "P"]. If you ever lose your schema, or a human needs to read a payload in a log, you're staring at unlabeled arrays. Protobuf binary has the same problem but nobody markets binary as "easy to inspect with standard tools." The "serialize now, deserialize in 100 years" claim has a real asterisk. Compatibility checking requires you to opt into stable record IDs and maintain snapshots. If you skip that (and the docs' own examples often do), the CLI literally warns you: "breaking changes cannot be detected." So it's less "built-in safety" and more "safety available if you follow the discipline." Which is... also what Protobuf offers. The Rust-style enum unification is genuinely cleaner than Protobuf's enum/oneof split. No notes there, that's just better language design. Minor thing that bothered me disproportionately: the constant syntax in the docs (x = 600) doesn't match what the parser actually accepts (x: 600). The weirdest thing that bugged the heck out of me was the tagline, "like protos but better", that's doing the project no favors. I think this would land better if it were positioned as "Protobuf, but fresh" rather than "Protobuf, but better." The interesting conversation is which opinions are right, not whether one tool is universally superior. Quite frankly, I don't use protobuf because it seems like an unapproachable monolith, and I'm not at FAANG anymore, just a solo dev. No one's gonna complain if I don't. But I do love the idea of something simpler thats easy to wrap my mind around. That's why "but fresh" hits nice to me, and I have a feeling it might be more appealing than you'd think - ex. it's hard to believe a 2 month old project is strictly better than whatever mess and history protobufs gone through with tons of engineers paid to use and work on it. It is easy to believe it covers 99% of what Protobuf does already, and any crazy edge cases that pop up (they always do, eventually :), will be easy to understand and fix.
- vineyardmike 7mo ago> Minor thing that bothered me disproportionately: the constant syntax in the docs (x = 600) doesn't match what the parser actually accepts (x: 600). You’re a better man than me. If the docs can’t even get the syntax right, that’s a hard no from me. Also, fwiw, you’ve got a few points wrong about protos. Inspecting the binary data is hard, but the tag numbers are present. You need the schema, but at least you can identify each element. Also, I disagree on the constructor front. Proto forces you to grapple with the reality that a field may be missing. In a production system, when adding a new field, there will be a point where that field isn’t present on only one side of the network call. The compiler isn’t saving you. Fresh is more honest than better, and personally, I wouldn’t change it.
- nazgu1 7mo agoIf I may suggest, Swift support will be more than appreciated, to consider it for a viable protocol for connecting backend with mobile applications.
- gepheum 7mo agoYeah totally fair. I targeted Dart because of Flutter, but I think I will include Swift in the next wave of languages, after Rust, Go and C#.
- chocolatkey 7mo agoThat “compact JSON” format reminds me if the special protobufs JSON format that Google uses in their APIs that has very little public documentation. Does anyone happen to know why Google uses that, and to OP, were you inspired by that format?
- gepheum 7mo agoI think you may be referring to JSPB. It's used internally at Google but has little support in the open-source. I know about it, but I wouldn't say I was inspired by it. It's particularly unreadable, because it needs to account for field numbers being possible sparse. Google built it for frontend-backend communication, when both the frontend and the backend use Protobuf dataclasses, as it's more efficient than sending a large JSON object and also it's faster to deserialize than deserializing a binary string on the browser side. I think it's mostly deprecated nowadays.
- woadwarrior01 7mo agoAlso, flexbuffers.
- kevincox 7mo agoI don't know but if I had to guess. 1. Google uses protobufs everywhere, so having something that behaves equivalently is very valuable. For example in protobuf renaming fields is safe, so if they used field names in the JSON it would be protobuf incompatible. 2. It is usually more efficient because you don't send field names. (Unless the struct is very sparse it is probably smaller on the wire, serialized JS usage may be harder to evaluate since JS engines are probably more optimized for structs than heterogeneous arrays). 3. Presumably the ability to use the native JSON parsing is beneficial over a binary parser in many cases (smaller code size and probably faster until the code gets very hot and JITed).
- rivetfasten 7mo agoI don't know the reason TextFormat was invented, but in practice it's way easier to work with TextFormat than JSON in the context of Protos. Consider numeric types - JSON: number aka 64-bit IEEE 754 floating point Proto: signed and unsigned int 8, 16, 32, 64-bit, float, double I can only imagine the carnage saved by not accidentally chopping of the top 10 bits (or something similar) of every int64 identifier when it happens to get processed by a perfectly normal, standards compliant JSON processor. It's true that most int64 fields could be just fine with int54. It's also true that some fields actually use those bits in practice. Also, the JSPB format references tag numbers rather than field names. It's not really readable. For TextProto it might be a log output, or a config file, or a test, which are all have ways of catching field name discrepancies (or it doesn't matter). For the primary transport layer to the browser, the field name isn't a forward compatible/safe way to reference the schema. So oddly the engineers complaining about the multiple text formats are also saved from a fair number of bugs by being forced to use tools more suited to their specific situation.
- argon81 7mo agoHow does this compare with https://connectrpc.com/ https://connectrpc.com/ as that project seems to share similar goals
- elvin_d 7mo agoLike this but zero copy, easy migration/versioning, Rust and WASM support.
- lasgawe 7mo agoLooks nice. But what are the use cases of this? I'm still trying to figure that out.
- gepheum 7mo agoThanks! Main use case (similarly to Protobuf) is when you need to exchange data types between systems written in different languages. Like Protobuf, it can also be used in a mono-linguistic system, when you want to serialize systems and have strong guarantees that you will be able to deserialize your data in the future (when you use classic serialization libraries like Pydantic, Java Serialization etc., it's easy to accidentally modify a schema and break the ability to deserialize old data.)
- il-b 7mo agoImpressive. Some interop with established standards such as OpenAPI or gRPC would make it an easier sell for non-greenfield projects
- gepheum 7mo agoThanks. Definitely agree, will try to think about what that could look like.
- shalabhc 7mo agoDid you look at other formats like Avro, Ion etc? Some feedback: 1. Dense json Interesting idea. You can also just keep the compact binary if you just tag each payload with a schema id (see Avro). This also allows a generic reader to decode any binary format by reading the schema and then interpreting the binary payload, which is really useful. A secondary benefit is you never ever misinterpret a payload. I have seen bugs with protobufs misinterpreted since there is no connection handshake and interpretation is akin to 'cast'. 2. Compatibility checks +100 there's not reason to allow breaking changes by default 3. Adding fields to a type: should you have to update all call sites? I'm not so sure this is the right default. If I add a field to a core type used by 10 services, this requires rebuilding and deploying all of them. 4. enum looks great. what about backcompat when adding new enum fields? or sometimes when you need to 'upgrade' an atomic to an enum?
- gepheum 7mo agoThanks for the feedback. 0. Yes, I looked at Avro, Ion. I like Protobuf much better because I think using field numbers for field identity, meaning being able to rename fields freely, is a must. 1. Yes. Skir also supports that with binary format (you can serialize and deserialize a Skir schema to JSON, which then allows you to convert from binary format to readable JSON). It just requires to build many layers of extra tooling which can be painful. For example, if you store your data in some SQL engine X, you won't be able to quickly visualize your data with a simple SELECT statement, you need to build the tooling which will allow you to visualize the data. Now dense JSON is obviously not idea for this use case, because you don't see the field names, but for quick debugging I find it's "good enough". 3. I agree there are definitely cases where it can be painful, but I think the cases where it actually is helpful are more numerous. One thing worth noting is that you can "opt-out" of this feature by using `ClassName.partial(...)` instead of `ClassName()` at construction time. See for example `User.partial(...)` here: https://skir.build/docs/python#frozen-structs https://skir.build/docs/python#frozen-structs I mostly added this feature for unit tests, where you want to easily create some objects with only some fields set and not be bothered if new fields are added to the schema. 4. Good question. I guess you mean "forward compatibility": you add a new field to the enum, not all binaries are deployed at the same time, and some old binary encounters the new enum it doesn't know about? I do like Protobuf does: I default to the UNKNOWN enum. More on this: - https://skir.build/docs/schema-evolution#adding-variants-to-an-enum https://skir.build/docs/schema-evolution#adding-variants-to-... - https://skir.build/docs/schema-evolution#default-behavior-drop https://skir.build/docs/schema-evolution#default-behavior-dr... - https://skir.build/docs/protobuf#implicit-unknown-variant https://skir.build/docs/protobuf#implicit-unknown-variant
- curtisf 7mo ago> For optional types, 0 is decoded as the default value of the underlying type (e.g. string? decodes 0 as "", not null). In the "dense JSON" format, isn't representing removed/absent struct fields with `0` and not `null` backwards incompatible? If you remove or are unaware of a `int32?` field, old consumers will suddenly think the value is present as a "default" value rather than absent
- gepheum 7mo agoThat is correct and that is a good catch, the idea though is that when you remove a field you typically do that after having made sure that all code no longer read from the removed field and that all binaries have been deployed.
- joshuamorton 7mo agoHow does this work if, for example, you persist the data in a database?
- gepheum 7mo agoLet's imagine you have this: ``` struct User { id: int64; email: string?; name: string; } ``` You store some users in a database: [10,"john@gmail.com""john"], [11,"jane",null,"john@gmail.com"] You remove the email field later: ``` struct User { id: int64; name: string; removed; } ``` Supposedly you remove a field after you have migrated all code that uses the field and you have deployed all binaries. In your DB, you still have [10,john@gmail.com","john"], [11,null,"jane"], which you are able to deserialize fine (the email field is ignored). New values that you serialize are stored as [12,0,"jack"]. If you happen to have old binaries which still use the old email field and which are still running (which you shouldn't, but let's imagine you accidentally didn't deploy all your binaries before you removed the field), these new binaries will indeed decode the email field for new values (Jack) as an empty string instead of null.
- maxloh 7mo agoLooks really like Prisma to me: https://www.prisma.io/docs/orm/prisma-schema/overview#example https://www.prisma.io/docs/orm/prisma-schema/overview#exampl... Why build another language instead of extending an existing one?
- gepheum 7mo agoI looked at Prisma, I very much prefer the Protobuf/Thrift model of using numbers to identify fields, which allows 2 important things: fields to be renamed without breaking backward compatibility, and a compact wire format. I think the Protobuf language (which Skir is heavily influenced by) has some flaws in its core design, e.g. the enum/oneof mess, the fact that it allows spare field numbers which makes the "dense JSON" format (core feature of Skir) harder to get, the fact that it does not allow users to optionally specify a stable identifier to a message to get compatibility checks to work. I get your point about "why building another language", but also that point taken too far means that we would all be programming in Haskell.
- karteum 7mo agoApart from the comparison with Protobuf, how does it compare to flatbuffers, capnproto, messagepack, jsonbinpack... ?
- gepheum 7mo agoflatbuffers and capnproto are in the game of trying to make serialization to binary format as efficient as possible. Their goal is trying to beat benchmarks: how long it takes to convert an object to bytes and vice-versa. It's cool, but I personally think that for most use cases (not all), serialization efficiency shouldn't be the primary goal: serialization time is often negligible compared to time it takes to send data over the wire, and it's less important than other features (e.g. quality of the generated API) that some of these techs might neglect. I have an example to illustrate this. With Proto3, Google decided that when encoding a `string` field in C++, it would not perform UTF-8 validation. This leads to better benchmark metrics. This has also been a horrible mistake that led to many bugs which have costed so much in eng hours, since for example the same protobuf C++ API fails at deserialization when it encounters an invalid UTF-8 string. As per messagepack, jsonbinpack, these seem to be layers on top of JSON to make JSON more compact. They still use field names for field identity, which I think can be problematic for long-term data persistence since it prevents renaming fields. I think the Protobuf/Thrift approach of using meaningless field numbers in serialization forms is better.
- kentonv 7mo ago> flatbuffers and capnproto are in the game of trying to make serialization to binary format as efficient as possible. Little-understood fact about Cap'n Proto: Serialization is not the game at all. The RPC system is the whole game, the serialization was just done as a sort of stunt. Indeed, unless you are mmap()ing huge files, the serialization speed doesn't really matter. Though I would say the implementation of Cap'n Proto is quite a bit simpler than Protobuf due to the serialization format just being simpler, and that in itself is a nice benefit. The recently-released Cap'n Web jettisons the whole serialization side and focuses just on the RPC system: https://blog.cloudflare.com/capnweb-javascript-rpc-library/ https://blog.cloudflare.com/capnweb-javascript-rpc-library/ (I'm the author of Cap'n Proto and Cap'n Web.)
- battery8318 7mo agoI found that our work is somewhat similar. However, I mainly focus on HTTP and JSONRPC: https://news.ycombinator.com/item?id=47306983 https://news.ycombinator.com/item?id=47306983 Unfortunately, I really like postfix types, but IDL itself doesn't support them.
- ndr 7mo agoThis seems a Chesterton's fence fail. protobuf solved serialization with schema evolution back/forward compatibility. Skir seems to have great devex for the codegen part, but that's the least interesting aspect of protobufs. I don't see how the serialization this proposes fixes it without the numerical tagging equivalent.
- gepheum 7mo agoHey, Skir does have numerical tagging, see https://skir.build/docs/language-reference#structs https://skir.build/docs/language-reference#structs
- ndr 7mo agoThis seems new and retrofit. The implicit version is brittle design for backwards compatibility. People/LLMs will keep adding fields out of order and whatever has been serialised (both in client/server interaction, and stored in dbs) will be broken.
- anentropic 7mo agoThis vs JTD?
- marvin-hansen 7mo agoI had my fair share of frustration with proto as well. I appreciate in Skir GH style import. This is a big one I wish proto had in the first place. The entire idea of a proto registry feels reactive to me when, ideally, you want to pull in a versioned shared file to import that is verified by the compiler long before serve or client verifies the payload schema. Schema validation and compatibility checks on CI. Again a big one and critical to catch issues early. Enums done right... No further comment required. I think with some more attention to details e.g. hammering out the gaps some other comments have identified and more language support e.g. Rust, Go, C# this can actually work out over time. Here is an idea to contemplate as a side gig with your favorite Ai assistant: A tool to convert proto to Skir. Or at least as much as possible. As someone who had to maintain larger and complex proto files, a lot of proto specific pain points are addressed. The only concern i have is timing. Ten years ago this would have been a smash hit. These days, we have Thrift and similar meaning the bar is definitely higher. That's not necessarily bad, but one needs to be mindful about differentiation to the existing proto alternatives. I hope this project gains trajectory and community especially from the frustrated proto folks.
- gepheum 7mo agoHey, thanks a lot for the comment! I share your frustration with protobuf: although I think it's great, it carries a few design flaws which are hard to fix at this point and they create pain points which are not going away. I completely agree with you about timing, wish I had done this 10 years ago :) "Here is an idea to contemplate as a side gig with your favorite Ai assistant: A tool to convert proto to Skir. Or at least as much as possible. As someone who had to maintain larger and complex proto files, a lot of proto specific pain points are addressed." < I tried asking Claude: "Migrate this project from protobuf to Skir, see https://skir.build/ https://skir.build/" and it works pretty well. I created http://skir.build/llms.txt http://skir.build/llms.txt which helps with this. The pain point is data migration though, and as much as I want Skir to succeed: I cannot recommend people migrating from protobuf to Skir if they have some persisted data to migrate, the effort is probably not worth it.
- cyberax 7mo agoDefinitely interesting and seems to be a nice improvement over Protobuf. Especially for Python, Protobuf bindings for Python were made probably after taking a lot of hallucinogenic drugs. I like constants, great addition. Things that I'll miss: 1. Oneof fields. There are enums, but it looks like it's not possible to have ad-hoc onefos? 2. Streaming requests/responses. 3. Introspection and annotations. 4. Go bindings.
- gepheum 7mo agoThanks for the comment! Agree with you about the horrible Protobuf-to-Python bidding, it was a big frustration and definitely contributed to me wanting to build Skir. 1. You can create an enum with just "wrapper" fields, that's exactly like a oneof 2. Totally fair, I'm planning to work on this later this year, probably Q3 (priority is adding support to 4 more languages, and then I'll get to it) 3. So there is introspection in the 6 targeted languages, and I think I did it a bit better than protobuf because it generally has better type safety. Example in C++: https://github.com/gepheum/skir-cc-example/blob/main/string_capitalizer.h https://github.com/gepheum/skir-cc-example/blob/main/string_...; Typescript: https://skir.build/docs/typescript#reflection https://skir.build/docs/typescript#reflection I realize I haven't documented it in Python (although it is available and generally the same API as Typescript), will fix that However, you're right that there is no support yet for annotations. Still trying to gauge whether that's needed 4. Assuming you mean Go language: working on that now, hoping to have C#, Go, Rust and Swift in the next 2-3 months.
- cyberax 7mo ago> So there is introspection in the 6 targeted languages, and I think I did it a bit better than protobuf because it generally has better type safety. A bit different kind of introspection. In Protobuf I can write a code generator that loads the compiled PB descriptions and then generates whatever it needs. For example, I'm using it to generate SQL-serializing wrappers for my Protobuf types for Go. Oh yeah, also having a standardized pretty-printer would be great.