11 ms·
I made my own Git
- kgeist 8mo ago>The hardest part about this project was actually just parsing. How about using sqlite for this? Then you wouldn't need to parse anything, just read/update tables. Fast indexing out of the box, too.
- deleted 8mo ago[deleted]
- grenran 8mo agothat would be what https://fossil-scm.org/ https://fossil-scm.org/ is
- TonyStr 8mo agoVery interesting. Looks like fossil has made some unique design choices that differ from git[0]. Has anyone here used it? I'd love to hear how it compares. [0] https://fossil-scm.org/home/doc/trunk/www/fossil-v-git.wiki#history https://fossil-scm.org/home/doc/trunk/www/fossil-v-git.wiki#...
- embedding-shape 8mo agoUsed it on and off mainly to check it out, but always in a personal/experimental capacity. Never managed to convince any teams to give it a try, mostly because git don't tend to get in the way, so hard to justify to learn something completely new. I really enjoy how local-first it is, as someone who sometimes work without internet connection. That the data around "work" is part of the SCM as well, not just the code, makes a lot of sense to me at a high-level, and many times I wish git worked the same...
- usrbinbash 8mo agoI mean, git is just as "local-first" (a git repo is just a directory after all), and the standard git-toolchain includes a server, so... But yeah, fossil is interesting, and it's a crying shame its not more well known, for the exact reasons you point out.
- embedding-shape 8mo ago> I mean, git is just as "local-first" (a git repo is just a directory after all), and the standard git-toolchain includes a server, so... It isn't though, Fossil integrates all the data around the code too in the "repository", so issues, wiki, documentation, notes and so on are all together, not like in git where most commonly you have those things on another platform, or you use something like `git notes` which has maybe 10% of the features of the respective Fossil feature. It might be useful to scan through the list of features of Fossil and dig into it, because it does a lot more than you seem to think :) https://fossil-scm.org/home/doc/trunk/www/index.wiki https://fossil-scm.org/home/doc/trunk/www/index.wiki
- adastra22 8mo agoThose things exist for git too, e.g. git-bug. But the first-class to do it in git is email.
- embedding-shape 8mo agoEmail isn't a wiki, bug tracking, documentation and all the other stuff Fossil offers as part of their core design. The point is for it to be in one place, and local-first. If you don't trust me, read the list of features and give it a try yourself: https://fossil-scm.org/home/doc/trunk/www/index.wiki https://fossil-scm.org/home/doc/trunk/www/index.wiki
- adastra22 8mo agoI am aware of fossil. Did you look up git-bug?
- smartmic 8mo agoI use Fossil extensively, but only for personal projects. There are specific design conditions, such as no rebasing [0], and overall, it is simpler yet more useful to me. However, I think Fossil is better suited for projects governed under the cathedral model than the bazaar model. It's great for self-hosting, and the web UI is excellent not only for version control, but also for managing a software development project. However, if you want a low barrier to integrating contributions, Fossil is not as good as the various Git forges out there. You have to either receive patches or Fossil bundles via email or forum, or onboard/register contributors as developers with quite wide repo permissions. [0]: https://fossil-scm.org/home/doc/trunk/www/rebaseharm.md https://fossil-scm.org/home/doc/trunk/www/rebaseharm.md
- graemep 8mo agoI like it but the problem is everyone else already knows git and everything integrates with git. It is very easy to self host. Not having staging is awkward at first but works well once you get used to it. I prefer it for personal projects. In think its better for small teams if people are willing to adjust but have not had enough opportunities to try it.
- TonyStr 8mo agoIs it possible to commit individual files, or specific lines, without a staging area? I guess this might be against Fossil's ethos, and you're supposed to just commit everything every time?
- jact 8mo agoI use Fossil extensively for all my personal projects and find it superior for the general case. As others said it’s more suited for small projects. I also use Fossil for lots of weird things. I created a forum game using Fossil’s ticket and forum features because it’s so easy to spin up and for my friends to sign in to. At work we ended up using Fossil in production to manage configuration and deployment in a highly locked down customer environment where its ability to run as a single static binary, talk over HTTP without external dependencies, etc. was essential. It was a poor man’s deployment tool, but it performed admirably. Fossil even works well as a blogging platform.
- dchest 8mo agoWhile Fossil uses SQLite for underlying storage (instead of the filesystem directly) and various support infrastructure, its actual format is not based on SQLite: https://fossil-scm.org/home/doc/trunk/www/fileformat.wiki https://fossil-scm.org/home/doc/trunk/www/fileformat.wiki It's basically plaintext. Even deltas are plaintext for text files. Reason: "The global state of a fossil repository is kept simple so that it can endure in useful form for decades or centuries. A fossil repository is intended to be readable, searchable, and extensible by people not yet born."
- storystarling 8mo agoSQLite solves the storage layer but I suspect you run into a pretty big impedance mismatch on the graph traversals. For heavy DAG operations like history rewriting, a custom structure seems way more efficient than trying to model that relationally.
- SQLite 8mo agoThe Common Table Expression feature of SQL is very good at walking graphs. See, for example <https://sqlite.org/lang_with.html#queries_against_a_graph https://sqlite.org/lang_with.html#queries_against_a_graph>.
- prakhar1144 8mo agoI was also playing around with the ".git" directory - ended up writing: "What's inside .git ?" - https://prakharpratyush.com/blog/7/ https://prakharpratyush.com/blog/7/
- sluongng 8mo agoZstd dictionary compression is essentially how Meta's Mercurial fork (Sapling VCS) stores blobs https://sapling-scm.com/docs/dev/internals/zstdelta https://sapling-scm.com/docs/dev/internals/zstdelta. The source code is available in GitHub if folks want to study the tradeoffs vs git delta-compressed packfiles. I think theoratically, Git delta-compression is still a lot more optimized for smaller repos. But for bigger repos where sharding storaged is required, path-based delta dictionary compression does much better. Git recently (in the last 1 year) got something called "path-walk" which is fairly similar though.
- darkryder 8mo agoGreat writeup! It's always fun to learn the details of the tools we use daily. For others, I highly recommend Git from the Bottom Up[1]. It is a very well-written piece on internal data structures and does a great job of demystifying the opaque git commands that most beginners blindly follow. Best thing you'll learn in 20ish minutes. 1. https://jwiegley.github.io/git-from-the-bottom-up/ https://jwiegley.github.io/git-from-the-bottom-up/
- spuz 8mo agoThanks - I think this is the article I was thinking of that really helped me to understand git when I first started using it back in the day. I tried to find it again and couldn't.
- MarsIronPI 8mo agoOh, I hadn't ever seen that one. I "grokked" Git thanks to The Git Parable[0] several years ago. [0]: https://tom.preston-werner.com/2009/05/19/the-git-parable https://tom.preston-werner.com/2009/05/19/the-git-parable
- sanufar 8mo agoOoh, this looks fun! I didn’t know you could cat-file on a hash id, that’s actually quite cool.
- heckelson 8mo agogentle reminder to set your website's `<title>` to something descriptive :)
- TonyStr 8mo agohaha, thank you. Added now :-)
- teiferer 8mo agoIf you ever wonder how coding agents know how to plan things etc, this is the kind of article they get this training from. Ends up being circular if the author used LLM help for this writeup though there are no obvious signs of that.
- wasmainiac 8mo agoMaybe we can poison LLMs with loops of 2 or more self referencing blogs.
- jama211 8mo agoI see the AI hating part of HN has come out again
- jdiff 8mo agoOnly need one, they're not thinking critically about the media they consume during training.
- falcor84 8mo agoHere's a sad prediction: over the coming few years, AIs will get significantly better at critical evaluation of sources, while humans will get even worse at it.
- topaz0 8mo agoMy sad prediction is that LLMs and humans will both get worse. Humans might get worse faster though.
- whstl 8mo agoI wish I could disagree with you, but what I'm seeing on average (especially at work) is exactly that: people asking stuff to ChatGPT and accepting hallucinations as fact, and then fighting me when I say it's not true.
- sneela 8mo ago> If you want to look at the code, it's available on github. Why not tvc-hub :P Jokes aside, great write up!
- TonyStr 8mo agohaha, maybe that's the next project. It did feel weird to make git commits at the same time as I was making tvc commits
- igorw 8mo agoRandom but y'all might enjoy. Git client in PHP, supports reading packfiles, reftables, diff via LCS. Written by hand. https://github.com/igorwwwwwwwwwwwwwwwwwwww/gipht-horse https://github.com/igorwwwwwwwwwwwwwwwwwwww/gipht-horse
- nasretdinov 8mo agoNice! This repo is a huge W for PHP I'd say. P.S. Didn't know that plain '@' can be used instead of HEAD, but I guess it makes sense since you can omit both left and right parts of the expressions separated by '@'
- h1fra 8mo agoLearning git internals was definitely the moment it became clear to me how efficient and smart git is. And this way of versionning can be reused in other fields, as soon as have some kind of graph of data that can be modified independently but read all together then it makes sense.
- black_13 8mo ago[dead]
- p4bl0 8mo agoNice post :). It made me think of ugit: DIY Git in Python [1] which is still by far my favorite of this kind of posts. It really goes deep into Git internals while managing to stay easy to follow along the way. [1] https://www.leshenko.net/p/ugit/ https://www.leshenko.net/p/ugit/
- eru 8mo ago> These objects are also compressed to save space, so writing to and reading from .git/objects/ will always involve running a compression algoritm. Git uses zlib to compress objects, but looking at competitors, zstd seemed more promising: That's a weird thing to put so close to the start. Compression is about the least interesting aspect of Git's design.
- alphabetag675 8mo agoWhen you are learning, everything is important. I think it is okay to cut the person some slack regarding this.
- eru 8mo agoYes, probably. It's just that git does a much more interesting job with compression, actually. Lot's more to learn. They don't compress the snapshots via something like zstd directly, that comes much later after a delta step. (Interestingly, that delta compression step doesn't use the diffs that `git show` shows you for your commits.)
- deleted 8mo ago[deleted]
- nasretdinov 8mo agoNice work! On a complete tangent, Git is the only SCM known to me that supports recursive merge strategy [1] (instead of the regular 3-way merge), which essentially always remembers resolved conflicts without you needing to do anything. This is a very underrated feature of Git and somehow people still manage to choose rebase over it. If you ever get to implementing merges, please make sure you have a mechanism for remembering the conflict resolution history :). [1] https://stackoverflow.com/questions/55998614/merge-made-by-recursive-strategy https://stackoverflow.com/questions/55998614/merge-made-by-r...
- arunix 8mo agoI remember in a previous job having to enable git rerere, otherwise it wouldn't remember previously resolved conflicts. https://git-scm.com/book/en/v2/Git-Tools-Rerere https://git-scm.com/book/en/v2/Git-Tools-Rerere
- nasretdinov 8mo agoI believe rerere is a local cache, so you'd still have to resolve the conflicts again on another machine. The recursive merge doesn't have this issue — the conflict resolution inside the merge commits is effectively remembered (although due to how Git operates it actually never even considers it a conflict to be remembered — just a snapshot of the closest state to the merged branches)
- Guvante 8mo agoAre people repeatedly handling merge conflicts on multiple machines? If there was a better way to handle "I needed to merge in the middle of my PR work" without introducing reverse merged permanently in the history I wouldn't mind merge commits. But tools will sometimes skip over others work if you `git pull` a change into your local repo due to getting confused which leg of the merge to follow.
- nasretdinov 8mo agoOne place where it mattered was when I was working on a large PHP web site, where backend devs and frontend devs would be working in the same branch — this way you don't have to go back and forth to get the new API, and this workflow was quite unique and, in my mind, quite efficient. The branchs also could live for some time (e.g. in case of large refactorings), and it's a good idea to merge in the master branch frequently, so recursive merge was really nice. Nowadays, of course, you design the API for your frontend, mobile, etc, upfront, so there's little reason to do that anymore.
- jrockway 8mo agosha256 is a very slow algorithm, even with hardware acceleration. BLAKE3 would probably make a noticeable performance difference. Some reading from 2021: https://jolynch.github.io/posts/use_fast_data_algorithms/ https://jolynch.github.io/posts/use_fast_data_algorithms/ It is really hard to describe how slow sha256 is. Go sha256 some big files. Do you think it's disk IO that's making it take so long? It's not, you have a super fast SSD. It's sha256 that's slow.
- grumbelbart2 8mo agoIs that even when using the SHA256 hardware extensions? https://en.wikipedia.org/wiki/SHA_instruction_set https://en.wikipedia.org/wiki/SHA_instruction_set
- oconnor663 8mo agoIt's mixed. You get something in the neighborhood of a 3-4x speedup with SHA-NI, but the algorithm is fundamentally serial. Fully parallel algorithms like BLAKE3 and K12, which can use wide vector extensions like AVX-512, can be substantially faster (10x+) even on one core. And multithreading compounds with that, if you have enough input to keep a lot of cores occupied. On the other hand, if you're limited to one thread and older/smaller vector extensions (SSE, NEON), hardware-accelerated SHA-256 can win. It can also win in the short input regime where parallelism isn't possible (< 4 KiB for BLAKE3).
- EdSchouten 8mo agoIt depends on the architecture. On ARM64, SHA-256 tends to be faster than BLAKE3. The reasons being that most modern ARM64 CPUs have native SHA-256 instructions, and lack an equivalent of AVX-512. Furthermore, if your input files are large enough that parallelizing across multiple cores makes sense, then it's generally better to change your data model to eliminate the existence of the large inputs altogether. For example, Git is somewhat primitive in that every file is a single object. In retrospect it would have been smarter to decompose large files into chunks using a Content Defined Chunking (CDC) algorithm, and model large files as a manifest of chunks. That way you get better deduplication. The resulting chunks can then be hashed in parallel, using a single-threaded algorithm.
- mg794613 8mo ago"Though I suck at it, my go-to language for side-projects is always Rust" Hmm, dont be so hard on yourself! proceeds to call ls from rust Ok nevermind, although I dont think rust is the issue here. (Tony I'm joking, thanks for the article)
- sublinear 8mo ago> If I were to do this again, I would probably use a well-defined language like yaml or json to store object information. I know this is only meant to be an educational project, but please avoid yaml (especially for anything generated). It may be a superset of json, but that should strongly suggest that json is enough. I am aware I'm making a decade old complaint now, but we already have such an absurd mess with every tool that decided to prefer yaml (docker/k8s, swagger, etc.) and it never got any better. Let's not make that mistake again. People just learned to cope or avoid yaml where they can, and luckily these are such widely used tools that we have plenty of boilerplate examples to cheat from. A new tool lacking docs or examples that only accepts yaml would be anywhere from mildly frustrating to borderline unusable.
- holoduke 8mo agoI wonder if in the near future there will be no tools anymore in the sense we know it. you will maybe describe the tool you need and its created on the fly.
- ofou 8mo agobtw, you can change the hashing algorithm in git easily
- justabrowser 8mo ago[flagged]
- adzm 8mo agoSecond time today I've read and agreed with most of your comment only to eyeroll and downvote once seeing your ridiculous and immature edit.
- smangold 8mo agoTony nice work!
- b1temy 8mo agoNice work, it's always interesting to see how one would design their own VCS from scratch, and see if they fall into problems existing implementations fell into in the past and if the same solution was naturally reached. The `tvc ls` command seems to always recompute the hash for every non-ignored file in the directory and its children. Based on the description in the blog post, it seems the same/similar thing is happening during commits as well. I imagine such an operation would become expensive in a giant monorepo with many many files, and perhaps a few large binary files thrown in. I'm not sure how git handles it (if it even does, but I'm sure it must). Perhaps it caches the hash somewhere in the `.git`directory, and only updates it if it senses the file hash changed (Hm... If it can't detect this by re-hashing the file and comparing it with a known value, perhaps by the timestamp the file was last edited?). > Git uses SHA-1, which is an old and cryptographically broken algorithm. This doesn't actually matter to me though, since I'll only be using hashes to identify files by their content; not to protect any secrets This _should_ matter to you in any case, even if it is "just to identify files". If hash collisions (See: SHAttered, dating back to 2017) were to occur, an attacker could, for example, have two scripts uploaded in a repository, one a clean benign script, and another malicious script with the same hash, perhaps hidden away in some deeply nested directory, and a user pulling the script might see the benign script but actually pull in the malicious script. In practice, I don't think this attack has ever happened in git, even with SHA-1. Interestingly, it seems that git itself is considering switching to SHA-256 as of a few months ago https://lwn.net/Articles/1042172/ https://lwn.net/Articles/1042172/ I've not personally heard of the process of hashing to also be known as digesting, though I don't doubt that it is the case. I've mostly familiar of the resulting hash being referred to as the message digest. Perhaps it's to differentiate between the verb 'hash' (the process of hashing) with the output 'hash' (the ` result of hashing). And naming the function `sha256::try_digest`makes it more explicit that it is returning the hash/digest. But it is a bit of a reach, perhaps that are just synonyms to be used interchangeably as you said. On a tangent, why were TOML files not considered at the end? I've no skin in the game and don't really mind either way, but I'm just curious since I often see Rust developers gravitate to that over YAML or JSON, presumably because it is what Cargo uses for its manifest. -- Also, obligatory mention of jujutsu/jj since it seems to always be mentioned when talking of a VCS in HN.
- 8mo ago
- athrowaway3z 8mo agoI do wonder if the compression step makes sense at this layer instead of the filesystem layer.
- aabbcc1241 8mo agoInteresting take. I'm using btrfs (instead of ext4) with compression enabled (using zstd), so most of the files are compressed "transparently" - the files appear as normal files to the applications, but on disk it is compressed, and the application don't need to do the compress/decompress.
- quijoteuniv 8mo agoNow … if you reinvent Linux you are closer to be compared to LT
- oldestofsports 8mo agoNice job, great article! I had a go at it as well a while back, I call it "shit" https://github.com/emanueldonalds/shit https://github.com/emanueldonalds/shit
- tpoacher 8mo agoTHE shit, in fact.
- hahahahhaah 8mo agoFast Useful Change Keeper
- astinashler 8mo agoDoes this git include empty folder? I always annoy that it's not track empty folder.
- TonyStr 8mo agoyep! Had to check to be sure: Finished `dev` profile [unoptimized + debuginfo] target(s) in 0.02s Running `target/debug/tvc decompress f854e0b307caf47dee5c09c34641c41b8d5135461fcb26096af030f80d23b0e5` === args === decompress f854e0b307caf47dee5c09c34641c41b8d5135461fcb26096af030f80d23b0e5 === tvcignore === ./target ./.git ./.tvc === subcommand === decompress ------------------ tree ./src/empty-folder e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 blob ./src/main.rs fdc4ccaa3a6dcc0d5451f8e5ca8aeac0f5a6566fe32e76125d627af4edf2db97
- woodrowbarlow 8mo agohuh, cool. what happens if you use vanilla-git to clone a repo that contains empty folders? and do forges like github display them properly?
- lucasoshiro 8mo agoActually, the Git data model supports empty directories, however, the index doesn't since it only maps names to files but not to directories. You can even create a commit with a root directory using --allow-empty, and it will use the hardcoded empty tree object (4b825dc642cb6eb9a060e54bf8d69288fbee4904).
- brendoncarroll 8mo agoMe too. Version control is great, it should get more use outside of software. https://github.com/gotvc/got https://github.com/gotvc/got Notable differences: E2E encryption, parallel imports (Got will light up all your cores), and a data structure that supports large files and directories.
- DASD 8mo agoNice! Not sure if you're aware of Got(Game of Trees) that appears to pre-date your Got. https://gameoftrees.org/index.html https://gameoftrees.org/index.html
- brendoncarroll 8mo agoYes the author reached out. There has not yet been a confusion among real users that I am aware of. https://github.com/gotvc/got/issues/20 https://github.com/gotvc/got/issues/20
- rtkwe 8mo agoThe problem is when you move beyond text files it gets hard to tell what changes between two versions without opening both versions in whatever program they come from and comparing.
- brendoncarroll 8mo ago> The problem is when you move beyond text files it gets hard to tell what changes between two versions without opening both versions in whatever program they come from and comparing. Yeah, totally agree. Got has not solved conflict resolution for arbitrary files. However, we can tell the user where the files differ, and that the file has changed. There is still value in being able to import files and directories of arbitrary sizes, and having the data encrypted. This is the necessary infrastructure to be able to do distributed version control on large amounts of private data. You can't do that easily with Git. It's very clunky even with remote helpers and LFS. I talk about that in the Why Got? section of the docs. https://github.com/gotvc/got/blob/master/doc/1.1_Why_Got.md https://github.com/gotvc/got/blob/master/doc/1.1_Why_Got.md
- direwolf20 8mo agoCool. When you reimplement something, it forces you to see the fractal complexity of it.
- temporallobe 8mo agoReminds me of when I tried to invent a SPA framework. So much hidden complexity I hadn’t thought of and I found myself going down rabbit holes that I am sure the creators of React and Angular went down. Git seems to be like this and I am often reminded of how impressive it is at hiding underlying complexity.
- alsetmusic 8mo ago> at hiding underlying complexity. It's only in the context of recreating Git that this comment makes sense.
- lasgawe 8mo agonice work! This is one of the best ways to deeply learn something, reinvent the wheel yourself.
- gkbrk 8mo agoCodeCrafters has an amazing "Build your own Git" [1] tutorial too. Jon Gjengset has a nice video [2] doing this challenge live with Rust. [1]: https://app.codecrafters.io/courses/git/overview https://app.codecrafters.io/courses/git/overview [2]: https://www.youtube.com/watch?v=u0VotuGzD_w https://www.youtube.com/watch?v=u0VotuGzD_w
- smekta 8mo ago...with blackjacks, and hookers
- KolmogorovComp 8mo agoIt’s really a shame git storage use files as the unit for storage. That’s what makes it improper for usage with many of small files, or large files. Content-based chunking like Xethub uses really should become the default. It’s not like it’s new either, rsync is based on it. https://huggingface.co/blog/xethub-joins-hf https://huggingface.co/blog/xethub-joins-hf
- jonny_eh 8mo agoWhy introduce yet another ignore file? Can you have it read .gitignore if .tvcignore is missing?
- bryan2 8mo agoFtr you can make repos with sha256 now. I wonder if signing sha-1 mitigates the threat of using an outdated hash.