6 ms·
Redis array: short story of a long development process
- SuperV1234 5mo agoClosely matches my own experiences with current SOTA AI. Extremely useful collaborator, far from being a replacement for human intelligence and creativity.
- antirez 5mo agoThere are projects that I develop mostly not looking at the code, but owning the concepts, algorithms and ideas asking questions and giving hints, and owning especially the product. But, not for Redis, not yet at least. When in the future this will be possible, server software, the way it is developed today, will be over. I bet there will be still projects and repositories, as accumulation of features, fixes and experiences will still be worth it, but the role of programmers will be very similar to what Linus did so far for the kernel. And for certain projects I'm developing, like the DeepSeek v4 inference engine, I'l already working like that.
- foobarian 5mo agoI like to say, AI is the duck programming duck I always wanted
- bonesss 5mo agoLLMs are the insensitive Asmovian robots I’ve always wanted, who translate and do the hardest part of my job: ensuring my emails are polite and none of my true thoughts or feelings are revealed… Now I just need a way to protect my chats from any potential discovery, and <pew pew> business’ll be easy.
- genghisjahn 5mo agoI occasionally type into slack "Future lawyers, the previous conversation is a joke. No one is doing cocaine to get through writing requirements docs."
- imadethis 5mo agoWe have a “don’t get the slack subpoenaed” emoji that gets frequent use. Incidentally, a lawyer doing discovery in the future could just search for uses of that emoji to find what they’re looking for.
- gbalduzzi 5mo agoIs it possible to see the specification file you created and used for AI assisted development? Very cool anyway! Can I expect a youtube video about this soon?
- antirez 5mo agoYep I will release it, it is a bit out of sync at this point, but will do a pass of updating and will release it.
- nojvek 5mo agoIt’s always a great HN thread when an author of a widely used lib/app engages on a technical level. antirez - you inspire a generation of devs. Thanks for all you do.
- leetrout 5mo agoOn safari mobile it's a page with the title header and a footer. Theres no content rendering.
- jdw64 5mo agoIt feels like Redis is becoming a small database, which seems to make it more convenient to use. Could you add more examples that clarify where the boundary should be?
- antirez 5mo agoWell, Redis is a data structures server, and has very complicated and edgy data structures like the HyperLogLog, so I have very little doubts that a fundamental data type like the Array will fit :) Also the actual complexity added is mostly two C files that are quite commented and understandable. wc -l t_array.c sparsearray.c 2012 t_array.c 2063 sparsearray.c 4075 total (including comments) Sure there are also the AOF / RDB glues, the tests, the vendored TRE library for ARGREP. But all in all it's self contained complexity with little interactions with the rest of the server. A quick note: if we focus only on that part of the implementation, skipping tests and persistence code which is not huge, 4075 lines in 4 months are an average of 33 lines per day, which is quite low.
- jdw64 5mo agoI’m a big fan of your work, and I honestly didn’t expect to receive a reply from you. Thank you. Also, thank you for pointing out exactly where I was misunderstanding the issue. In the past, I used Redis for temperature measurements in a smart farm project. I used Hashes back then, but it seems like Array would fit that use case much better. This looks like a very useful feature. Thank you again for the reply.
- antirez 5mo agoI appreciate your kind reply as well :)
- localhoster 5mo agoLet's make it very clear - this is the original creator of redis, or one of them. He is not "your avg dev" and it took him 4 months with llm. This is not a seal of approval for you to go and command all your developers to move to Claude code/codex/any other ai coding tool fully. I'm looking at you - any avg CEO of a startup.
- simonw 5mo agoIt's a pretty strong endorsement for the idea that coding agents, used skillfully by experienced developers, can further amplify their expertise.
- zozbot234 5mo agoSure but the OP suggests that these were minor gains, and that this limited scope for gains was necessary in order to preserve the quality standard that's long been expected in that FLOSS community. We aren't talking about either a 10x productivity gain or one-shotting entire new features from scratch. This is arguably a key quote: "Then, it was time to read all the code, line by line. ... I found many small inefficiencies or design errors ... so I started a process of manual and AI-assisted rewrite of many modules." We should not underestimate that step: reading code line by line might easily require more time than writing it from scratch.
- simonw 5mo agoRight, and those of us who advocate for a sensible approach to agentic engineering don't talk about 10x productivity gains or one-shotting entire new (production-ready) features from scratch either. I remain unconvinced by the "faster to write it by hand than read it" arguments though. My experience throughout my career is that most people, myself included, top out at a couple of hundred lines of tested, production-ready code per day. I can productively review a couple of thousand.
- FEELmyAGI 5mo ago"top out at a couple of hundred lines of tested, production-ready code per day" + " productively review a couple of thousand." + LLM agents that write code for you = apparent contradiction with your first paragraph.
- ok123456 5mo agoIs this an apologia since the PR is +22,212 -34?
- antirez 5mo agoHaha, ~5000 LOC with comments. The rest is tests + TRE code + TRE tests.
- wood_spirit 5mo agoSharing my current MO: I start with a high level design md doc which an AI helps write. Then I ask another AI - whether the same model without the context, or another model - to critique it and spot bugs, gaps and omissions. It always finds obvious in hindsight stuff. So I ask it to summarize its findings and I paste that into the first AI and ask its opinions. We form an agreed change and make it and carry on this adversarial round robin until no model can suggest anything that seems weighty. I then ask the AI to make a plan. And I round robin that through a bunch of AIs adversarially as well. In the end, the plan looks solid. Then the end to end test cases plan and so on. By the end of the first day or week or month - depending on the scale of the system - we are ready to code. And as code gets made I paste that into other AIs with the spec and plan and ask them to spot bugs, omissions and gaps too and so on. Continually using other AI to check on the main one implementing. And of course you have to go read the code because I have found it that AI misses polishes.
- deleted 5mo ago[deleted]
- lovasoa 5mo agoHow much faster/slower are you with that process compared to writing code yourself?
- tracker1 5mo agoCan't speak for GP or OP, but I see about 10x the output and 2-4x the value of what I would be able to get by hand. Within the gap between 2-4x and the 10x is really a lot of design documents, user/dev documentation and testing that I might not have rolled to nearly the extent that I do/get when using AI. I haven't been using multiple AIs adversarially as OP, but might consider giving it a try with Codex and Opus. That said, my AI workflow has been pretty similar... lots of iterations on just design, then iterations on documentation, testing, etc... then iterations on implementation, testing, validation and human review in the mix. My analogy is that it's really close to working with a foreign dev team, but your turnaround is in minutes instead of days, where it's much more interactive.
- dsecurity49 5mo agoAI is a fantastic co-pilot, but you still need to know how to fly the plane when the edge cases start hitting the fan.
- epolanski 5mo agoGot few questions: - the project essentially spans almost 3 different (albeit minor) generations of LLMs. Have you noticed major differences in their personas, behavior, output for that specific use case? - when using AI for feedback, have you ever considered giving it different "personalities"? I have few skills that role play as very different reviewers with their own different (by design conflicting) personalities. I found this to improve the output, but also to be extremely tiring and to often have high noise ratio. - when did you, if ever, felt that AI was slowing you down massively compared to just doing it yourself (e.g. some specific bug or performance or design fix)? Are there recurring patterns? - conversely, how often did AI had moments where it genuinely gave you feedback or ideas that would've not come to you? - last: do you have specific prompts, skills, setups, etc to work on specific repositories?
- antirez 5mo ago1. The huge jump from from Opus to GPT 5.3. Game changer. GPT 5.4, 5.5, were better but only incrementally better. 2. Nope I don't give much personalities, but I use subtle prompt differences to maximize certain responses I want, to make the model focusing in a given detail or acting in a specific kind of engineering mindset. 3. It never happened that the AI was slowing me down since I always had the full context and code detail in mind of what was happening. I believe that this happens more when you don't have a clear idea. Also GPT >= 5.3/4 is not the past generation of models, it is very hard to trap it into a situation where it seems unable to understand what you mean. 4. A few times the AI provided fresh insights that I really liked. Most of the times it was the other way around. Certain implementations were written by the AI at a very impressive level of quality. 5. I don't use general skills, I build skills with deep search when needed for specific projects, and build an AGENT.md that works as a knowledge base as I work with the AI. One thing that I use a lot is, when there is a very complex problem, to tell GPT that I have a friend called Machiavelli that is an incredible computer scientist. To write him an email in /tmp/letter.md with the problem we are facing, and I'll try to get a reply. Then I ask GPT 5.5 Pro on the web with extensive reasoning set on. It will take sometimes 30 minutes or more to reply. Often times after I feed back the reply, the agent will be able to see things a lot more clearly.
- sylvinus 5mo agoThanks for the write up. Always interesting to see how very senior developers interact with AI these days. @antirez: Introducing a regex feature that late into the project for a seemingly unrelated feature feels a bit weird? Can you explain more your rationale on that? thanks!
- antirez 5mo agoOnce I realized arrays were a great fit for text files, many use cases I could conceive were always limited by the fact we need to grep on files. So I thought: what is the AROP equivalent for files? ARGREP. Then I made sure to add both fast, exact and regexp matching so that depending on the use case the best tool could be used. I then discovered that for many OR-ed strings regexps could be the faster way if we'll optimized. And then I specialized TRE a bit.
- ftruzzi 5mo ago[dead]
- simonw 5mo agoI vibe coded up an interactive playground against a WebAssembly build of the new array features: https://tools.simonwillison.net/redis-array https://tools.simonwillison.net/redis-array
- deleted 5mo ago[deleted]
- gurgeous 5mo agoThanks for adding this. Excited about array/regex, also very interested in your experience using LLMs to stretch your abilities. There are many of us laboring quietly on various projects attempting the same. "Vibe coding" (and the backlash) doesn't really capture how we work.
- tracker1 5mo agoI definitely don't consider how I've used agents as vibe coding at all... I'm much too involved and validate/verify/review everything.
- epolanski 5mo agoThe problem with "vibe coding" is that the author who coined the term gave it a very specific definition (after all, it's his term): writing software without looking at the code, just "vibing". Then it quickly lost its original meaning as people started using it for virtually all forms of AI-assisted coding.
- nitwit005 5mo agoThe use of C stdlib localization functions (toupper, mbrtowc, etc), makes me suspect if there will be some regex behavior differences between systems or locales.
- antirez 5mo agoRedis sets the locale at startup to avoid issues so should be ok but we will document that for instance è will not match È when nocase is used.
- ericpauley 5mo agoCouldn't some of the use cases presented for this be accomplished with ZSETs? I get the performance angle, but it seems that this could have been accomplished without the new API surface by selectively optimizing ZSET storage for dense values (in the same way that Arrays selectively use sparse representations). The RE component is interesting, but as commentary here has noted it seems orthogonal to the array data structure (i.e., usable on others as well). Does this not make more sense to accomplish with Lua scripting? Or if performance of Lua is an issue perhaps abstracting OP to be composable on top of any command that returns a range of values. I say this with reverence for Antirez as the expert in this space, but some of this new feature set feels like the sort of solution that I tend to see arise from LLM-driven development; namely creation of new functionality instead of enhancement of existing, plus overcomplicating features when composition with others might be more effective.
- antirez 5mo agoUnfortunately not, sorted sets are actually a bit in the other side of the spectrum: they are semantically sound, but absolutely wasteful because of the combined skiplist + array. Also, if the underlying representation is not an array, range queries and ring buffers will never be as efficient and compact as they should. In theory you can do everything with everything, but segmenting what each API can do allows you to exploit the use cases to provide the best underlying implementation.
- tibbar 5mo agoReviewing 22,000 lines of code, even from antirez, with this complex of a feature set and minimal PR description sounds like a nightmare. One starts to see why major open-source software like Postgres tends to be developed on a mailing list, with intermediate design decisions discussed by the community, separate patches for different related features, incremental review, and then a spaced release cadence.
- epolanski 5mo agoPostgres and Redis are dramatically different projects with radically different stories, contributions and development team. Virtually all major Redis features are a solo job of the post author. By the way reviewers are paid good money for this and know the setup.
- tibbar 5mo agoOh wow, I didn't realize that Redis is still mostly just authored by antirez! (My understanding is that he had left for some time and then returned to the project.) That is, honestly, pretty amazing. Well, redis is great and clearly it's worked out.
- antirez 5mo agoThe code is 5000 lines of code in total, comments included: 2000 lines the sparse array. 2000 lines the t_array commands and upper layer implementation. ~500 lines of AOF / RDB code. All the other stuff is tests, JSON command descriptions, TRE library under "deps".
- fancy_pantser 5mo agoI think the point GP is making is this is a PR that smells like a solo dev working on their own project and not how a community-driven project adds major new functionality, although I'm sure there are docs and descriptions (or at least a discussion of tradeoffs and design decisions if not ADRs) are somewhere, but not linked handily to the PR. There is a lot of explanation in the blog post and PR, but it's unilateral-looking. c.f. valkey and others
- thallavajhula 5mo agoSalvatore really wants to popularize the term Automatic Programming/Coding it seems. (https://antirez.com/news/159 https://antirez.com/news/159)
- beyti 5mo agoI keep finding myself to minimize the words to describe the same thing as well, since we are finding ourselves doing "that" operation more and more over time. maybe shortening the term to "auto-code" would help tho.
- zozbot234 5mo agohttps://en.wikipedia.org/wiki/Automatic_programming https://en.wikipedia.org/wiki/Automatic_programming It's an acknowledged term in computer science, describing any mechanism whatsoever of auto-generating code from a description at a higher level of abstraction. Of course LLM's are highly unusual in being non-deterministic and having a surprisingly broad scope, but this does not make the term inapplicable.
- shay_ker 5mo agoantirez: i'm curious, with the final code, have you experimented with effectively one-shotting the final result? i wonder if we can get there with GEPA, and maybe there's something we can learn in how to elicit/prompt these models to get what we want. or maybe the conclusion is that model providers need to clean up their training data!
- srinikhilr 5mo agoAnyone know how to get the specification mentioned in the blog post? Don't see one in the linked PR.
- jaunt7 5mo agoIn short, Redis can't be trusted any more. Who is going to do an LLM free fork?
- dontdoxxme 5mo agoYour comment is not constructive, why can't it be trusted? If every user of an LLM took this much care and attention, many people would have fewer issues with LLM assisted coding. In this case the author has demonstrated they can write plenty of code without an LLM, so why not use it carefully to benefit their productivity?
- zdkaster 5mo agoSadly, we can't really expect any rational like a sane person from this kind of anonymous random user created 14 hours go.
- deleted 5mo ago[deleted]
- ozozozd 5mo agoThat was too short a story @antirez!
- ardline 5mo agoSolid work. The devil's in the operational complexity, but this looks manageable.
- asG1298 5mo agoRedis wants to get in on the vector database market that is popular in AI. That is all there is to it. That is the reason why the Redis author keeps boosting AI. To the point where he even uses Redis to demonstrate how many bugs AI has found. Not every software is as buggy as Redis. It is all an advertisement by someone extremely adept at manipulating techies.