4 ms·
Show HN: QUALITY.md – open format/specification, agent skill, and CLI
Hello all, I created QUALITY.md to help build a holistic quality evaluation process for my projects. Turns out it's also ideal for loop engineering. I'm hoping this provides a valuable contribution to the conversation around quality and craft and having AI help us in the effort. I hope to shift the mindset from a reactive/review/repair mindset to a proactive care mindset.
Give it a go. I look forward to your thoughts/comments/feedback!
Website: https://getquality.md https://getquality.md
GitHub: https://github.com/qualitymd/quality.md https://github.com/qualitymd/quality.md
- dofm 3mo agoThe one thing I do not understand is that here you say: "Ensure stakeholders are aligned on what matters most and why" But it is instructions for LLMs, right? A way to describe something that the humans know and the LLMs don't. LLMs literally cannot be stakeholders, by definition.
- chrisweekly 3mo agoNot OP, but it seems to me the idea is that stakeholders can collaborate and come to consensus on the contents of QUALITY.md.
- craigsmitham 3mo agoTHe problem is that humans often don't know - this is as much about encouraging getting the humans aligned as the agents. Completely agree agents really aren't stakeholders. Fine point. I'll update description to clarify ... thank you!
- MisterKent 3mo agoIs this really where we've landed? I refuse to believe that any of this markdown insanity will continue indefinitely.
- willcodeforfoo 3mo agoI thought the same about Yaml and Kubernetes/Helm…
- nextaccountic 3mo agoit's looking like llms are interpreters, and markdown plus english text is the language of choice to run non deterministic programs on it
- blooalien 3mo ago> it's looking like llms are interpreters, and markdown plus english text is the language of choice to run non deterministic programs on it That's actually a pretty good clear way of putting it for the typical nerdy "programmer minded" individual.
- cyanydeez 3mo agoit is until we define real consistent deterministic gates and protocols. It really is a symptom of the lack of concerted effort. Everyone has a personal preference on how to shove the context and most of them are just "here's some good text I've found to work in my context"
- blooalien 3mo ago> define real consistent deterministic gates and protocols I've been experimenting with doing kinda exactly that with the "routing layer" / "harness" level of things, before the "main" LLM itself ever receives the user's input, by getting "user intent" (as a little JSON packet) really quickly from an ultra-lightweight model first and deciding from there in deterministic code what "context" to inject into the user message template, which system prompt to use, and which model to route the assembled context "packet" to for the final response. These LLMs really are fun to play with once you get a feel for which ones do what well, and where each falls short so you can use them each around their individual strengths. :)
- 8cvor6j844qw_d6 3mo ago
- bellowsgulch 3mo agoWhat?
- formerly_proven 3mo agoPure slop.
- athrowaway3z 3mo agoWhats the revenue model for this NBPaaS? (No Bugs Please As A Service)
- craigsmitham 3mo agoNo commercial offering associated with this effort. But a lot of potential for others to incorporate the QUALITY.md standard into products that assess/evaluate quality at varying levels of a loop stack. The agent skill/CLI that's provided generate a quality evaluation report with recommendations for handoff (ideal for loop engineering) is just one example of how the QUALITY.md file can be used. It's easy to imagine a SaaS that does the same that provides better eval, reporting, and integration capabilities.
- bironran 3mo agoThis is perfectly encapsulated in xkcd's "Standards" strip [https://xkcd.com/927/ https://xkcd.com/927/].
- stronglikedan 3mo agoSo is every proposal to standardize new things, and eventually the cream rises to the top, even though some people are perfectly happy sticking with milk.
- craigsmitham 3mo agoI'd really like feedback on the standard/specification. In short, it defines a quality model of quality factors/characteristics (which you can define as security, reliability, etc), requirements (how those qualities are assessed), a customizable rating scale, and "areas" to have different attributes/requirements for different areas of your project (e.g. frontend/backend, tests, specs, etc). That's basically it - and it follows a consistent pattern of how quality models have been practiced for decades. They are simple and powerful, but - until AI - kind of a pain/toil to get started with. Simple but not easy - until now. QUALITY.md + AI makes it easy. However, you still have to put in the work/care/attention to what goes into your QUALITY.md so you can get maximum leverage from it.
- hiAndrewQuinn 3mo agoI'm less interested in this than in what people are willing to aggressively trade off against in order to get the stuff they truly care about. For example, readability. Where are the developers out there saying "I am very willing to sacrifice a lot of readability to get even a small improvement on e.g. abstraction cleanliness", and sticking with it? Or "performance can take a huge hit at the cost of being dead easy to read and reason about". Coming up with a list of abstractly good-sounding qualities is just prosocial signaling without knowing what you're willing to sacrifice. There should be a FUCKIT.md that enumerates these.
- craigsmitham 3mo agoOP here. You're spot on. Trade-offs matter. The trade-offs are implied by the selection of what quality factors/attributes are selected and their requirements. A statement like "performance can take a huge hit at the cost of being dead easy to read and reason about" can sit right there in the QUALITY.md as a comment or in the markdown body.
- Leewen 3mo agoUseful Nice
- LiamPowell 3mo agoHere's the question I ask about every project that claims to make a LLMs output so much better: If it works so well then why would the model provider not just put it in the system prompt? Or in the case of interactive skills, why would Claude Code/Codex not make it a core part of the product? On top of that, if your magic markdown file really does work then where's the evidence showing that? These projects never include even basic benchmarks. At best they're entirely vibe based, however more often they're completely untested. Give us a proper benchmark, even a single prompt and it's output with and without your skill in use would be better than every other project out there.
- craigsmitham 3mo agoNo magic. QUALITY.md describes what is unique and valueble to your proejct context that model providers won't have insight into.
- istvan0 3mo agoYou can already put these into the AGENTS.md(or CLAUDE.md) or if it’s too big, you can put it into a SKILL — no need to reinvent the wheel.