5 ms·
Show HN: CommerceTXT – An open standard for AI shopping context (like llms.txt)
Hi HN, author here.
I built CommerceTXT because I got tired of the fragility of extracting pricing and inventory data from HTML. AI agents currently waste ~8k tokens just to parse a product page, only to hallucinate the price or miss the fact that it's "Out of Stock".
CommerceTXT is a strict, read-only text protocol (CC0 Public Domain) designed to give agents deterministic ground truth. Think of it as `robots.txt` + `llms.txt` but structured specifically for transactions.
Key technical decisions v1.0:
1. *Fractal Architecture:* Root -> Category -> Product files. Agents only fetch what they need (saves bandwidth/tokens).
2. *Strictly Read-Only:* v1.0 intentionally excludes transactions/actions to avoid security nightmares. It's purely context.
3. *Token Efficiency:* A typical product definition is ~380 tokens vs ~8,500 for the HTML equivalent.
4. *Anti-Hallucination:* Includes directives like @INVENTORY with timestamps and @REVIEWS with verification sources.
The spec is live and open. I'd love your feedback on the directive structure and especially on the "Trust & Verification" concepts we're exploring.
Spec: https://github.com/commercetxt/commercetxt https://github.com/commercetxt/commercetxt
Website: https://commercetxt.org https://commercetxt.org
- deleted 10mo ago[deleted]
- reddalo 10mo agoWe should stop polluting website roots with these files (including llms.txt). All these files should be registered with IANA and put under the .well-known namespace. https://en.wikipedia.org/wiki/Well-known_URI https://en.wikipedia.org/wiki/Well-known_URI
- tsazan 10mo agoI understand the theoretical argument. We follow the precedent of robots.txt, ads.txt, and llms.txt. The reason is friction. Platforms like Shopify and Wix make .well-known folders difficult or impossible for merchants to configure. Root files work everywhere. Adoption matters more than namespace hygiene.
- JimDabell 10mo agoHow about following the precedent of all of these users of /.well-known/ https://en.wikipedia.org/wiki/Well-known_URI#List_of_well-known_URIs https://en.wikipedia.org/wiki/Well-known_URI#List_of_well-kn... robots.txt was created three decades ago, when we didn’t know any better. Moving llms.txt to /.well-known/ is literally issue #2 for llms.txt https://github.com/AnswerDotAI/llms-txt/issues/2 https://github.com/AnswerDotAI/llms-txt/issues/2 Please stop polluting the web.
- tsazan 10mo agoI prioritize simplicity and adoption for non-technical users over strict IETF compliance right now. My goal is to make this work for a shop owner on Shopify and Wix, not just for sysadmins. That said, I am open to supporting .well-known as a secondary location in v1.1 if the community wants it.
- xemdetia 10mo agoHow is using a standard path 'just for sysadmins' again? You are introducing something new today.
- tsazan 10mo agoTry uploading a file to /.well-known/ on Shopify or Wix. You cannot. Their file managers block hidden directories (starting with a dot). To do it, you need a custom app, a meta-field hack, or a reverse proxy. That is sysadmin work. Uploading a file to the root is user work. That is the difference.
- amitav1 10mo agoWait, am I dumb, or did the authors hallucinate? @INVENTORY says that 42 are in stock, but the text says "Only 3 left". Am I misunderstanding this or does stock mean something else?
- tsazan 10mo agoGood eye. This demonstrates the protocol’s core feature. The raw data shows 42. We used @SEMANTIC_LOGIC to force a limit of 3. The AI obeys the developer's rules, not just the CSV. We failed to mention this context. It causes confusion. We are changing it to 42.
- nebezb 10mo agoAh, so dark patterns then. Baked right into your standard.
- tsazan 10mo agoNot dark patterns. Operational logic. Physical stock rarely equals sellable stock. Items sit in abandoned carts. Or are held as safety buffers. If you have 42 items and 39 are reserved, telling the user "42 available" is the lie. It causes overselling. The protocol allows the developer to define the sellable reality. Crucially, we anticipated abuse. See Section 9: Cross-Verification. If an agent detects systematic manipulation (fake urgency that contradicts checkout data), the merchant suffers a Trust Score penalty. The protocol is designed to penalize dark patterns, not enable them.
- hrimfaxi 10mo agoWho maintains this trust score? How is it communicated to other agents?
- tsazan 10mo agoThere is no central authority. The Trust Score is a conceptual framework, not a shared database. Each AI platform (OpenAI, Anthropic, Google) builds its own model. They retain full discretion. Agents do not talk to each other. They talk to users. If a score is low, the agent warns the user. It adds caveats or drops the recommendation. It does not broadcast to other bots.
- duskdozer 10mo agoI'm not sure I understand the point of this as opposed to something like a json file, and also, assuming there is any type of structured format, why one would use an LLM for this task instead of a normal parser.
- tsazan 10mo agoYou assume JSON is a standalone file. It rarely is. Even if it were, JSON is verbose. Every bracket and quote costs tokens. In reality, the data is buried in 1MB+ of HTML. You download a haystack to find a needle. We fetch a standalone text file. It cuts the syntax tax. It is pure signal.
- xemdetia 10mo agoI believe what the commenter is suggesting is that since this is supposed to be machine readable then why not start with a common format like JSON similar to how things like MCP serve what functions are available or an OpenAPI spec. Generate the JSON and serve that from the well known directory. People serve plain JSON all the time. This proposed standard is essentially a structured file anyway.. why not YAML? Why not INI? Getting away from bespoke unicorn file formats has been good for everyone.
- tsazan 10mo agoJSON is great for code. It is heavy and deeply nested for Agents. The constraint is the context window. Brackets, quotes, and nesting are token tax. YAML is brittle. Whitespace errors break parsers. We chose the robots.txt model. It is dense and resilient. It is not a unicorn. It is a workhorse.
- throwaway_20357 10mo agoCan shops not just embed Schema/JSON-LD in the page if they want their information to be machine readable?
- tsazan 10mo agoThat is the current standard. But it is hard for agents to read efficiently. To access JSON-LD, an agent must download the entire HTML page. This creates a haystack problem where you download 2MB of noise just to find 5KB of data. Even then, you pay a syntax tax. JSON is verbose. Brackets and quotes waste valuable context window. Furthermore, the standard lacks behavior. JSON-LD lists facts but lacks instructions on how to sell (like @SEMANTIC_LOGIC). CommerceTXT is a fast lane. It does not replace JSON-LD. It optimizes it.
- inerte 10mo agoWouldn't be easier on everybody (servers and clients) to just expose Structured Data in a text file then? And add the 1 or 2 things it doesn't have?
- tsazan 10mo agoThat solves bandwidth. It fails on tokens. JSON syntax is heavy. Brackets and quotes consume context window. More importantly, Schema.org is a dictionary of facts. It lacks behavior. It defines what a product is, but not how to sell it. It has no concept of @SEMANTIC_LOGIC or @BRAND_VOICE. We need a format that carries both data and instructions efficiently. JSON-LD is too verbose and too static for that.
- captn3m0 10mo agoHow is "schema.org compatibility" related to Legal Compliance?
- tsazan 10mo agoSchema.org is the dictionary for facts. We map strictly to Schema.org for all transactional data (Price, Inventory, Policies). This ensures legal interoperability. But Schema.org describes what a product is, not how to sell it. So we extend it. We added directives like @SEMANTIC_LOGIC for agent behavior. We combine standard definitions for safety with new extensions for capability.
- captn3m0 10mo agoIs there a specific regulation this is for? What’s the compliance bit?
- tsazan 10mo agoIt targets Consumer Protection and Truth-in-Advertising laws globally. The 'compliance bit' is Price Transparency. If an AI quotes a price as 'final' but checkout adds hidden fees or tax, that is a deceptive practice. Our spec enforces fields like TaxIncluded and TaxNote. It instructs the Agent to disclose whether the price is net or gross. It prevents the AI from accidentally committing fraud via misleading omissions.
- hrimfaxi 10mo agoHow do you avoid downloading the whole haystack to search through the data? How does the hierarchy work? I have to keep a bunch of .txt files updated in my web root? Doesn't this require essentially mirroring the inventory db as text files (if the intent is for accurate counts of items, etc they would need to be updated in real time)?
- tsazan 10mo agoYou do not download the haystack. You traverse it. The architecture is fractal. The agent reads the Root. If the user wants "Headphones", it follows that specific link. It ignores the rest. It is lazy loading for context. Do not mirror your DB manually. For real stores, generate the files dynamically. It is a view layer, just like HTML or sitemap.xml. Real-time? Yes. Since it is a dynamic response, it reflects the DB state instantly. Cache-Control headers handle the freshness.
- pdntspa 10mo agoThis would have been great if it was adopted while I was still working on shopping site scrapers
- tsazan 10mo agoIt definitely lowers the barrier. But relying on messy HTML as a defense against competitors is 'security through obscurity'. It does not stop them; it just costs you server CPU. The data is public. If you put it on the screen, a scraper can read it. CommerceTXT just ensures that the good bots (AI Agents bringing customers) get it efficiently, while you can still block the bad ones via WAF.
- pdntspa 10mo agoIf it delivers accurate data then I can hit that instead of scraping the full HTML. Everybody wins. What I have found, however, with existing standardization of this kind of data (yours is not the first!), is that shopping sites (big ones) will lie, and you still need to read the HTML as ground truth.
- tsazan 10mo agoYou are right. Standardization often drifts from reality. That is why we built Section 9: Cross-Verification. The HTML remains the audit layer. The Agent does not trust blindly. It spot-checks. If commerce.txt says $50 but the HTML says $100, the merchant gets a Trust Score penalty. We do not replace the ground truth. We cache it, and we audit the cache to ensure it matches.
- pdntspa 10mo agoThen why bother with commerce.txt?
- tsazan 10mo agoBecause you don't need to audit every single transaction. Think of it like a cache. You use the commerce.txt for 99% of your agentic workflows because it’s 30% cheaper in tokens and 95% faster than parsing a 2MB HTML haystack. You only 'bother' with the HTML for periodic spot-checks or when a high-value transaction requires absolute verification. Without CommerceTXT, you are forced to pay the 'HTML tax' on every single interaction. With it, you get a high-speed fast lane for context, while keeping the HTML as a decentralized source of truth for when trust needs to be verified. It’s about moving the baseline from 'expensive and fragile' to 'efficient and auditable'.
- theturtletalks 10mo agoI’m working on a decentralized marketplace and for now, we tap into the store’s e-commerce platform API to get the inventory, handle cart creation, etc. I commend you for trying to start a standard. Letting the established players establish standards and protocols just gives them a bigger moat and more influence. Pay very close attention to e-commerce and conversational commerce, rent seekers are pushing protocols.
- tsazan 10mo agoAPIs are toll roads. If you need an API key just to read a price, it is not the Open Web. It is a walled garden. We designed this to be permissionless. A text file has no gatekeeper. It bypasses the rent seekers entirely. The standard must belong to the commons, or it becomes just another extraction layer. Keep fighting the good fight.
- theturtletalks 10mo agoI'm working on a Shopify alternative[0] as part of this decentralized marketplace. If adding support for CommerceTXT is not too difficult, I wouldn't mind adding it. 0. https://github.com/openshiporg/openfront https://github.com/openshiporg/openfront
- tsazan 10mo agoThat would be a fantastic first implementation. Openship is exactly the kind of architecture CommerceTXT is built for. Integration is straightforward: it’s essentially just a new 'View' layer. Instead of rendering HTML, you render a .txt endpoint that maps your existing product DB to our fields. I'll head over to your repo and open an Issue to discuss how we can map Openfront's data to the spec. I'd be happy to guide the implementation myself. Let's get this moving!
- theturtletalks 10mo agoSure, that sounds good! Happy to hop on a call to get things moving.
- dehugger 10mo agoBetter idea, how about you just put a link to a csv dump of your inventory data and label it "AI Agents/Scrapers, click here to get all the inventory data", embed that on every page, then call it a day? When you are being scraper there are two possible reactions: 1 - good, because someone scraping your data is going to help you make a sale (discoverability) 2 - bad, work to obfuscate/block/prevent access. In the first case, introducing a complex new standard that few if any will adopt achieves nothing compared to "here's a link for all the data in one spot, now leave my site alone. cheers". In the second case, you actively don't want your data scraped, so why would you ever adopt this? If you are reading all the inventory data into context then you are doing it wrong. Use your LLM to analyze the website and build a mapping for the HTML data, then parse using traditional methods (bs4 works nicely). You'll save yourself a gajillion tokens and get more consistent and accurate results at 1000x the speed.
- tsazan 10mo agoA CSV is a dump of facts. CommerceTXT is a layer of intent and logic. If you give an AI a giant CSV of your whole inventory, you blow the context window before the conversation even starts. If you serve a CSV per product, you still pay for headers and commas without getting any behavioral control. Our spec handles this via @SEMANTIC_LOGIC and @BRAND_VOICE. It’s about how the AI represents your brand, not just the raw numbers. Regarding bs4: mapping HTML to a thousand different store layouts is exactly what we are trying to escape. That is the 'fragility tax'. We are proposing a deterministic fast-lane that bypasses the need for custom scrapers for every single store. You don't want the AI to 'guess' your data. You want it to 'know' your data.
- IgorPartola 10mo agoMeh. I would rather just have the ability to query any given products catalog in a machine-readable way. Any tool or protocol specifically designed for an LLM to consume is in my opinion a design smell. We should instead design proper APIs and protocols usable by all kinds of program and the LLMs can adapt. You are also solving a business problem with a technical solution. Shopify recently announced that they will open up their entire catalog via an easy to use API to a select few enterprise partners. Amazon is doing a similar thing. This is because they do not want you and I to have the ability to programmatically query their catalog. They want to extract money out of specific partners who are trying to enshittify AI chat apps by throwing tons of ads in there. The big movers in the industry could have already easily adopted a similar standard but they are not going to on purpose. On top of you technical issues other commenters are pointing out, I don’t see why this should be in use at all.
- deleted 10mo ago[deleted]