5 ms·
Show HN: Create-LLM – Train your own LLM in 60 seconds
https://medium.com/@theaniketgiri/three-months-ago-i-wanted-to-train-my-own-llm-b796aae9aa94 https://medium.com/@theaniketgiri/three-months-ago-i-wanted-...
- kk58 1y agoDoes this work on mac
- theaniketgiri 1y agoYep, works fine on Mac. Try the nano or tiny templates if you want quicker training runs
- efilife 1y ago2 questions: how much of this project is AI generated and how much of only the readme is AI generated?
- theaniketgiri 1y agoMostly the repetitive stuff like README generation and pushing code with meaningful commit messages was handled by AI. The actual work and logic were done by me.
- joshribakoff 1y agoWhat about the commit that added tens of thousands of lines of markdown claiming to be an AI summary? Or the meaningful commit message of “.” And the commit editing 1,000s of lines of python code mislabeled as a docs change?
- theaniketgiri 1y agoTotally fair question! Docs / Markdown: AI handled repetitive stuff like READMEs and summaries. Core logic / Python: fully written by me. Commit messages: some minimal ones just for quick iterations — the real work is in the code. AI helped with boilerplate so I could ship faster; all functionality is hand-crafted.
- computerthings 1y ago[dead]
- joshribakoff 1y agoIf the AI did the boilerplate that implies it was not fully written by you. The “meaningful commit messages” — again are a single period as the message for a single commit for the entire python portion of the codebase. My question was rhetorical. Whether the AI did it or a human did, it burns credibility to refer to things that don’t exist (like “meaningful commit messages”)
- teruakohatu 11mo agoHacker News is a better place when we don’t attack people sharing their work. Your point was made. Well done to the author for shipping code. I look forward to trying it out.
- Grimblewald 11mo ago> for sharing their work If it was their work your point would hold.
- theaniketgiri 11mo agoTo clarify the AI question once and for all: What AI did: - Generated README templates (boilerplate markdown) - Suggested commit messages (I didn't always edit them) - Helped with documentation structure What I wrote: - All Python training logic (train.py, trainer.py, callbacks) - All model architectures (gpt.py, tiny.py, small.py, etc.) - Tokenizer integration - Data pipeline - CLI scaffolding (
- biinjo 11mo agoDon’t feed the trolls. This was your idea and you made something that works. Who cares if its (partially) done by AI. Whomever is taking offense by people using AI for coding, is just having a hard time adapting to the current state of affairs. It’s here, it’s happening. Try the project, if you like it thats great, if you don’t then move on. And if you don’t intent to try it for whatever reason that’s fine as well but don’t be salty to the OP for sharing their passion project.
- darepublic 1y agoI don't quite understand how you get from this: > I wanted to understand how these things work by building one myself. Directly to this: What if training an LLM was as easy as npx create-next-app? I mean that the second thought seems to be the opposite of the first (what if the entirety of training llm was abstracted behind a simple command)
- theaniketgiri 1y agoGreat question - I should've been clearer. When I started, I wanted to understand LLMs deeply. But I hit a wall: tutorials were either "hello world" toys or "here's 500 lines of setup before you start." What I needed was: "give me working code quickly, THEN let me modify and learn." That's what create-llm does. It scaffolds the boilerplate (like create-next-app), so you can spend time learning the interesting parts: - Why does vocab size matter? (adjust config, see results) - What causes overfitting? (train on small data, see it happen) - How do different architectures perform? (swap templates, compare) It's "easy to start, deep to master." The abstraction gets you running in 60 seconds, then you dig into the code
- seg_lol 1y agoThe blogpost is some of the best LLM greentext I have seen for targeting the hn hivemind. Everything about this is :chefs kiss:
- theaniketgiri 11mo agoThanks! The blog post is just my honest journey - spent way too much time trying to understand LLMs, figured others had the same frustration. If you try create-llm, would love your feedback. Always looking to make it better.
- 3abiton 11mo agoHow does this differ from nanochat?
- theaniketgiri 11mo agoGood question! I think you mean nanoGPT (Karpathy's minimal GPT implementation)? Key differences: nanoGPT: - Minimal reference implementation (~300 lines) - Educational code for understanding transformers - Requires manual setup and configuration - Great for learning the internals create-llm: - Production-ready scaffolding tool (like create-next-app) - One command: npx create-llm → complete project ready - Multiple templates (nano/tiny/small/base) - Built-in validation (warns about overfitting, vocab mismatches) - Includes tokenizer training, evaluation, deployment tools - Auto-detects issues before you waste GPU time Think of it as: nanoGPT is the reference, create-llm is the framework. nanoGPT teaches you HOW it works. create-llm lets you BUILD with what you learned. You can actually use nanoGPT's architecture in create-llm templates - they're complementary tools!
- Grimblewald 11mo agoUnlike nanochat this is purely vibe-coded, improving vibes by 110%, with 112x more emoji. A key innovation that gets to the heart of the problem is that this project stores python files as strings in typscript files to help improve workflows. I imagine the author solved this engineering challenge to overcome existing limitations\emdash more efficient, interpretable, and maintainable code\emdash in existing projects.
- theaniketgiri 11mo agoThe Python-in-TS bit made me smile But to clarify, it’s a standard TypeScript CLI — no such hacks involved, just template-based generation.
- Grimblewald 11mo agoOk, but there is no reason to bake it into the TS scripts. You could write the python scripts and package them using standard tools. In my experience only an LLM would do that, since it makes sense to generate the code and templates to insert in one go. However, if a human were to do it, the python scripts would be their own files and they would be bundled / read in as strings when/as required. A gigantic lump of text in a string makes no sense in human paradigms, even if it makes perfect sense for an LLM to do it. For humans it is incredibly hostile to update and maintain. As a side note, without looking it up, on your device, what is the process for typing an emdash?
- theaniketgiri 11mo agoThanks everyone for the feedback and discussion. For those asking technical questions - happy to help! The tool works on Mac/Linux/Windows, check the README for setup. For those concerned about the architecture - it follows standard scaffolding patterns (create-next-app, etc). TypeScript CLI generates Python projects. 82+ stars in 24 hours - grateful for everyone trying it out. Keep the feedback coming!
- potamic 11mo agoHow did you test this? Did you train something?
- theaniketgiri 11mo agoYeah, I did! I trained a few small ones — mostly the “nano” and “tiny” templates (a few million params) on datasets like Shakespeare and Alpaca. The goal was to make sure the training loop, tokenizer, and evaluation all worked smoothly. Didn’t go for massive models — more about making the whole setup process quick and reliable. You can actually train the nano one on CPU in a few minutes just to see it working.
- mkrishnan 11mo agoThis is great initiative. don't let anyone here discourage from doing something great like this.
- theaniketgiri 11mo agoThanks a lot, really appreciate that Building something new always gets mixed reactions, but messages like yours keep me going.