3 ms·
Cheaper/faster is coming for sure. Model on a custom silicon: https://chatjimmy.ai/ https://chatjimmy.ai/ 1-bit models that run on a CPU: https://github.com/m
by ForHackernews 1mo ago
Cheaper/faster is coming for sure.
Model on a custom silicon: https://chatjimmy.ai/ https://chatjimmy.ai/
1-bit models that run on a CPU: https://github.com/microsoft/BitNet https://github.com/microsoft/BitNet
- deleted 1mo ago[deleted]
- saturn8601 1mo agoBlazing fast...but terrible. Put Sol on silicon but will still need access to the internet...so it will be somewhat slow anyway
- Tuna-Fish 1mo agoIt's terrible because it's Llama 3.1 8B. It's such a crappy model because HC1 was a relatively low budget proof of concept. The team that built is working on a better implementation.
- saturn8601 1mo agoNot sure its that to be honest. It seems like maybe its not installed correctly or is like GPT-1/GPT-2 quality? I asked it who is [famous actress] and it started talking about some random person from Mexico with a completely different name. The speed is intoxicating but i'd like for it to actually answer based on what I asked. Thats why I think something might be wrong in implementation on this site. Edit: I went back and retested it. It revealed that its data is from July 2021 which explains partially why It couldn't talk about the actress I asked about (she exploded in popularity in 2026 but was still a professional actress in 2021 so idk). I then went and asked questions about a very popular actress and movie in 2010. It got it much better but still hallucinated a ton of details about her. I guess I didn't fully understand what you were saying. Sorry about that! I look forward to their next releases because upon thinking about what I experienced here, I am super excited to see this progress further!
- Tuna-Fish 1mo agoIt has no internet access and only 8B 3-bit parameters. That's simply not enough for it to compress all that much knowledge.