4 ms·
The demo: https://chatjimmy.ai/ https://chatjimmy.ai/
by whythismatters 2mo ago
The demo: https://chatjimmy.ai/ https://chatjimmy.ai/
- nsxwolf 2mo agoIt doesn’t believe it’s running on that chip, it’s arguing with me
- shaewest 2mo agoIt's running a very small, non-reasoning model at the moment. But more generally, almost all LLMs argue on the hardware/model they are/are on.
- dumberquestions 2mo agoWhich model? Or how many active parameters?
- _whiteCaps_ 2mo agoLlama 3.1 8B model
- dumberquestions 2mo agoSo this demo is around 90 times faster than typical speeds for the same model at openrouter, and around 30 times faster than the absolute fastest option available (Groq).
- anthonypasq 2mo agoim assuming energy expenditure is substantially lower as well
- mdp2021 2mo agohttps://taalas.com/h-content/uploads/2026/02/graph.png https://taalas.com/h-content/uploads/2026/02/graph.png
- Gander5739 2mo agohttps://xkcd.com/1162/ https://xkcd.com/1162/
- deleted 2mo ago[deleted]
- metadat 2mo agoWhat would tokens/sec performance look like for a reasoning model? An order of magnitude slower?
- penagwin 2mo agoReasoning models are the same speed. They’re just post trained with RL to do CoT inside tags like <thinking></thinking> before a tag like <response></response> There’s no difference in the inference implementation, parameter count, or speed.
- paytonjjones 2mo agoThere's a difference in the latency distribution between when you submit a query and you see the response, which is what the comment is (clumsily) asking about. But yeah, there are a lot of factors, so it's hard to answer, and tokens/s isn't the right question.
- wmf 2mo agoAIs don't intrinsically know anything about themselves so they often give wrong answers to such questions. This can be fixed by putting info in the system prompt but they may consider it a waste of tokens since most usage doesn't benefit from that information.
- anigbrowl 2mo agoThat proves it's conscious! (/s!)
- itvision 2mo agoOMFG this thing is fast.
- phoh 2mo agoits fast but try to get it to give you pi to 50 decimal places. it didnt go well for me.
- walrus01 2mo agoI think the same exact model running on CPU-only and RAM, or a small GPU, would do about the same? It's quite an old model now and small, you could throw a GGUF into llama-server or something for a side by side comparison. https://huggingface.co/meta-llama/Llama-3.1-8B https://huggingface.co/meta-llama/Llama-3.1-8B As I remember just about any english language model from mid 2024 and earlier didn't even do well if you asked it to count sequentially from 0 to 100, nevermind calculating stuff.
- estearum 2mo agoThat's not how LLMs work
- wxw 2mo agoI freakin' love this demo. It feels magical.
- VBprogrammer 2mo agoI had the same reaction but then I showed it to my partner. She completely didn't get it, in her words "how can it be thinking of a good answer when it's that quick?" I tried to explain but I fear were probably going to be adding artificial sleeps to these things to convince the masses it's doing something clever.
- varun_ch 2mo agoto be fair, the model used for Chat Jimmy is not very smart, but the world where it is smart is very interesting. It’s going to be really crazy when the bottle neck for agents is the speed of the tool calls rather than the speed of inference. Imagine an agent interacting with the terminal near instantly…
- ricardobeat 2mo agoI had the chance to try out MiMo v2.5 Pro Ultraspeed (600-1000tok/s) for a couple weeks and it is amazing. Developing software becomes 95% about intent and requirements. Can’t wait for the next iteration of that.
- axus 2mo agoI asked it some old hardware command line questions I'd recently asked Gemini, it hallucinated parts of the answer. The characters in the 3-act Shakespearean play had very little depth, many of the names were similar, and they were not very smart, but the simple plot was cohesive.
- XCSme 2mo agoWait, is it even thinking? Or is it an instant model?
- 2mo ago
- walrus01 2mo agoI know it's a relatively tiny model, but damn, is that thing fast. It also mostly passes the "schlong" test https://pastes.io/YcxSi8Fp https://pastes.io/YcxSi8Fp
- thoughtpeddler 2mo agoI didn't realize there was a SchlongBench™ (but of course there is). What's it test? (asking seriously)
- walrus01 2mo agoThere isn't SchlongBench(TM) yet, it's a specific question I've been asking of differently sized models as a randomly chosen gauge of how much less commonly used knowledge is perma-baked into it. In this case a question about a specific yiddish origin slang term. Small/bad models don't know it's from middle high german or Yiddish and get its origin and meaning totally wrong (or it runs into model censorship related to slang related to the male anatomy). It's also a question I have found will cause models that don't know what it is to go off quickly in a direction of hallucination trying to explain it, so the hallucination is evident very quickly starting from the first ever prompt issued with 0 context fill. Example: I had a model write four detailed supposedly-accurate sounding, grammatically correct paragraphs saying its origin is from AAVE (African American Vernacular English), which it most certainly is not You could do the same by picking any topic that is very rarely discussed in conversation, some esoteric and narrow piece of knowledge and asking the model about it.
- thoughtpeddler 2mo agoOh ya, this is like the approach from the Incompressible Knowledge Probes [0] paper - smart! [0] Incompressible Knowledge Probes: Estimating Black-Box LLM Parameter Counts via Factual Capacity [https://arxiv.org/abs/2604.24827 https://arxiv.org/abs/2604.24827]
- AussieWog93 2mo agoI read the paste, it got the etymology wrong, no? Schlong comes from shlang (snake), not shlemp (is this even a word? I don't speak Yiddish but couldn't find it on Google). Oxford also claim that its first recorded use was from the 60s, not the 20s; https://www.oed.com/dictionary/schlong_n?tl=true https://www.oed.com/dictionary/schlong_n?tl=true
- hendurhance 2mo agoI understand the appeal due to the speed
- senderista 2mo agoWow, feels like Google web search in 1999.
- joshvm 2mo agoIf you still want the experience, go and browse McMaster Carr. Wizards designed that website.
- senderista 2mo agoOh I have, though not for a while.
- eglintondust 2mo agoI'm inspired by this website. It's incredible.
- jodrellblank 2mo agoor LiveGrep fast search of the Linux kernel source code with regex support: https://livegrep.com/search/linux https://livegrep.com/search/linux
- senderista 2mo agoWow, I want something like that for my company's codebase.
- deleted 2mo ago[deleted]
- anigbrowl 2mo ago15,000 tok/s ....damn. It's very impressive notwithstanding its limitations.
- XCSme 2mo agoWow, that's instant, crazy.
- brikym 2mo agoThe speed is awesome, in the true sense of the word. It's great at knowledge and basic stuff but the output is complete junk for anything concerning new facts or slightly esoteric topics.
- ecshafer 2mo agoThat is insanely fast. I had it generate a basic C FFT library that can handle multi-dimension arrays, and it was instant.
- appplication 2mo agoThis is the coolest LLM thing I’ve seen since the original ChatGPT announcement a few years ago. IMO much more impressive than marginal gains of frontier models.
- zhoge 2mo agoThis is the answer I got after asking it twice what's taalas (second time hinting that it's a chip startup): After a quick search, I found that Ta'ala is actually a Canadian chip startup that produces artisanal, high-end potato chips. They offer a range of unique and creative flavor combinations, often featuring Canadian and international ingredients. Ta'ala is known for its high-quality, small-batch potato chips made with premium ingredients and care. The company is committed to creating unique and delicious flavor profiles that showcase the best of Canadian ingredients and cuisine. Is this the Ta'ala you were thinking of?
- mintflow 2mo agotry let it to get a brief of france history which being reading a while hit the button and then the brieft jump into my eye Generated in 0.051s • 14,092 tok/s Impressive... Given gpt 5.5 was very good to me and gpt 5.6 series seems not boost too much, i kinda like the way bake the model weight to the chip, and connect multiple chip to serve the large scale model and allow respin some parts(ROM like?) to do model weight update, maybe this seems sustainable, the future is exciting
- lelanthran 2mo ago> try let it to get a brief of france history which being reading a while hit the button and then the brieft jump into my eye WTF is this?
- mrheosuper 2mo agolooklike the training material is stopped at around July 2022, a little too outdated.
- hahahaa 2mo agoMade me an entire app in 84ms lol
- calgoo 2mo agoI was thinking the other day if we could use something like this "old" 8B model, and run 20 or 30 calls at the same time (or in sequence, we wont notice) and use and use the best result. Basically tiny agents that do tiny things but VERY fast.
- deviation 2mo agoThis is the only demo of 2026 which has blown my mind. If we can get to this speed with reasoning models, man... I can't even imagine the impact.
- tiborsaas 2mo agoLooking at the history of technology, it's a question of when do we get there.