Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
XCSme
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
91.
▲
by
XCSme
2mo ago
With default config via Ollama and 65k context I get 50tps on a 3090.
92.
▲
by
XCSme
2mo ago
They both have horizontal scroll on mobile...
93.
▲
by
XCSme
2mo ago
So where does the guarantee stop? At the GPU driver level? Firmware level? What if the GPU has a custom bios flash that somehow logs the unencrypted prompts?
94.
▲
by
XCSme
2mo ago
Why do I feel like there haven't been any hardware releases since the Macbook Neo? Market is so slow, probably due to the DRAM shortages. No new GPUs, CPUs, technologies, etc...
95.
▲
by
XCSme
2mo ago
Yeah, Sol models have higher in/out cost but are incredibly token efficient. Also, in those tests Sol Low did better, but you can also compare the price vs Sol High, then it's getting a bit closer. So Grok 4.6 is still not the bes
96.
▲
by
XCSme
2mo ago
Open source doesn't matter if someone else is running it, right? They can change it? As long as the prompt is not encrypted at some point, and I don't think LLMs can run on encrypted prompts, then it can be read.
97.
▲
by
XCSme
2mo ago
This is not a solution for now, just a moonshot solution for maybe 5-10 years from now. There's no immediate fix for the RAM supply issue, but as long as it's manufacturing issue, not a raw materials issue, time always solves it.
98.
▲
by
XCSme
2mo ago
Grok 4.6 vs Sol 5.6 vs Opus 5: https://aibenchy.com/compare/openai-gpt-5-6-sol-low/x-ai-gro...
99.
▲
by
XCSme
2mo ago
Does really well and ~2x cheaper than Qwen3.8 2.4T, they have same pricing but grok is around 2x more token efficient: https://aibenchy.com/compare/qwen-qwen3-8-2-4t-a95b-low/x-ai...
100.
▲
by
XCSme
2mo ago
Interestingly, the high variant does a lot worse and failed to generate a valid SVG (and the low variant use more tokens than the high one, so maybe their reasoning efforts are not working properly). The solar system animation is also the c
101.
▲
by
XCSme
2mo ago
That's a really cool hamster [0], unfortunately it's really expensive now, 2x more expensive than Grok 4.6[1]. [0]: https://aibenchy.com/compare/x-ai-grok-4-6-high/bytedance-se... [1]: https://
102.
▲
by
XCSme
2mo ago
I don't understand how this works? Is it another proxy on top? What stops the provider from reading/storing the prompts at the LLM execution level?
103.
▲
by
XCSme
2mo ago
Well, before AI the RAM prices were cheapest ever, even if cloud providers were still in demand for hosting and cloud computing. I hope the supply will increase at some point, I doubt cloud providers will be able to buy everything, especial
104.
▲
by
XCSme
2mo ago
I "trust" what they say on OpenRouter for the provider, for some it says they retain prompts, for other that they retain but can also use them for training. It's not any crazy IP, just my own benchmarks/tests, once they
105.
▲
by
XCSme
2mo ago
Again, I will wait until there's a provider that doesn't train on prompts before I will benchmark.
106.
▲
by
XCSme
2mo ago
Thanks for sharing. Yeah, that was my experience too (2x or 3x is indeed considerably faster), but not workflow-changing faster at 45tps baseline, especially for asynchronous tasks (which is my goal with a local 3090, to just let it do thin
107.
▲
by
XCSme
2mo ago
True, but the ad is to my own product, there are no advertisers, maybe there will be at some point, but if they were, that would barely cover the costs of testing the models, and likely never get a ROI on the hundreds of hours I've spe
108.
▲
by
XCSme
2mo ago
> Your browser is automatically speaking HTTP for you so that you don't have to. Yes, but it's not filling in the forms or clicking the buttons for me. HTTP is just infrastructure. Are LLMs infrastructure? Are we too maybe infr
109.
▲
by
XCSme
2mo ago
Hazardous waste? My idea was more like you get some pre-made chips, that you can maybe assemble together configure at home with your desired models. Maybe each one of this chip is a layer, so you can stack as many layers as you want.
110.
▲
by
XCSme
2mo ago
So, on a crude calculation, for a 3090 with 936.2 GB/s, a model that has 20 GB of active params would run at 45tps and one with 3GB active params at 300tps? In practice, I don't think I saw over 100tps on a 3090, for a local 20-30
111.
▲
by
XCSme
2mo ago
I remember running both qwen 30b-a3b and 27b on my 3090, and on the initial test, the 27b was only like 2x slower.
112.
▲
by
XCSme
2mo ago
Definitely, LLMs are highly ineficient now. The diffusion models are interesting, but those also seem hacky. I think the next form of AIs will be simpler and more abstract. The building blocks of our brain don't have the notion of a &q
113.
▲
by
XCSme
2mo ago
Isn't a phone call like 10x more taxing than an email reply?
114.
▲
by
XCSme
2mo ago
Will we have any way to know if a submission was done by a human or a bot?
115.
▲
by
XCSme
2mo ago
The eternal fight between bots and anti-bot systems. The difference now is that big companies themselves promote/offer bots, but they also don't like to be scraped and use captchas. What do we do now? Is it allowed to use automate
116.
▲
by
XCSme
2mo ago
I am asking mostly for running on a 3090. I think the tps difference between them (both fitting in vram) won't be more than 2x in practice. I would happily take 20tps over 40tps, if the model gets 3x more correct answers.
117.
▲
by
XCSme
2mo ago
This is just temporary though, right? With the benefit of LLMs already being proven, in a couple of years we will have vastly better hardware for inference I guess. I feel like now hardware is stagnating a bit, because the software side has
118.
▲
by
XCSme
2mo ago
Is that the case? If the entire model fits in vram, won't the tps be comparable?
119.
▲
by
XCSme
2mo ago
Oh, good to know, I just quickly tested and published the results. I will add model sizes (total/active params) for each model, good point.
120.
▲
by
XCSme
2mo ago
I did, avertisment is a big word, as I gain nothing from the traffic, I run the website for myself, and some other people find it useful too. Happy to hear what would make the website more useful.
More ›