Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
throwdbaaway
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
61.
▲
by
throwdbaaway
1y ago
> But I did not expect something as core as the service managing the leases on the physical EC2 nodes to not have recovery procedure. I guess they don't have a recovery procedure for the "congestive collapse" edge case. I
62.
▲
by
throwdbaaway
1y ago
Actually I am keen to know how Roblox got impacted. Following the terrible Halloween Outage in 2021, they posted 2 years ago about migrating to a cell based architecture in https://corp.roblox.com/newsroom/2023/12&
63.
▲
by
throwdbaaway
1y ago
This looks impressive. As someone who is not familiar with ML, I do have a question -- surely in 2025 there must be a way to schedule a large pytorch job across multiple k8s clusters? EKS and GKE already provide VPC native flat network by d
64.
▲
by
throwdbaaway
1y ago
Huh I just moved to AMDVLK yesterday, after learning that it has 50% more PP on llama.cpp compared to RADV: https://www.reddit.com/r/LocalLLaMA/comments/1nabcek/comment...
65.
▲
by
throwdbaaway
1y ago
That makes sense, thanks for the info. Here's a quick recap of the recent MoE models based on the criteria.. correct activated params: * DeepSeek V3/R1 series * Kimi K2 * GPT-OSS series undercount activated params:
66.
▲
by
throwdbaaway
1y ago
So GLM-4.5 series omits the embedding layer and the output layer when counting both the total parameters and the active parameters: > When counting parameters, for GLM-4.5 and GLM-4.5-Air, we include the parameters of MTP layers but not
67.
▲
by
throwdbaaway
1y ago
You support logprobs, that's wonderful! Fireworks, Synthetic, (ik_)llama.cpp, now I have a quorum.
68.
▲
by
throwdbaaway
1y ago
> use repomap tools on dependencies if creating new code that leverages those dependencies, and have that in context for that work. It seems to me that currently there are 2 schools of thought: 1. Use repomap and/or LSP to help the
69.
▲
by
throwdbaaway
1y ago
https://llm-tracker.info/_TOORG/Strix-Halo has very comprehensive test results for running llama.cpp with Strix Halo. This one is particularly interesting: > But when we switch to longer context, we see something in
70.
▲
by
throwdbaaway
1y ago
With GLM-4.5, we can disable thinking by appending /nothink to the user message. Is that how we suppose to use Octofriend, e.g. "Plan looks good. Go ahead and implement feature X. /nothink"
71.
▲
by
throwdbaaway
1y ago
Actually, you can combine them. When compared to Mac Studio, the main advantage of these Strix Halo boxes is that you still add a bunch of egpu over usb4/oculink, for better PP.
72.
▲
by
throwdbaaway
1y ago
> how the scout -> maverick -> behemoth doesn't scale sparsity according to any formula (less sparse -> sparse -> less sparse) Ah I see. I didn't notice that behemoth has the same sparsity as scout. That seems quite
73.
▲
by
throwdbaaway
1y ago
There is no way that gpt-oss-120b can beat the much larger Kimi-K2-Instruct, Qwen3 Coder/Instruct/Thinking, or GLM-4.5. How did you arrive at this rather ridiculous conclusion? The current sentiment in r/LocalLLaMA is that gp
74.
▲
by
throwdbaaway
1y ago
I thought Kimi K2 uses 8 active experts out of 384? Sparsity should be 48:1. Indeed Llama4 Maverick is the only one that has 128:1 sparsity.
75.
▲
by
throwdbaaway
1y ago
> The bubble itself becomes increasingly intellectually dishonest, increasingly unserious, as it inflates. You started to sound like Dario, who likes to accuse others as intellectually dishonest and unserious. Anyway, perhaps the strict
76.
▲
by
throwdbaaway
1y ago
Another problem is that it works like a slot machine -- sometimes the code is good, most of the time the code is mediocre and full of bugs. Last Friday I spent about $15 in 1 hour using claude code with API key, and the code doesn't re
77.
▲
by
throwdbaaway
1y ago
> The models are then trained to (hopefully) treat the system prompt delimited tokens as more influential on how the rest of the input is treated. I can't find any study that compares putting the same initial prompt in the system ro
78.
▲
by
throwdbaaway
1y ago
I just learnt the `nl -ba` trick from Codex. Claude Code is most likely doing the same.
79.
▲
by
throwdbaaway
1y ago
As shared by Simon in https://news.ycombinator.com/item?id=44176523 , a better agent will prepend the line numbers as a workaround, e.g. Claude Code: 54 def dicts_to_table_string( 55 headings: List[str], dicts:
80.
▲
by
throwdbaaway
1y ago
> Here's the technical takeaway: Never use CASCADE deletes on critical foreign keys. Set them to NULL or use soft deletes instead. It's fine for UPDATE operations, but it's too dangerous for DELETE ones. The convenience of
81.
▲
by
throwdbaaway
1y ago
It is possible that the tasks you gave to the model previously were just about easy enough for it to handle, while the few failing tasks you gave recently were a bit too tough for the model, thus it had to cheat. For the exact same task, so
82.
▲
by
throwdbaaway
1y ago
Ah right, my bad. Somehow I thought the prompt was only: Say just 'hi' while the "without any extra words or explanations" part was for the readers of your comment. Perhaps kubb also made a similar mistake. I us
83.
▲
by
throwdbaaway
1y ago
o3-mini-2025-01-31 with high reasoning effort replied with "Hi" after 448 reasoning tokens. gpt-4.5-preview-2025-02-27 replied with "Hi!"
84.
▲
by
throwdbaaway
1y ago
For the code review use case, maybe can try to create the diff with something like `git diff -U99999`, and then send only the diff.
85.
▲
by
throwdbaaway
1y ago
This prompt works fine with Qwen2.5-Coder-32B-Instruct-Q4_K_M: Add a line number prefix to each line, stopping at line 27. What's on line 27 of this program?
86.
▲
by
throwdbaaway
1y ago
https://github.com/google/mysql-tools/tree/master/old/mysql-... - Google used to do that too
87.
▲
by
throwdbaaway
2y ago
With gentoo, if you allocate let's say 20G to / on ext4, then you can quite easily run into this issue. /usr/src/linux will use about 30% of the space and 10% of the inodes. /var/db/repos/gentoo
88.
▲
by
throwdbaaway
2y ago
I am running the base model of Qwen2.5-Coder-32B with llama.cpp. It can only do completion, it can't chat. Where did you get that information from?
89.
▲
by
throwdbaaway
2y ago
Using https://github.com/kvcache-ai/ktransformers/ , an intel/amd laptop with 128GB RAM and 16GB VRAM can run the IQ4_XS quant and decode about 4-7 token/s, depending on RAM speed and context size. Using
90.
▲
by
throwdbaaway
2y ago
This is where the base open models can really shine, before they got lobotomized by the instruction fine-tuning. For example, this is the completion I get with DeepSeek-Coder-V2-Base and greedy decoding: Chat: On the day of June 4th 1989, i
More ›