Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
lhl
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
11 ms
·
61.
▲
by
lhl
2y ago
Yeah, it seems to not be able to execute the tool calling properly. Maybe it's a bad interaction w/ it's own async calling ability or something else (eg, how search and code interpreter can't seem to run at the same time
62.
▲
by
lhl
2y ago
I've been using GenSpark.ai for the past month to do research (its agents usually does ~20 minutes, but I've seen it go up to almost 2 hours on a task) - it uses a Mixture of Agents approach using GPT-4o, Claude 3.5 Sonnet and Gem
63.
▲
by
lhl
2y ago
Besides not having ES markings, It is a retail serial and stepping in dmidecode, so that's unlikely.
64.
▲
by
lhl
2y ago
1: 33586.74 2: 47371.93 4: 65870.07 With `likwid-bench -i 100 -t load -w M0:5GB:1 -w M1:5GB:1 -w M2:5GB:1 -w M3:5GB:1 -w M4:5GB:1 -w M5:5GB:1 -w M6:5GB:1 -w M7:5GB:1` we get 187976.60 Obvious there's a bottleneck either going on somewh
65.
▲
by
lhl
2y ago
w/ likwid-bench S0:5GB:8:1:2, 129136.28 MB/s . At S0:5GB:16:1:2 184734.43 MB/s (this is the max, S0:5GB:12:1:2 is 186228.62 and S0:5GB:48:1:2 is 183598.29 MB/s) - According to lstopo my 9274F has 8 dies with 3 cores on e
66.
▲
by
lhl
2y ago
Yeah, that channels are populated correctly. As you can see from the mlc-results.txt, the latency looks fine: mlc --idle_latency Intel(R) Memory Latency Checker - v3.11b Command line parameters: --idle_latency Using buffer size
67.
▲
by
lhl
2y ago
Last fall I built a new workstation with an EPYC 9274F (24C Zen4 4.1-4.3GHz, $2400), 384GB 12 x 32GB DDR5-4800 RDIMM ($1600), and a Gigabyte MZ33-AR0 motherboard. I'm slowly populating with GPUs (including using C-Payne MCIO gen5 adapt
68.
▲
by
lhl
2y ago
There's been a lot of discussion in https://www.reddit.com/r/DataHoarder/ Here's documentation on independent backup efforts of various government websites: https://www.reddit.com/r/
69.
▲
by
lhl
2y ago
Give Section 3 of the DeepSeek-V3 paper a read. The discuss their HAI-LLM framework and have a pretty in-depth description of their DualPipe algorithm and how it compares to other pipeline bubbles. They also describe how they work around NV
70.
▲
by
lhl
2y ago
I think you haven't been looking too hard in that case. Here is the R1 paper: https://arxiv.org/abs/2501.12948 You can find more papers from the attached author: https://arxiv.org/search/cs?se
71.
▲
by
lhl
2y ago
> Looking at his recent track record[1] One might argue he's had a pattern for even longer. While he did do some early hypervisor glitching, even his PS3 root key release was basically just applying fail0verflow's ECDSA exploit
72.
▲
by
lhl
2y ago
Fireworks, Together, and Hyperbolic all offer DeepSeek V3 API access at reasonable prices (and full 128K output) and none of them will retain/train on user submitted data. Hyperbolic's pricing is $0.25/M tokens, which is actu
73.
▲
by
lhl
2y ago
Yeah, I think anyone w/ old Jetsons knows what it's like to be left high and dry by Nvidia's embedded software support. Older models are basically just ewaste. Since the Digits won't be out until May, I guess there'
74.
▲
by
lhl
2y ago
But you pay triple damages if you knowingly vs unknowingly violate a patent (35 U.S.C. § 284). Of course, everything is patented, so, engineers are just told to not read patents.
75.
▲
by
lhl
2y ago
Yeah, the M4 Max actually has pretty decent MBW - 546 GB/s (cheapest config is $4.7K on a 14" MBP atm, but maybe there will be a Mac Studio at some point). The big weakness for the Mac is actually the lack of TFLOPS on the GPU - t
76.
▲
by
lhl
2y ago
I don't know of any non-soldered memory Strix Halo devices, but both HP and Asus have announced 128GB SKUs (availability unknown). For LLM inference, basically everything works w/ ROCm on RDNA3 now (well, Flash Attention is via Tr
77.
▲
by
lhl
2y ago
Getting to 50 tok/s for a big model requires not just memory, but also memory bandwidth. Currently, 1TB/s of MBW will get a 70B Q4 (~40GB) model to about 20-25 tok/s. The good thing is models continue to get smarter - today&#
78.
▲
by
lhl
2y ago
I published this as a comment as well, but it's probably worth nothing that the ChatGPT water/power numbers cited (the one that is most widely cited in these discussions) comes from an April 2023 paper (Li et al, arXiv:2304.03271)
79.
▲
by
lhl
2y ago
I tested Phi-4 with a Japanese functional test suite and it scored much better than prior Phis (and comparable to much larger models, basically in the top tier atm). [1] The one red-flag w/ Phi-4 is that it's IFEval score is relat
80.
▲
by
lhl
2y ago
I've been subscribed to Kagi since trying it out early on in 2023 and have been paying the $10/mo for unlimited usage since they introduced that tier (if you are a heavy search user, I'd agree the $5/tier "limited&q
81.
▲
by
lhl
2y ago
The memory bandwidth has not been announced for this device. It's probably going to be more appropriate to compare vs a 128GB M4 Max (410-546GB/s MBW) or an AMD Ryzen AI Max+ 395 (yes, that's its real name) at 256GB/s of
82.
▲
by
lhl
2y ago
So, I gave this to ChatGPT-4o, changing the initial part of the prompt to: "Write Python code to solve this problem. Use the code interpreter to test the code and print how long the code takes to process:" I then iterated 4 times
83.
▲
by
lhl
2y ago
Sort of funny for the lack of local options considering that Mozilla funds llamafile even. Hopefully they allow some API integration, if they are using standard OpenAI API calls, it should be easy to enable swapping the endpoint. Also, whil
84.
▲
by
lhl
2y ago
For those interested in some more testing of Qwen's censorship (including testing dataset, testing to compare english vs chinese responses, and a refusal-orthoganlized version of Qwen2): https://huggingface.co/blog/
85.
▲
by
lhl
2y ago
It depends on what you mean by "this." MLC's catch is that you need to define/compile models for it with TVM. Here is the list of supported model architectures: https://github.com/mlc-ai/mlc-llm/
86.
▲
by
lhl
2y ago
My testing has been w/ a Lunar Lake Core 258V chip (Xe2 - Arc 140V) on Arch Linux. It sounds like you've tried a lot of things already, but case it helps, my notes for installing llama.cpp and PyTorch: https://llm-track
87.
▲
by
lhl
2y ago
I've recently been poking around with Intel oneAPI and IPEX-LLM. While there are things that I find refreshing (like their ability to actually respond to bug reports in a timely manner, or at all) on a whole, support/maturity actu
88.
▲
by
lhl
2y ago
Just an FYI, this is writeup from August 2023 and a lot has changed (for the better!) for RDNA3 AI/ML support. That being said, I did some very recent inference testing on an W7900 (using the same testing methodology used by Embedded L
89.
▲
by
lhl
2y ago
> The iGPU on Strix Point can do 9tflops of good old fp32. Unsure what the TFlOps drops to on the new Jetson here as it goes from INT8->FP32. But probably still has a very healthy lead. Not as much as you think. Per their developer bl
90.
▲
by
lhl
2y ago
The safetensors are in the phi-4 folder of the very repo you linked in your OP.
More ›