Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
coder543
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
11 ms
·
241.
▲
by
coder543
2y ago
Altera seems to have some level of independence from Intel now: https://www.datacenterdynamics.com/en/news/fpga-company-alte...
242.
▲
by
coder543
2y ago
As I mentioned in my comment, you can also observe the escape analysis from the compiler and know whether your code will allocate or not, and you can make adjustments to the code based on the escape analysis. I was making the point that y
243.
▲
by
coder543
2y ago
> unless using a totally separate toolchain w/ "CGo". CGo is built into the primary Go toolchain... it's not a 'totally separate toolchain' at all, unless you're referring to the C compiler used by CGo
244.
▲
by
coder543
2y ago
tl;dr Async/await is what the Python community has settled on, so I don’t think there’s any point in resisting it now, but the fact that it happened is a historical curiosity from my point of view. ----- I’m not the person you replied
245.
▲
by
coder543
2y ago
How exactly are you trying to train and deploy this YOLO model? What kind of accuracy are you seeing against the validation set at the end of the training process?
246.
▲
by
coder543
2y ago
Most of the RAM usage would likely just be executable files that are mmap’d from disk.. not “real” RAM usage. But, also, the 1000 users in question wouldn’t all be connected at the same time… and I honestly doubt they would all be assigned
247.
▲
by
coder543
2y ago
Preventing people from using their preferred tools — tools which are extremely widely used in the real world — does not seem like a useful application of time and effort.
248.
▲
by
coder543
2y ago
No… it’s not. To quote the message earlier in the thread, that message said “everyone with >100MB of disk usage on the class server was a VSCode user.” 100MB * 1000 users is how the person I responded to calculated 100GB, which is storag
249.
▲
by
coder543
2y ago
100GB of storage is… not much for 1000 users. A 1TB NVMe SSD is $60. So 100GB is a total of $6 of storage… or about half a penny per user. And that’s for SSD storage… an enterprise-grade 14TB hard drive is only $18/TB on Amazon right
250.
▲
by
coder543
2y ago
I'm somewhat surprised neither this article nor the previous one mention anything about the Florence-2 model series. I had thought that Florence-2 was not just surprisingly capable for this kind of work, but also easily fine-tunable
251.
▲
by
coder543
2y ago
I can't imagine why anyone would run it unquantized, but there are some laptops with the more than 70GB of RAM that would be required. It's not that it can't be done... it's just that quantizing to at least 8-bit seems t
252.
▲
by
coder543
2y ago
32B models are easy to run on 24GB of RAM at a 4-bit quant. It sounds like you need to play with some of the existing 32B models with better documentation on how to run them if you're having trouble, but it is entirely plausible to run
253.
▲
by
coder543
2y ago
That's been happening consistently for over a year now. Small models today are better than big models from a year or two ago.
254.
▲
by
coder543
2y ago
I used Continue before Cursor. Cursor’s “agent” composer mode is so much better than what Continue offered. The agent can automatically grep the codebase for relevant files and then read them. It can create entirely new files from scratch.
255.
▲
by
coder543
2y ago
Location: Birmingham, AL Remote: Yes Willing to relocate: Yes Technologies: Go, Rust, TypeScript, React, Postgres, Kafka, AWS, GCP, etc. Résumé/CV: https://drive.google.com/file/d/1VNC272B3n7ZEfppM
256.
▲
by
coder543
2y ago
I have never had a chance to use C# professionally, but it was one of the first languages I taught myself when I was first learning programming when I was a kid. I have a lot of fond memories of it, and I hear so many positive things about
257.
▲
by
coder543
2y ago
I think they wanted to define free functions, not figure out a way to use functions without a class name, but I could be wrong.
258.
▲
by
coder543
2y ago
Mistral Small 3 is roughly comparable in capabilities to 4o-mini (apart from 4o-mini's support for multimodality)... o1-mini was already better than GPT-4o (full size) for tasks like writing code, and this is supposedly better than o1
259.
▲
by
coder543
2y ago
> 2.0 is the full model Not quite. "2.0 Flash" is also called 2.0. The "Pro" models are the full models. But, I love how they have both "gemini-exp-1206" and "gemini-2.0-flash-thinking-exp-01-21".
260.
▲
by
coder543
2y ago
> seems likely that its a 70B-llama distillation since that's what AWS offers on Bedrock I think you misread something. AWS mainly offers the full size model on Bedrock: https://aws.amazon.com/blogs/aws/dee
261.
▲
by
coder543
2y ago
This doesn't prove anything at all . Of course the toolchain has to be built somehow. Some toolchains use make to do that, rather than depending on the previous version of the toolchain's build system. Some toolchains are written
262.
▲
by
coder543
2y ago
Are those Makefiles doing anything more than calling "go build" and "cargo build"? Because if they're still using the language-specific build tools and dependency management systems, then I think you would find that
263.
▲
by
coder543
2y ago
Make is primarily used with C and C++. It is not commonly used in Java, Rust, Go, NodeJS, or hardly anything besides C and C++. Make is not "generally used with other languages".
264.
▲
by
coder543
2y ago
No problem. I haven’t had a chance to do anything hardware related for a long time, so it’s fun to think about hardware problems again. On the topic of extending battery life mentioned in the article, one relatively straightforward thing to
265.
▲
by
coder543
2y ago
Wow! This is some phenomenal engineering! > Another area that stumped me is how to shut the power off 100% on the device, so that it can remain “off” for weeks or months. This is actually a pretty solvable problem... https://c
266.
▲
by
coder543
2y ago
Sure, but that default batch size would only matter if the person in question was actually generating and measuring parallel requests, not just measuring the straight line performance of sequential requests... and I have no confidence they
267.
▲
by
coder543
2y ago
I’m not an expert on at-scale inference, but they surely can’t have been running at a batch size of more than 1 if they were getting performance that bad on 4xH100… and I’m not even sure how they were getting performance that low even at ba
268.
▲
by
coder543
2y ago
They’re all listed here: https://ollama.com/library/deepseek-r1/tags
269.
▲
by
coder543
2y ago
> By default, this downloads the main DeepSeek R1 model (which is large). If you’re interested in a specific distilled variant (e.g., 1.5B, 7B, 14B), just specify its tag No… it downloads the 7B model by default. If you think that is l
270.
▲
by
coder543
2y ago
The most cost-effective way is arguably to run it off of any 1TB SSD (~$55) attached to whatever computer you already have. I was able to get 1 token every 6 or 7 seconds (approximately 10 words per minute) on a 400GB quant of the model,
More ›