Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
coder543
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
61.
▲
by
coder543
6mo ago
There are issues with the chat template right now[0], so tool calling does not work reliably[1]. Every time people try to rush to judge open models on launch day... it never goes well. There are ~always bugs on launch day. [0]: https:/
62.
▲
by
coder543
6mo ago
For the many DGX Spark and Strix Halo users with 128GB of memory, I believe the ideal model size would probably be a MoE with close to 200B total parameters and a low active count of 3B to 10B. I would personally love to see a super sparse
63.
▲
by
coder543
6mo ago
> Wild differences in ELO compared to tfa's graph Because those are two different, completely independent Elos... the one you linked is for LMArena, not Codeforces.
64.
▲
by
coder543
6mo ago
If you visit https://chatgpt.com/codex/settings/connectors , you're saying you don't have GitHub connected? Plugins are a new feature as of this past week, so Codex "helpfully" installs the GitH
65.
▲
by
coder543
6mo ago
That Codex one comes from the new `github` plugin, which includes a `github:yeet` skill. There are several ways to disable it: you can disconnect github from codex entirely, or uninstall the plugin, or add this to your config.toml: [[
66.
▲
by
coder543
7mo ago
From my point of view, Parakeet is not very good at formatting the output, so it would be nice if a small model focused on having nicely formatted (and correct) text, not just the lowest WER score. Rewarding the model for inserting logical
67.
▲
by
coder543
7mo ago
I do not see prompt processing, only some kind of nebulous “throughput” that could be output or input+output, but definitely not input only.
68.
▲
by
coder543
7mo ago
I wish someone would also thoroughly measure prompt processing speeds across the major providers too. Output speeds are useful too, but more commonly measured.
69.
▲
by
coder543
7mo ago
Parakeet is insanely fast and much more accurate, and it doesn't really matter that Whisper requires hacks to work live when those hacks have existed for years and work great. (The Hello Transcribe app on iOS is a great example of ho
70.
▲
by
coder543
7mo ago
Neutral here means not strongly identifiable as any particular regional American accent. Some people have very strong regional accents, some don’t. It is still clearly an American accent, not British or anything else.
71.
▲
by
coder543
7mo ago
“A few years ago” sounds like it could be before the modern era of STT, as defined by when Whisper was released. Your comment says TTS, which is different from what I’m discussing, though, so there might be some confusion.
72.
▲
by
coder543
7mo ago
Terrible relative to everything else that exists today. I have a neutral American accent. Maybe you just don’t know what you’re missing? Google’s default speech to text is still bad compared to Whisper and Parakeet, but even Google’s is mar
73.
▲
by
coder543
7mo ago
> Siri/iOS-Dictation is truly good when it comes to understanding the speech. What...? It is terrible, even compared to Whisper Tiny , which was released years ago under an Apache 2.0 license so Apple could have adopted it instan
74.
▲
by
coder543
7mo ago
Why on earth would Anthropic commit to interoperability? That is the company that doesn't interoperate with the standard LLM APIs that OpenAI developed, which everyone else in the industry has adopted and uses. Whether OpenAI's AP
75.
▲
by
coder543
7mo ago
The A18 Pro performs about on par with an M4 in terms of single threaded performance, and a little better than M1 in terms of multi threaded performance. The MacBook Neo has one of the fastest processors on the market for single threaded ta
76.
▲
by
coder543
7mo ago
No... benchmarks are not always "fishy." That is just a defense people use when they have nothing else to point to. I already said the benchmarks aren't perfect, but they are much better than claiming vibes are a more objecti
77.
▲
by
coder543
7mo ago
I would not say a full year... not even close to a year: GLM-5 is very close to the frontier: https://artificialanalysis.ai/ Artificial Analysis isn't perfect, but it is an independent third party that actually runs th
78.
▲
by
coder543
8mo ago
> Are there leaderboards that you follow or trust? Not for OCR. Regardless of how much some people complain about them, I really do appreciate the effort Artificial Analysis puts into consistently running standardized benchmarks for LLMs
79.
▲
by
coder543
8mo ago
If someone releases a benchmark/dataset, I'm sure that significantly increases the chances of one of these AI labs training on the task.
80.
▲
by
coder543
8mo ago
It is missing both models that I mentioned, so yes, I would say one reason it is not accurate is because it is so incomplete. It also doesn't provide error bars on the ELO, so models that only have tens of battles are being listed alon
81.
▲
by
coder543
8mo ago
It's not that they can't do multiple pages... but did you compare against doing one page at a time? How many pages did you try in a single request? 5? 50? 500? I fully believe that 5 pages of input works just fine, but this does
82.
▲
by
coder543
8mo ago
If you want OCR with the big LLM providers, you should probably be passing one page per request. Having the model focus on OCR for only a single page at a time seemed to help a lot in my anecdotal testing a few months ago. You can even pass
83.
▲
by
coder543
8mo ago
There are a bunch of new OCR models. I’ve also heard very good things about these two in particular: - LightOnOCR-2-1B: https://huggingface.co/lightonai/LightOnOCR-2-1B - PaddleOCR-VL-1.5: https://huggingfac
84.
▲
by
coder543
8mo ago
> Do you have experience with that model No, I just heard about it this morning.
85.
▲
by
coder543
8mo ago
The diarization is on Voxtral Mini Transcribe V2, not Voxtral Mini 4B.
86.
▲
by
coder543
8mo ago
MoEs can be efficiently split between dense weights (attention/KV/etc) and sparse (MoE) weights. By running the dense weights on the GPU and offloading the sparse weights to slower CPU RAM, you can still get surprisingly decent pe
87.
▲
by
coder543
8mo ago
I wouldn't say "everyone" uses Air. I had never even heard of it, despite frequently developing in Go for close to a decade at this point. Modd is quite nice: https://github.com/cortesi/modd But, why cou
88.
▲
by
coder543
8mo ago
An LLM that can't understand the environment properly can't properly reason about which command to give in response to a user's request. Even if the LLM is a very inefficient way to pilot the thing, being able to pilot mean
89.
▲
by
coder543
8mo ago
Waymo is certainly interested in using LLMs/VLMs for this purpose. https://waymo.com/research/emma/ https://waymo.com/blog/2024/10/introducing-emma https://waymo.com
90.
▲
by
coder543
9mo ago
Another one is Soprano-1.1. It seems like it is being trained by one person, and it is surprisingly natural for such a small model. I remember when TTS always meant the most robotic, barely comprehensible voices. https://www.redd
More ›