Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
simedw
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
simedw
10mo ago
Surprisingly easy. If the language has a lot of conjugations (e.g., polite past verb forms), running each word through Snowball first makes the process a bit easier.
32.
▲
by
simedw
10mo ago
That's a really cool concept. Naively replacing words might work, but sometimes the context is needed. Maybe a model like gemini 2.5 flash lite would be fast enough but still maintain better context awareness?
33.
▲
by
simedw
10mo ago
Thanks for sharing; you clearly spent a lot of time making this easy to digest. I especially like the tokens-to-embedding visualisation. I recently had some trouble converting a HF transformer I trained with PyTorch to Core ML. I just could
34.
▲
by
simedw
10mo ago
Thanks! I think getting comfortable with characters fairly early is important, as it helps shift your mindset into the right place. That said, I don’t think this project really works until you’re comfortable with at least ~60 characters.
35.
▲
Show HN: Learning a Language Using Only Words You Know
(simedw.com)
84 points
by
simedw
10mo ago
|
29 comments
36.
▲
You Need an Ideation Channel
(simedw.com)
12 points
by
simedw
10mo ago
|
0 comments
37.
▲
by
simedw
1y ago
I noticed that I get far fewer refusals when I set my VPN to the USA.
38.
▲
by
simedw
1y ago
The model is only available in AI Studio when I set my VPN to the USA (I’m located in the UK).
39.
▲
by
simedw
1y ago
That's true, it's also why I didn't benchmark against any other model provider. It has been tuned so heavily on this specific format that even a tiny change, like switching the order in the `box_2d` format from `(ymin, xmin,
40.
▲
by
simedw
1y ago
That's really interesting, thanks for sharing! Are you using that approach in production for grounding when PDFs don't include embedded text, like in the case of scanned documents? I did some experiments for that use case, and it
41.
▲
by
simedw
1y ago
Thank you! Better training data is often the key to solving these issues, though it can be a costly solution. In some cases, running a model like SAM 2 on a loose bounding box can help refine the results. I usually add about 10% padding in
42.
▲
Is Gemini 2.5 good at bounding boxes?
(simedw.com)
280 points
by
simedw
1y ago
|
63 comments
43.
▲
by
simedw
1y ago
It was actually a bit worse than that the LLM never got the full recipe due to some truncation logic I had added. So it regurgitated the recipe from training, and apparently, it couldn't do both that and convert units at the same time
44.
▲
by
simedw
1y ago
Fantastic catch! It led me down a rabbit hole, and I finally found the root cause. The recipe site was so long that it got truncated before being sent to the LLM. Then, based on the first 8000 characters, Gemini hallucinated the rest of the
45.
▲
by
simedw
1y ago
You are right, it's Gemini 2.5 Flash Lite
46.
▲
by
simedw
1y ago
If the goal is to have a more consistent layout on each visit, I think we could save the last page's markdown and send it to the model as a one-shot example...
47.
▲
by
simedw
1y ago
Just had a look and three is quite a lot going into Firefox's reader mode. https://github.com/mozilla/readability
48.
▲
by
simedw
1y ago
We’ll probably have to add some custom code to log in, get an auth token, and then browse with it. Not sure if LinkedIn would like that, but I certainly would.
49.
▲
by
simedw
1y ago
That's a cool project. I think most of it comes down to Flash-Lite being really fast, and the fact that I'm only outputting markdown, which is fairly easy and streams well.
50.
▲
by
simedw
1y ago
Thank you. I was thinking of showing multiple tabs/views at the same time, but only from the same source. Maybe we could have one tab with the original content optimised for cli viewing, and another tab just doing fact checking (can gr
51.
▲
Show HN: Spegel, a Terminal Browser That Uses LLMs to Rewrite Webpages
(simedw.com)
426 points
by
simedw
1y ago
|
180 comments
52.
▲
by
simedw
2y ago
Very interesting breakdown, OP have you deep dived in pgvectorscale as well?
53.
▲
by
simedw
4y ago
After transcribing the speech into text, it seems to first translate it into English before send it over the GPT-3 (davinci). I wonder if that will lead to a bunch of mistakes, can't GPT-3 handle Mandarin directly? I did some experimen
54.
▲
Keynote: Clarity – Saša Jurić – ElixirConf EU 2021
(youtube.com)
3 points
by
simedw
5y ago
|
0 comments
55.
▲
by
simedw
6y ago
Useful tool! Do you not need to escape the ampersand in the second example? otherwise bash will run it in the background and skip the `filters=t2,m4` part.
56.
▲
Operating Systems via Development
(theerlangelist.com)
3 points
by
simedw
6y ago
|
0 comments
57.
▲
by
simedw
13y ago
Pretty cool but the LinkedIn Card gives me the following PHP exception: Catchable fatal error: Object of class stdClass could not be converted to string in /home/stephenou/socialbusinesscard.me/card_linkedin.php on line