Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
HenryNdubuaku
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
HenryNdubuaku
17d ago
thanks, we improved on False Negatives this time :)
2.
▲
by
HenryNdubuaku
17d ago
noted, we'd look into this, thanks
3.
▲
by
HenryNdubuaku
17d ago
thanks!, let us know if you ever build it out :)
4.
▲
by
HenryNdubuaku
17d ago
Thanks for testing Needle out! I'd be very interested in hearing more about the finetuning setup to see how we can make both the library's finetuning setup and the model better.
5.
▲
by
HenryNdubuaku
17d ago
Thanks for considering needle. Keep in mind that you can also fine-tune the model to fit your use case more. I think this illustrates the intended deployment pretty well, where both computational resources and compute credits can both be is
6.
▲
by
HenryNdubuaku
17d ago
Thank you fr this. "warm the house" now goes to the thermostat. It's fair that a more deterministic system with just action phrases would be easier to debug/interpret, but I think there is room for both a model that is t
7.
▲
by
HenryNdubuaku
17d ago
I think we might look into creating a baseline like this for our future models
8.
▲
by
HenryNdubuaku
17d ago
Thanks!
9.
▲
by
HenryNdubuaku
17d ago
The model can be sliced and perform the inference using a subset of its layers. The first 4 layers alone are 8 MB, all 20 are 29 MB. Fair point on the copy, tightened it.
10.
▲
by
HenryNdubuaku
17d ago
Thank you!
11.
▲
by
HenryNdubuaku
17d ago
I think that would be very useful for us! The best way to reach us is through the founders@cactuscompute.com email Thank you!
12.
▲
by
HenryNdubuaku
17d ago
This is extremely useful feedback for us, thanks! I think the easiest thing here that can be fixed with tool definitions is the number conversions. Additionally, the model tends to work better with fewer tools. We will definitely be focusin
13.
▲
by
HenryNdubuaku
17d ago
Thanks for the feedback! Implications and relations are hard for the model to understand (things like go to the living room, then the kitchen, and back), so yes the cleanest use cases involve direct language. Reasoning isn't true reas
14.
▲
by
HenryNdubuaku
17d ago
That's a really good point and I think it's not yet clear how well, say, 8-30MB worth of regexs with accompanying algorithmic structure would do on these tasks. I would imagine they do quite well on a well defined task, but it wou
15.
▲
by
HenryNdubuaku
17d ago
Thanks! Certainly giving the model more context on the task it needs to perform would help it. This was actually a part of training that we improved going from Needle 2 to Needle 3
16.
▲
by
HenryNdubuaku
17d ago
As far as I know Jev's architecture isn't public (though I might be mistaken!), but it is a coincidence :) Needle 3 has been in the making since Needle 2 launched early august, but we are very excited that Jev is bringing more att
17.
▲
by
HenryNdubuaku
17d ago
well certainly the environment on the website cannot be a full product, and it isn't claiming to be that. The model, while capable in many dimensions, is also limited by its size. The website is meant to show both the capabilities and
18.
▲
by
HenryNdubuaku
17d ago
Hey there! If you end up trying out needle on the test suite it would be very useful for us if you could share some failure modes of the model! We are always trying to understand where the model isn't doing good and where we can make i
19.
▲
by
HenryNdubuaku
17d ago
Hey, thanks for the feedback! I think this is a useful part of a demonstration so I added a 911 tool specifically to demonstrate this capability and the fact that you can guard it with triggers that make it so calling emergency is an unambi
20.
▲
by
HenryNdubuaku
17d ago
Hey there, yep we found that on Apple devices specifically running on CPU is fast enough that Metal support is not needed. Thanks for flagging this though, and if usecases that would benefit from Metal support come up we will be adding it t
21.
▲
by
HenryNdubuaku
17d ago
Hey! Yeah I think for labelling the model would need to have much better world knowledge than its current size allows. Jev really is a very good model, I think it has a very strong place in the upcoming tech stacks. Really good suggestion t
22.
▲
by
HenryNdubuaku
17d ago
Oh yeah really good use case! Definitely something to finetune the model for so that it gains better task-specific reliability, because rebooting the wrong server could easily be catastrophic.
23.
▲
by
HenryNdubuaku
17d ago
lol well i guess you can turn the whole kitchen into an oven with needle :) But for real usecases you are able to set explicit minimum and maximum values on the output range of numeric arguments, so that you can avoid situations like these.
24.
▲
by
HenryNdubuaku
17d ago
haha i think that's a good demonstration of how external guardrails could help ground tiny models like these to prevent issues from coming up. I wouldn't trust needle to be my autopilot either (:
25.
▲
by
HenryNdubuaku
17d ago
Hey, thanks a lot for this feedback, very useful and actionable for us! Quite a few of these came down to our tool definitions in the playground as well as out triggers. We updated them just now and these should be more reliable. Really thi
26.
▲
by
HenryNdubuaku
17d ago
Hi, thanks for the feedback! And yes absolutely less tools and better tool descriptions make a huge difference for this model.
27.
▲
by
HenryNdubuaku
17d ago
Yes, it was really difficult to compress meaningfully intelligence down to that, and there are many limitations we are aware off and still improving on.
28.
▲
by
HenryNdubuaku
17d ago
100%, data was honestly most of the work, Needle 3 is trained on 360B tokens of structured data and we spend way more time on the generation pipeline than on the model. On Led Zeppelin, Needle doesn't actually need to know it, argument
29.
▲
by
HenryNdubuaku
17d ago
Makes a lot of sense! declare one record with the ~20 keys as fields, cuisine and friends as enums with the common values, and the grammar can't produce a value outside the set; fields with no evidence come back empty, so the output is
30.
▲
by
HenryNdubuaku
17d ago
presets updated for you now :)
More ›