Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
parthsareen
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
parthsareen
2mo ago
Hi! From Ollama here - you can run: ollama run qwen3.8 (or if on mac qwen3.8:27b-mlx)
2.
▲
by
parthsareen
8mo ago
Also recently added ollama launch claude if you want to connect to cloud models from there :)
3.
▲
by
parthsareen
8mo ago
Hey! One of the maintainers of Ollama. 8GB of VRAM is a bit tight for coding agents since their prompts are quite large. You could try playing with qwen3 and at least 16k context length to see how it works.
4.
▲
Building Reliable AI Agents
(parthsareen.com)
1 points
by
parthsareen
9mo ago
|
0 comments
5.
▲
by
parthsareen
10mo ago
How much ram are you running with? Qwen3 and gpt-oss:20b punch a good bit above their weight. Personally use it for small agents.
6.
▲
by
parthsareen
10mo ago
You're welcome to go through the source: https://github.com/ollama/ollama/
7.
▲
by
parthsareen
10mo ago
Desktop app is open-source now.
8.
▲
How to Train an LLM: Part 1
(omkaark.com)
20 points
by
parthsareen
11mo ago
|
3 comments
9.
▲
by
parthsareen
1y ago
Since we shipped web search with gpt-oss in the Ollama app I've personally been using that a lot more especially for research heavy tasks that I can shoot off. Plus with a 5090 or the new macs it's super fast.
10.
▲
by
parthsareen
1y ago
Hi - author of the post. Yes it does! The "build a search agent" example can be used with a local model. I'd recommend trying qwen3 or gpt-oss
11.
▲
by
parthsareen
1y ago
Hey! Author of the blogpost and I also work on Ollama's tool calling. There has been a big push on tool calling over the last year to improve the parsing. What's the issues you're running into with local tool use? What models
12.
▲
by
parthsareen
1y ago
That's a great idea. Going to try this next :)
13.
▲
by
parthsareen
1y ago
Hey! I'm the author of the post. We haven't optimized sampling yet so it's running linearly on the CPU. A lot of SOTA work either does this while the model is running the forward pass or does the masking on the GPU. The greed
14.
▲
by
parthsareen
1y ago
Thank you! Maybe not "perfect" but near-perfect is something we can expect. Models like the Osmosis structure which just structure data inspired some of that thinking ( https://ollama.com/Osmosis/Osmosis-Struct
15.
▲
by
parthsareen
1y ago
Thanks for posting! Didn't expect this to get picked up – it was a bit of a draft haha. Happy to answer questions around structured outputs :)
16.
▲
by
parthsareen
2y ago
Yes! I have checked guidance out, as well as a few others. Planning to refactor sampling in the near future which would include improving using grammars for sampling as well. Thanks for sharing!
17.
▲
by
parthsareen
2y ago
The constraints will always be met. It’s the data inside that might be inaccurate. YMMV with smaller models in that sense.
18.
▲
by
parthsareen
2y ago
Hey! Author of the blog here. The current implementation uses llama.cpp GBNF which has allowed for a quick implementation. The biggest value-add at this time was getting the feature out. With the newer research - outlines/xgrammar comi
19.
▲
by
parthsareen
2y ago
Hey! Author of the post and one of the maintainers here. I agree - we (maintainers) got to this late and in general want to encourage more contributions. Hoping to be more on top of community PRs and get them merged in the coming year.
20.
▲
by
parthsareen
2y ago
This looks really useful. Thank you!
21.
▲
by
parthsareen
2y ago
I authored the blog with some other contributors and worked on the feature (PR: https://github.com/ollama/ollama/pull/7900 ). The current implementation uses llama.cpp GBNF grammars. The more recent research (
22.
▲
by
parthsareen
2y ago
We’ve been keeping a close eye on this as well as research is coming out. We’re looking into improving sampling as a whole on both speed and accuracy. Hopefully with those changes we might also enable general structure generation not only l
23.
▲
by
parthsareen
2y ago
Hey! Author of the blog post here. Yes you should be able to use any model. Your mileage may vary with the smaller models but asking them to “return x in json” tends to help with accuracy (anecdotally).
24.
▲
Learnings from Building a Graph AI Agent Framework
(parthsareen.com)
1 points
by
parthsareen
2y ago
|
0 comments
25.
▲
by
parthsareen
5y ago
The first few of these are my fav: https://dive.sh/thread/81vD2RhjxF