Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
aliljet
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
91.
▲
by
aliljet
1y ago
I wonder how this can be used with an LLM to provide interesting tax advice? I'd love to regularly ask questions of the tax code...
92.
▲
by
aliljet
1y ago
What is the use case for these tiny models? Is it speed? Is it to move on device somewhere? Or is it to provide some relief in pricing somewhere in the API? It seems like most use is through the Claude subscription and therefore the use cas
93.
▲
by
aliljet
1y ago
These benchmarks in real world work remain remarkably weak. If you're using this for day-to-day work, the eval that really matters is how the model handles a ten step action. Context and focus are absolutely king in real world work. To
94.
▲
by
aliljet
1y ago
For me? It's simple. Completely empty the context and rebuild focused on the new task at hand. It's painful, but very effective.
95.
▲
by
aliljet
1y ago
There's a misunderstanding here broadly. Context could be infinite, but the real bottleneck is understanding intent late in a multi-step operation. A human can effectively discard or disregard prior information as the narrow window of
96.
▲
by
aliljet
1y ago
I work deep in the weeds on startups using datasets like these and I'm always curious about how people drive value from this raw data. Who is designing things that use this data for anything interesting in the United States? The whole
97.
▲
by
aliljet
1y ago
Is it? Across what metric?
98.
▲
by
aliljet
1y ago
For better or worse, I find Python beautiful. And here's an easter egg for everyone (it comes up in the Documentary too): import this Even on HN, indentation reigns king.
99.
▲
by
aliljet
1y ago
Having played a LOT with browser use, playwright, and puppeteer (all via MCP integrations and pythonic test cases), it's incredibly clear how quickly Claude (in particular) loses the thread as it starts to interact with the browser. Th
100.
▲
by
aliljet
1y ago
I don't want to be insulting here, but have you sat down with a partner at a VC before? You may be surprised to discover their skill is rarely deeply technical...
101.
▲
by
aliljet
1y ago
This article is a great way to showcase A16Z standinging head and shoulders above other VCs with REAL technical expertise in the partnership. Love reading this kind of stuff, but the article really needs a price to put this in perspective.
102.
▲
by
aliljet
1y ago
How do you know that?
103.
▲
by
aliljet
1y ago
This is definitely one of my CORE problem as I use these tools for "professional software engineering." I really desperately need LLMs to maintain extremely effective context and it's not actually that interesting to see a n
104.
▲
by
aliljet
1y ago
For the simple minded among us, can someone explain why this would be worth 34.5 billion dollars? Wouldn't the fork ( https://www.perplexity.ai/comet ) be sufficient?
105.
▲
by
aliljet
1y ago
I'm curious what platform people are using to test GPT-5? I'm so deep into the claude code world that I'm actually unsure what the best option is outside of claude code...
106.
▲
by
aliljet
1y ago
Between Opus aand GPT-5, it's not clear there's a substantial difference in software development expertise. The metric that I can't seem to get past in my attempts to use the systems is context awareness over long-running tas
107.
▲
by
aliljet
1y ago
The eval bar I want to see here is simple: over a complex objective (e.g., deploy to prod using a git workflow), how many tasks can GPT-5 stay on track with before it falls off the train. Context is king and it's the most obvious and g
108.
▲
by
aliljet
1y ago
Honestly, you're probably right. It's quickly become a pretty weak eval, but the guy that's running that eval is excellent. I'd much rather the evals people were using to test these things looked more like classic/b
109.
▲
by
aliljet
1y ago
It's very unclear if OpenAI has been casually leaking things to create buzz, but a few days ago there was a pretty stunning pelican on a bike attempt: https://old.reddit.com/r/OpenAI/comments/1mettre/
110.
▲
by
aliljet
1y ago
Maybe this is an unpopular opinion, but it seems like Anthropic has quietly 4x'd the real cost of the Pro plan. There are 168 hours in a week, and if I'm able to (safely) bet on 40 hours of use, realistically, I just lost 75% of t
111.
▲
by
aliljet
1y ago
I see the term 'local inference' everywhere. It's an absurd misnomer without hardware and cost defined. I can also run a coal fired power plant in my backyard, but in practice, there's no reasonable way to make that econ
112.
▲
by
aliljet
1y ago
What is the likelihood today that Trump and his allies in Congress may improve the ability for H1B holders to obtain green cards? This was repeatedly raised during his campaign and his first weeks in office and it sounded like this was goin
113.
▲
by
aliljet
1y ago
I'm a little confused about how pricing works here. What is an 'agentic interaction' and how does that translate to dollars? And how does this work with models that are differently priced???
114.
▲
by
aliljet
1y ago
If the SWE Bench results are to be believed... this looks best in class right now for a local LLM. To be fair, show me the guy who is running this locally...
115.
▲
by
aliljet
1y ago
[edit to focus on pricing, leaving praise of Simon's post out despite being deserved] Simon claims, 'Grok 4 is competitively priced. It's $3/million for input tokens and $15/million for output tokens - the same pric
116.
▲
by
aliljet
1y ago
How much of Berkshire actually relied on Buffet toward this announcement? I'm increasingly suspect of 90+ or 80+ adults operating the machinery of massive entities. Lots of examples of this.
117.
▲
by
aliljet
2y ago
I am so curious about taking over running the services that perform this for my car? Shouldn't I be able to issue commands to my car myself?
118.
▲
by
aliljet
2y ago
I genuinely wonder if this kind of use case is going to completely upend UX testing. There's something dynamic, perhaps poorly so, that would change the way prescriptive, painfully locked UX tests work...
119.
▲
by
aliljet
2y ago
I'm curious about whether anyone is running this locally using ollama?
120.
▲
by
aliljet
2y ago
This is fantastic news. I've been using Qwen2.5-Coder-32B-Instruct with Ollama locally and it's honestly such a breathe of fresh air. I wonder if any of you have had a moment to try this newer context length locally? BTW, I fail t
More ›