Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
alach11
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
61.
▲
by
alach11
1y ago
Ironically, every paper published about monitoring chain-of-thought reduces the likelihood of this technique being effective against strong AI models.
62.
▲
by
alach11
1y ago
A lower-key variant of this frequently comes into play with consulting or other sales pitches. "You spend <big number> per year on this <necessary business expense>. Our service will easily shave 2% off this, making the cos
63.
▲
by
alach11
1y ago
This is starting to encroach on Lovable, right? I do suspect the effect of these "vibe coded" apps on the SaaS market will be smaller than expected. Heavier-featured apps will have all sorts of functionality and polish a user won&
64.
▲
by
alach11
1y ago
You nailed it. Microsoft should have a huge advantage with depth of integration, but for some reason treats Copilot in office as a glorified chat iframe. It's a huge missed opportunity.
65.
▲
by
alach11
1y ago
Interesting to see them announce "Record Mode", where ChatGPT can listen in on your meetings and use that to inform future conversations. I suspect the idea of an ever-present all-hearing AI is going to be a growing trend. Especia
66.
▲
Updates to ChatGPT for Business (internal connectors, record mode, and more) [video]
(youtube.com)
2 points
by
alach11
1y ago
|
1 comments
67.
▲
by
alach11
1y ago
I'm trying to think of practical use cases for this. Is surveillance one of them? Could you drop a bunch of these marked tardigrades on an object or on money and later identify it? Doesn't seem like a very efficient way to accompl
68.
▲
by
alach11
1y ago
I think this is really sick. And it's against the mission of Meta. It's possible, despite what many of these comments say, to build friendships as an adult. It just takes investment. What has helped me the most is: - Bringing conn
69.
▲
by
alach11
2y ago
This is basically a big dunk on OpenAI, right? OpenAI made a big show out of hiding their reasoning traces and using them for alignment purposes [0]. Anthropic has demonstrated (via their mech interp research) that this isn't a reliabl
70.
▲
by
alach11
2y ago
These business can easily Anthropic models through AWS Bedrock. All it requires it a simple clickthrough EULA. That's what we do at the F500 non-tech company where I work. The same is true with OpenAI models in Azure. I can't imag
71.
▲
by
alach11
2y ago
Fascinating papers. Could deliberately suppressing memorization during pretraining help force models to develop stronger first-principles reasoning?
72.
▲
by
alach11
2y ago
It's incredible that this took 316 days to be released since it was initially announced. I do appreciate the emphasis in the presentation on how this can be useful beyond just being a cool/fun toy, as it seems most image generatio
73.
▲
by
alach11
2y ago
It's incredible that this took 316 days to be released since it was initially announced. I do appreciate the emphasis in the presentation on how this can be useful beyond just being a cool/fun toy, as it seems most image generat
74.
▲
by
alach11
2y ago
The realtime API can be used to call tools [0], but I agree with your general point on the flexibility of working directly with text. [0] https://github.com/openai/openai-realtime-agents
75.
▲
by
alach11
2y ago
It's interesting that they pitch this for agent development. The realtime API provides a much simpler architecture for developing agents. Why would you want to string together STT -> LLM -> TTS when you could have a consolidated
76.
▲
by
alach11
2y ago
> Can't you just [store the power] If you find a cheap solution to power storage, you can make a lot of money. This is the key enabler for a 100% renewable grid.
77.
▲
Is It an AWS EC2 Instance or a US Visa?
(rahmatashari.com)
39 points
by
alach11
2y ago
|
6 comments
78.
▲
by
alach11
2y ago
Even OpenAI's cookbook is out of date in many places. This is a really difficult and common problem!
79.
▲
by
alach11
2y ago
Very interesting announcement. GPT-4.5 being the last non-reasoning model seems like a tacit admission that the old scaling paradigms are exhausted. Also, I wonder if any of these decisions around GPT-5 are intended to make it harder for co
80.
▲
by
alach11
2y ago
Yes, those are the same.
81.
▲
by
alach11
2y ago
I've been switching back and forth a bit at work recently, and I find Cursor still has a slight edge.
82.
▲
by
alach11
2y ago
If we assume distillation remains viable, the game theory implications are huge. It’s going to shift the market of how foundation models are used. Companies creating models will be incentivized to vertically integrate, owning the full stack
83.
▲
by
alach11
2y ago
Sounds like it's time for someone to set up an email service that offers the same functionality without the +. There would be some headaches and it would limit the degrees of freedom users have with base email addresses, but I'd u
84.
▲
by
alach11
2y ago
> the approach where "agents" accomplish things by using the browser/desktop always seemed off to me It's certainly a much more difficult approach, but it scales so much better. There's such a long-tail of small
85.
▲
by
alach11
2y ago
Make sure to check out their system card [0]. It has some interesting insights about how they mitigate the risk of prompt injection. There's a separate "Supervisor" model watching the Operator and looking out for prompt injec
86.
▲
by
alach11
2y ago
I don't know if I'm ready to hand over my grocery shopping (or date night planning) to an agent. But if pricing is reasonable, this could be a powerful alternative to normal RPA. Instead of hardcoding some automation using Seleniu
87.
▲
by
alach11
2y ago
This is a very impressive result. OpenAI was able to achieve 72% with o3, but that's at a very high compute cost at inference-time. I'd be interested for Aide to release more metrics on token counts, total expenditure, etc. to bet
88.
▲
by
alach11
2y ago
This is completely logical from a compensation/effort maximization perspective. But I find it deeply unfulfilling and could never work like this. If I'm spending 1/3 of my conscious hours on something, I want to feel like it
89.
▲
by
alach11
2y ago
> Yemeni Coffee Shops in Texas Here in Houston, law enforcement is pretty strict on homeless people causing problems. Maybe that's the reason?
90.
▲
by
alach11
2y ago
> I don't think any of the big boys are working on how to get an LLM to design a better LLM Not sure if you count this as "working on it", but this is something Anthropic tests for for safety evals on models. "If a mo
More ›