3 ms·
Earlier today I was playing around with the "Union Alpha" stealth model (which I guess exited stealth later in the evening), and I noticed it had a habit of try
by saghm 10d ago
Earlier today I was playing around with the "Union Alpha" stealth model (which I guess exited stealth later in the evening), and I noticed it had a habit of trying to respond to the subagents it spawned while giving me an answer. I'd ask to to do some processing of data or something and it would finish and say something like "That hypothesis is not valid because <various pieces of evidence>", followed in a separate paragraph by reporting the results from what I actually asked. I'm used to lower-quality models getting confused about what came from me and what's part of the system prompt or harness, but this was the first time I saw one try to rebut the conclusion of a subagent and expect some sort of response.
- Bluestein 10d agoQuite the model I found this one to be. Disappointed when the trial ended.- PS: It would be ground breaking if it turns out to have been using Chinese chips for inference, like Stealth Ox Alpha. Unlikely though.-
- saghm 9d agoYeah, it seemed pretty good. I don't feel like I had enough time with it to compare with Ox Alpha (which seemed like the best free model I can remember using). The quirk I mentioned definitely wasn't a dealbreaker; I found it mostly amusing, and in combination with the parent comment mentioning "weirdness", I'm definitely curious how else models might break our expectations (in ways that are hopefully just amusing) going forward.
- Bluestein 9d agoAgree totally on Ox Alpha. It felt like pre-castration Fable.- Turned out to be these guys: - https://news.ycombinator.com/item?id=49751723 https://news.ycombinator.com/item?id=49751723
- saghm 9d agoYep, I saw last night when a few hours into using it I got a this response: > Error: Thank you for participating in the Stealth Union Alpha testing period. This model was Unbiased's Pareto. One of my friends quipped that "Unbiased Pareto" still sounded like the name of a stealth model. I'm not sure if I just noticed it later, but this definitely seemed to be a lot shorter than other stealth alphas I've tried. I wouldn't be shocked if this is more typical going forward though, or if stealth alphas entirely go away, since it certainly costs a bit of money to market this way.
- Bluestein 9d agoThis one was more "fly by night" than "stealth" :) We got drive-by modelled.-
- saghm 9d agoGiven that I don't pay for any subscriptions and just coast on the free tiers of OpenCode and OpenRouter (along with some judicious use of llama.cpp locally when things are scoped well enough enough for a local model), I'm fine with this. I had the $20/month Claude one for a few months starting in February, but one day it randomly started returning me errors claiming I needed to pay for more credits despite the usage showing 8% for the week and 20% for the session, and I figured if they couldn't even communicate to me the difference between them screwing up the check for hitting the limit or an outage, it wasn't worth it for me to keep paying them. I dislike OpenAI too much to want to pay them any of my personal money for anything, and when I tried out Mistral Vibe it did not work very well for me (it kept not following instructions and eventually when I kept trying to push it to handle things better it somehow spiraled into simulating some sort of existential crisis, culminating in gibberish and random characters being dumped on my screen infinitely until I killed the process; incredibly entertaining, but not worth paying for)
- Bluestein 9d ago> I figured if they couldn't even communicate to me the difference between them screwing up the check for hitting the limit or an outage, it wasn't worth it for me to keep paying them Sensible.-